RESEARCH · RESEARCH · #1417
Measuring the microtask eligibility gap for off-the-shelf SLMs (arXiv:2610.00025v1)
This paper builds a benchmark of four microtasks around LLM planners (auto-approving commands, writing memory, tool selection, ranking turns) with pre-specified, CI-backed eligibility thresholds and tests off-the-shelf small language models (Qwen3 0.6/1.7/4/8B, greedy FP16, no tuning). Across 16 model×task configurations, none met the practitioner-defined thresholds (0/16). The authors introduce a logprob decision-threshold diagnostic separating capability deficits from decoding-threshold failures, show 4-bit quantization methods (RTN/GPTQ/AWQ) generally worsen performance without enabling eligibility, replicate the gap on Llama-3.x (12/12 ineligible), and show results are robust to anchor choice and prompt paraphrases; they recommend placing SLMs behind a baseline that already meets the CI-backed threshold and only using SLMs where that baseline fails.
KEY POINTS
- This paper builds a benchmark of four microtasks around LLM planners (auto-approving commands, writing memory, tool selection, ranking turns) with pre-specified, CI-backed eligibility thresholds and tests off-the-shelf small language models (Qwen3 0.6/1.7/4/8B, greedy FP16, no tuning).
- Across 16 model×task configurations, none met the practitioner-defined thresholds (0/16).
- The authors introduce a logprob decision-threshold diagnostic separating capability deficits from decoding-threshold failures, show 4-bit quantization methods (RTN/GPTQ/AWQ) generally worsen performance without enabling eligibility, replicate the gap on Llama-3.x (12/12 ineligible), and show results are robust to anchor choice and prompt paraphrases; they recommend placing SLMs behind a baseline that already meets the CI-backed threshold and only using SLMs where that baseline fails.
WHY IT MATTERS
This matters because practitioners embedding SLMs into agent harnesses should not assume off-the-shelf small models meet operational thresholds: the paper quantifies where they fail and how quantization and decoding thresholds affect eligibility.