Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1417

Measuring the microtask eligibility gap for off-the-shelf SLMs (arXiv:2610.00025v1)

This paper builds a benchmark of four microtasks around LLM planners (auto-approving commands, writing memory, tool selection, ranking turns) with pre-specified, CI-backed eligibility thresholds and tests off-the-shelf small language models (Qwen3 0.6/1.7/4/8B, greedy FP16, no tuning). Across 16 model×task configurations, none met the practitioner-defined thresholds (0/16). The authors introduce a logprob decision-threshold diagnostic separating capability deficits from decoding-threshold failures, show 4-bit quantization methods (RTN/GPTQ/AWQ) generally worsen performance without enabling eligibility, replicate the gap on Llama-3.x (12/12 ineligible), and show results are robust to anchor choice and prompt paraphrases; they recommend placing SLMs behind a baseline that already meets the CI-backed threshold and only using SLMs where that baseline fails.

KEY POINTS

  1. This paper builds a benchmark of four microtasks around LLM planners (auto-approving commands, writing memory, tool selection, ranking turns) with pre-specified, CI-backed eligibility thresholds and tests off-the-shelf small language models (Qwen3 0.6/1.7/4/8B, greedy FP16, no tuning).
  2. Across 16 model×task configurations, none met the practitioner-defined thresholds (0/16).
  3. The authors introduce a logprob decision-threshold diagnostic separating capability deficits from decoding-threshold failures, show 4-bit quantization methods (RTN/GPTQ/AWQ) generally worsen performance without enabling eligibility, replicate the gap on Llama-3.x (12/12 ineligible), and show results are robust to anchor choice and prompt paraphrases; they recommend placing SLMs behind a baseline that already meets the CI-backed threshold and only using SLMs where that baseline fails.

WHY IT MATTERS

This matters because practitioners embedding SLMs into agent harnesses should not assume off-the-shelf small models meet operational thresholds: the paper quantifies where they fail and how quantization and decoding thresholds affect eligibility.

SOURCES & TIMELINE

1