RESEARCH · RESEARCH · #1326
NAQD-Env: benchmark for selective withdrawal in language agents
The paper (arXiv:2609.38460v1) introduces NAQD-Env, a synthetic deterministic benchmark that evaluates language-agent decisions to withdraw, preserve, and resume actions across 11 dependency families and explicit evidence/authorization inputs. Evaluating three open-weight instruction-tuned models (Qwen2.5-7B, Qwen2.5-3B, Llama-3.1-8B) under three prompt conditions on 350 scenarios (3,150 episodes) finds very low withdrawal recall (≤0.06), no valid resumptions at eligible opportunities, and only one episode matching the full reference policy; supervised fine-tuning raised Qwen2.5-3B decision accuracy from ~0.45–0.54 to ~0.83–0.92 but introduced other faults.
KEY POINTS
- The paper (arXiv:2609.38460v1) introduces NAQD-Env, a synthetic deterministic benchmark that evaluates language-agent decisions to withdraw, preserve, and resume actions across 11 dependency families and explicit evidence/authorization inputs.
- Evaluating three open-weight instruction-tuned models (Qwen2.5-7B, Qwen2.5-3B, Llama-3.1-8B) under three prompt conditions on 350 scenarios (3,150 episodes) finds very low withdrawal recall (≤0.06), no valid resumptions at eligible opportunities, and only one episode matching the full reference policy; supervised fine-tuning raised Qwen2.5-3B decision accuracy from ~0.45–0.54 to ~0.83–0.92 but introduced other faults.
- Selective withdrawal is a distinct reliability capability for agents; NAQD-Env exposes systematic failures (very low withdrawal/resumption rates) that matter for safe agent behavior and evaluation.
WHY IT MATTERS
Selective withdrawal is a distinct reliability capability for agents; NAQD-Env exposes systematic failures (very low withdrawal/resumption rates) that matter for safe agent behavior and evaluation.