Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1326

NAQD-Env: benchmark for selective withdrawal in language agents

The paper (arXiv:2609.38460v1) introduces NAQD-Env, a synthetic deterministic benchmark that evaluates language-agent decisions to withdraw, preserve, and resume actions across 11 dependency families and explicit evidence/authorization inputs. Evaluating three open-weight instruction-tuned models (Qwen2.5-7B, Qwen2.5-3B, Llama-3.1-8B) under three prompt conditions on 350 scenarios (3,150 episodes) finds very low withdrawal recall (≤0.06), no valid resumptions at eligible opportunities, and only one episode matching the full reference policy; supervised fine-tuning raised Qwen2.5-3B decision accuracy from ~0.45–0.54 to ~0.83–0.92 but introduced other faults.

KEY POINTS

  1. The paper (arXiv:2609.38460v1) introduces NAQD-Env, a synthetic deterministic benchmark that evaluates language-agent decisions to withdraw, preserve, and resume actions across 11 dependency families and explicit evidence/authorization inputs.
  2. Evaluating three open-weight instruction-tuned models (Qwen2.5-7B, Qwen2.5-3B, Llama-3.1-8B) under three prompt conditions on 350 scenarios (3,150 episodes) finds very low withdrawal recall (≤0.06), no valid resumptions at eligible opportunities, and only one episode matching the full reference policy; supervised fine-tuning raised Qwen2.5-3B decision accuracy from ~0.45–0.54 to ~0.83–0.92 but introduced other faults.
  3. Selective withdrawal is a distinct reliability capability for agents; NAQD-Env exposes systematic failures (very low withdrawal/resumption rates) that matter for safe agent behavior and evaluation.

WHY IT MATTERS

Selective withdrawal is a distinct reliability capability for agents; NAQD-Env exposes systematic failures (very low withdrawal/resumption rates) that matter for safe agent behavior and evaluation.

SOURCES & TIMELINE

1