Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #11641

Qwen2.5-3B

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

NAQD-Env: benchmark for selective withdrawal in language agents

The paper (arXiv:2609.38460v1) introduces NAQD-Env, a synthetic deterministic benchmark that evaluates language-agent decisions to withdraw, preserve, and resume actions across 11 dependency families and explicit evidence/authorization inputs. Evaluating three open-weight instruction-tuned models (Qwen2.5-7B, Qwen2.5-3B, Llama-3.1-8B) under three prompt conditions on 350 scenarios (3,150 episodes) finds very low withdrawal recall (≤0.06), no valid resumptions at eligible opportunities, and only one episode matching the full reference policy; supervised fine-tuning raised Qwen2.5-3B decision accuracy from ~0.45–0.54 to ~0.83–0.92 but introduced other faults.

6.0