Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #8437

SWE-bench Verified

Related event timeline, sources and context from the news index.

EVENT TIMELINE

4

RESEARCH · 1 SOURCE · Microsoft Research

Microsoft Research Asia open-sources Agent Lightning v1.0, a lightweight harnessed agentic RL framework

Microsoft Research Asia released Agent Lightning v1.0, a roughly 3,500-line open-source implementation of a Harnessed Agentic RL training paradigm that trains on the same agent harness used in deployment via an LLM proxy. The framework runs agents as native Kubernetes jobs, avoids reimplementing agent harnesses, and includes an end-to-end coding-agent example that raised Qwen3.5-9B’s Pass@1 on SWE-bench Verified from 41.8% to 56.4% using about 6,000 training samples.

7.0

RESEARCH · 1 SOURCE · Apple Machine Learning Research

SCLATE: a substrate for continual-learning agent training and evaluation

SCLATE is an execution substrate that unifies scheduling for benchmarks and unmodified agents via an open event scheduler and a hybrid simulated clock, compressing long multi-session scenarios into shorter runtime and recording every model call through an in-container proxy. The authors ported seven benchmarks and ran ten harness/memory configurations across ten models, finding that added memory systems do not consistently outperform native harness memory; they also post-trained Qwen3.5-4B using unmodified harnesses and memory, which reduced file reads, raised SWE-bench Verified pass rate by 16.7 points, and increased held-out MetaClaw accuracy by up to 11.8 points.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Study identifies frequent cost-inefficient behaviors in coding agents and evaluates mitigations

arXiv:2609.30725v1 analyzes 1,200 trajectories from Claude Code and Mini-SWE-Agent on SWE-bench Verified, identifying three recurring cost-inefficient behaviors (subsumed retrieval, similar script generation, test re-execution) that affect 79–98% of tasks and can account for up to 22.75% of task cost. The paper evaluates three mitigations (structure-aware retrieval, agent-synthesized skills, developer-designed skills) over 10k held-out trajectories and reports that structure-aware retrieval can sometimes increase costs (up to 28.14%), agent-synthesized skills give limited, low-level gains, while developer-designed skills reduce costs by up to 41.73%, about twice the maximum gain from agent-synthesized skills.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Privileged Self-Practice (PSP) for multi-turn LLM agents (arXiv:2609.29051v1)

This arXiv preprint identifies a failure mode of on-policy self-distillation (OPSD) in multi-turn agents—training yields overconfident behavior without the underlying information—and proposes Privileged Self-Practice (PSP). PSP keeps privileged information (PI) in the prompt/sampler rather than the loss: when a student fails rollouts, an analyzer model injects a short per-task instruction into the prompt, the task is re-sampled, and the result is trained with the same GRPO objective; across AppWorld and SWE-bench Verified and three student models PSP consistently outperforms OPSD and plain GRPO, improving task-goal completion by up to 65% on AppWorld and resolved rate by up to 61% on SWE-bench Verified.

7.0