RESEARCH · RESEARCH · #1521
Paired evaluation finds hosted Jev outperforms open-weight Laya on most System-1 agent decisions; audit revises claimed savings
The arXiv paper (arXiv:2610.02267v1) presents a paired, reproducible evaluation of two System‑1 decision models for agent harnesses—open-weight Laya and hosted Jev—across 11 decision points with 7,283 base cases and 6,640 robustness variants. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 percentage points); neither model beats chance on zero‑shot model routing and they tie on RAG relevance gating. The authors also self‑audit their pipeline and correct earlier deployment claims: a reported 23.9% pre-screen saving becomes 4.3% after accounting for omitted costs, gate accuracy was conflated with end‑to‑end quality (58% vs. 98%), thresholds were tuned in‑sample (held‑out misses up to 17%), and a channel‑native effect altered injection false‑positive rates; all raw outputs and analysis code are published at the linked GitHub repo.
KEY POINTS
- The arXiv paper (arXiv:2610.02267v1) presents a paired, reproducible evaluation of two System‑1 decision models for agent harnesses—open-weight Laya and hosted Jev—across 11 decision points with 7,283 base cases and 6,640 robustness variants.
- Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 percentage points); neither model beats chance on zero‑shot model routing and they tie on RAG relevance gating.
- The authors also self‑audit their pipeline and correct earlier deployment claims: a reported 23.9% pre-screen saving becomes 4.3% after accounting for omitted costs, gate accuracy was conflated with end‑to‑end quality (58% vs.
WHY IT MATTERS
This matters because System‑1 decision models are proposed to cut cost and latency in LLM agent pipelines, and the paper shows significant model performance differences, failure modes, and audit corrections that affect real deployment savings and safety assessments.