Paired evaluation finds hosted Jev outperforms open-weight Laya on most System-1 agent decisions; audit revises claimed savings
The arXiv paper (arXiv:2610.02267v1) presents a paired, reproducible evaluation of two System‑1 decision models for agent harnesses—open-weight Laya and hosted Jev—across 11 decision points with 7,283 base cases and 6,640 robustness variants. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 percentage points); neither model beats chance on zero‑shot model routing and they tie on RAG relevance gating. The authors also self‑audit their pipeline and correct earlier deployment claims: a reported 23.9% pre-screen saving becomes 4.3% after accounting for omitted costs, gate accuracy was conflated with end‑to‑end quality (58% vs. 98%), thresholds were tuned in‑sample (held‑out misses up to 17%), and a channel‑native effect altered injection false‑positive rates; all raw outputs and analysis code are published at the linked GitHub repo.