Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #12410

BioMysteryBench

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

Profession-specific 'Scientific Agents' prompts raise cost and tokens without improving accuracy

Researchers evaluated the open-source Scientific Agents corpus (503 profession-specific AGENTS.md profiles) using Gemini 3.8 Flash via OpenRouter in the Pi agent harness across nine text-based science benchmarks and 60 tool-using BioMysteryBench problems. Matched profession profiles showed no clear accuracy gain (mean profile–baseline difference −0.6 percentage points, 95% bootstrap interval [−1.5, +0.2]), produced 1.5–2.3× more output tokens and cost 2.2–4.5× more per successful call, and performed worse on BioMysteryBench (46.7% vs 56.7% solves, −10.0 pp) due to more token- and time-limit interruptions; longer prompts did improve robustness to API drops on SuperGPQA.

6.0