Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1413

Profession-specific 'Scientific Agents' prompts raise cost and tokens without improving accuracy

Researchers evaluated the open-source Scientific Agents corpus (503 profession-specific AGENTS.md profiles) using Gemini 3.8 Flash via OpenRouter in the Pi agent harness across nine text-based science benchmarks and 60 tool-using BioMysteryBench problems. Matched profession profiles showed no clear accuracy gain (mean profile–baseline difference −0.6 percentage points, 95% bootstrap interval [−1.5, +0.2]), produced 1.5–2.3× more output tokens and cost 2.2–4.5× more per successful call, and performed worse on BioMysteryBench (46.7% vs 56.7% solves, −10.0 pp) due to more token- and time-limit interruptions; longer prompts did improve robustness to API drops on SuperGPQA.

KEY POINTS

  1. Researchers evaluated the open-source Scientific Agents corpus (503 profession-specific AGENTS.md profiles) using Gemini 3.8 Flash via OpenRouter in the Pi agent harness across nine text-based science benchmarks and 60 tool-using BioMysteryBench problems.
  2. Matched profession profiles showed no clear accuracy gain (mean profile–baseline difference −0.6 percentage points, 95% bootstrap interval [−1.5, +0.2]), produced 1.5–2.3× more output tokens and cost 2.2–4.5× more per successful call, and performed worse on BioMysteryBench (46.7% vs 56.7% solves, −10.0 pp) due to more token- and time-limit interruptions; longer prompts did improve robustness to API drops on SuperGPQA.
  3. This shows that loading full profession-specific agent profiles by default can substantially raise cost and latency without improving — and sometimes reducing — task performance, guiding design choices for agent systems and prompt-retrieval strategies.

WHY IT MATTERS

This shows that loading full profession-specific agent profiles by default can substantially raise cost and latency without improving — and sometimes reducing — task performance, guiding design choices for agent systems and prompt-retrieval strategies.

SOURCES & TIMELINE

1