arXiv:2609.30705v1 — “The Price of Thought”: extra LLM reasoning does not reliably improve trading returns
A controlled study (arXiv:2609.30705v1) compares reasoning effort in representative LLMs from the DeepSeek, GPT, and Gemini families on one year of U.S. equities data across three input conditions (numerical, identifiable news, masked news). With over 800,000 asset predictions and repeated generations, the authors find that additional test-time reasoning does not produce a reliable improvement in net portfolio returns and can produce nonmonotonic or unstable treatment effects, motivating task-specific validation before deployment.