Propose, Don't Judge (arXiv:2609.27051v1): introduce a frozen, anytime-valid referee for LLM agents in factor research
The paper proposes an architecture where language-model agents freely propose investment factors while a frozen statistical referee—untouchable by the agent—judges candidates using betting on market outcomes revealed after submission, yielding an anytime-valid false-discovery guarantee. In synthetic tests and a ten-year walk-forward on the CSI 500 the frozen referee admitted 5–11× fewer sub-threshold factors under a scripted proposer, language-model proposers produced higher yield and custom probes, but certified true factors typically waited ~500 trading days and certified portfolios had lower realized Sharpe than ungated portfolios.