RESEARCH · RESEARCH · #935
Propose, Don't Judge (arXiv:2609.27051v1): introduce a frozen, anytime-valid referee for LLM agents in factor research
The paper proposes an architecture where language-model agents freely propose investment factors while a frozen statistical referee—untouchable by the agent—judges candidates using betting on market outcomes revealed after submission, yielding an anytime-valid false-discovery guarantee. In synthetic tests and a ten-year walk-forward on the CSI 500 the frozen referee admitted 5–11× fewer sub-threshold factors under a scripted proposer, language-model proposers produced higher yield and custom probes, but certified true factors typically waited ~500 trading days and certified portfolios had lower realized Sharpe than ungated portfolios.
KEY POINTS
- The paper proposes an architecture where language-model agents freely propose investment factors while a frozen statistical referee—untouchable by the agent—judges candidates using betting on market outcomes revealed after submission, yielding an anytime-valid false-discovery guarantee.
- In synthetic tests and a ten-year walk-forward on the CSI 500 the frozen referee admitted 5–11× fewer sub-threshold factors under a scripted proposer, language-model proposers produced higher yield and custom probes, but certified true factors typically waited ~500 trading days and certified portfolios had lower realized Sharpe than ungated portfolios.
- This matters because it separates proposing (agent creativity) from judging (statistical certification), offering a provable anytime-valid control that reduces false discoveries in automated factor research while quantifying the time/return trade-off.
WHY IT MATTERS
This matters because it separates proposing (agent creativity) from judging (statistical certification), offering a provable anytime-valid control that reduces false discoveries in automated factor research while quantifying the time/return trade-off.