Agents can propose investment factors, but a frozen referee has to judge them
A working paper splits factor research in two and finds that a referee the agent cannot touch admits 5 to 11 times fewer false factors.
Bo Qu, Mingguang Chen and Licheng Wang ask which parts of quantitative factor research a language-model agent should be allowed to run. Their answer is to let the agent propose factors and have a frozen statistical referee, one the agent cannot modify, decide which are admitted. The referee scores each candidate only on market outcomes revealed after submission, using a betting procedure, so its false-discovery guarantee holds at any stopping time and for any proposal policy.
They cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones. The tests run in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Under a scripted proposer, the frozen referee admits 5 to 11 times fewer sub-threshold factors than the leaky referees, and no proposer closes that gap. The language model beats the script on yield and matches the bandit.
The frozen referee is slow. An admitted true factor waits about 500 trading days, so the certified portfolio’s Sharpe ratio trails an ungated one. The market evidence comes from a single index.
Sources
- arXiv arxiv.org