Paper

Agents can propose investment factors, but a frozen referee has to judge them

A working paper splits factor research in two and finds that a referee the agent cannot touch admits 5 to 11 times fewer false factors.

AI Fin ResearchCovers ChinaWorking paper, not peer reviewed

Bo Qu, Mingguang Chen and Licheng Wang ask which parts of quantitative factor research a language-model agent should be allowed to run. Their answer is to let the agent propose factors and have a frozen statistical referee, one the agent cannot modify, decide which are admitted. The referee scores each candidate only on market outcomes revealed after submission, using a betting procedure, so its false-discovery guarantee holds at any stopping time and for any proposal policy.

They cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones. The tests run in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Under a scripted proposer, the frozen referee admits 5 to 11 times fewer sub-threshold factors than the leaky referees, and no proposer closes that gap. The language model beats the script on yield and matches the bandit.

The frozen referee is slow. An admitted true factor waits about 500 trading days, so the certified portfolio’s Sharpe ratio trails an ungated one. The market evidence comes from a single index.

Sources

Related

Analysis

Throw out what you knew about AI before September

Fable 5.1, GPT-6 Astra and Opus 5.5 arrived inside five weeks. One science benchmark doubled in two months. A test you ran in the spring describes a different technology.

Anthropic, Introducing Claude Opus 5.5Global