Self-evolving research agents show no consistent gain from accumulated skills in a factor test
Across 18 long-horizon alpha-research runs and 48 continuation branches, evolved capabilities did not reliably beat the starting set.
Siyuan Li and coauthors introduce EverMine, a framework for testing whether self-evolving agents get better at research as they accumulate skills, tools and rules. The setting is long-horizon alpha discovery, where each factor added to a portfolio changes the value of the next candidate.
EverMine splits the research state into history, the current factor portfolio and reusable capabilities. Under matched resource limits, the authors compare complete runs with fixed or evolving capabilities, and they swap capabilities while holding history and portfolio fixed. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from capability evolution. Across 48 continuation branches from shared states, accumulated capabilities do not consistently outperform the initial set. Parameter tuning of existing factor structures can still improve the portfolio.
In an exploratory replay of two screening batches, some screened-out candidates had positive marginal value, yet submitting all of them in sequence slightly lowered final portfolio IC. The abstract does not state the market or the sample period.
Sources
- arXiv arxiv.org