Guide

Automated research pipelines run in four fields, and you can point one at finance

Machine learning, mathematics, biomedicine and finance each have a system that runs most of the research loop. Two of them you can install today.

AI Fin ResearchWorking4 min read

“Fully automated” here means a system that takes a question or a field, then produces ideas, code, experiments and a written result without a person between the steps.

Four fields have a system that runs most of the loop

Field System Year How far it goes
Machine learning The AI Scientist and AI Scientist-v2 2024, 2025 Idea to written paper. One v2 manuscript scored above the average human acceptance threshold at a peer-reviewed workshop
Mathematics OpenAI internal model 2026 Results released with Lean proofs a computer can check
Biomedicine Co-Scientist 2025 Hypotheses only. People still run the experiments
Any field Agent Laboratory 2025 Literature review to report, with human feedback at each stage
Finance AI-Powered (Finance) Scholarship 2025 96 signals, three complete papers written for each
Finance RD-Agent(Q) 2025 Hypothesis, code and real-market backtest in a loop

Machine learning. The AI Scientist (2024) generates research ideas, writes code, runs experiments, makes figures, writes a full paper and then runs a simulated review. Its successor, The AI Scientist-v2 (2025), drops the human-written code templates the first version needed. The authors submitted three fully autonomous manuscripts to a peer-reviewed ICLR workshop, and one scored above the average human acceptance threshold.

Mathematics. On 6 October 2026 OpenAI released a set of results produced by an internal model, in a public repository, with many of the proofs formalized in Lean so a computer can check them. OpenAI reports that the average result used compute equal to roughly three hours of ChatGPT Pro thinking.

Biomedicine. Co-Scientist (2025) is a multi-agent system built on Gemini that generates, critiques and ranks hypotheses. Its validation is in three biomedical applications. It stops at the hypothesis: the experiments are still run by people.

Any field, with a person nearby. Agent Laboratory (2025) takes a human research idea through literature review, experimentation and report writing, and lets the user give feedback at each stage.

Finance. Novy-Marx and Velikov (2025) mined over 30,000 candidate return predictors from accounting data, kept the 96 that passed their testing protocol, and had LLMs write three complete papers for each one. The papers came with theoretical justifications and with citations that were, in the authors’ words, “on occasion, imagined.” They present the exercise as a cautionary tale about industrialized HARKing. RD-Agent(Q) (2025), from Microsoft, is an open-source multi-agent framework for quantitative strategies: a research stage proposes hypotheses, a development stage writes the code, and the code is run in real-market backtests whose results feed the next round.

The AI Scientist and RD-Agent install today

The AI Scientist. The README says the code is designed for Linux with NVIDIA GPUs. After cloning the repository:

conda create -n ai_scientist python=3.11
conda activate ai_scientist
sudo apt-get install texlive-full
pip install -r requirements.txt
export OPENAI_API_KEY="YOUR KEY HERE"
python launch_scientist.py --model "gpt-4o-2024-05-13" --experiment nanoGPT_lite --num-ideas 2

The README warns that the codebase executes code written by a language model and tells you to containerize it. Do that before the first run.

RD-Agent. The README says it supports Linux only and that Docker must be installed for most scenarios. Once the rdagent package is installed and configured with a model key:

rdagent health_check
rdagent fin_factor     # iterative factor proposal and implementation
rdagent fin_model      # iterative model proposal and implementation
rdagent fin_quant      # factors and models evolved jointly
rdagent fin_factor_report --report-folder=<folder of financial reports>

These scenarios run on Microsoft’s Qlib. The model names in the AI Scientist README date from 2024, so check each project’s current list before you copy a command.

A general system needs a template, a data pull and a judge

  1. A template the system can extend. The AI Scientist starts from an experiment folder such as nanoGPT_lite: a data loader, a baseline run and a metric. For an asset pricing question that means a script that loads your panel, a baseline specification, and one number that says whether a variant is better.
  2. Data the agent can pull by itself. If the pull needs a web form or a notebook cell run by hand, the loop stops there. See Make your data pull something an agent can run.
  3. A judge the agent cannot touch. An agent that both proposes and evaluates will find results. One recent paper lets the agent propose factors and gives the verdict to a frozen statistical referee, which admitted 5 to 11 times fewer false factors than leaky ones.

Then add the checks finance needs and machine learning mostly does not. Test for look-ahead bias if a language model reads historical text. Hold out a period the system never sees. Log every specification it tried, because the count is the denominator for any multiple-testing correction.

Automation is ahead of verification

In two controlled comparisons, no factor-mining method consistently won, and agents that accumulate their own skills did not reliably improve. Mathematics can attach a proof a computer checks. Empirical finance cannot, so the package that makes an automated result believable is plainer: the code, a manifest of the data, the list of everything tried and a test that was fixed before the search began.

References