Academic research is far behind the AI frontier, from finance to mathematics
A paper about a model is out of date before it clears review. In mathematics, the newest results come from a model no academic can run.
On 29 July 2026 OpenAI opened its academic program with a model it called GPT-5.6 Sol Pro. On 10 September its finance product named GPT-6 Astra. On 6 October it published mathematical results from an internal model that has not been released at all. One lab named three frontier systems in ten weeks.
Journals run on a slower clock. Hadavand, Hamermesh and Wilson document that publishing in economics proceeds much more slowly than in the natural sciences, and more slowly than in other social sciences and finance. They trace much of the lag to authors taking a long time to revise. Finance does better than economics on their measure, and it is still far too slow for a subject that is replaced every few months.
The published record is stale in every field
| Field | What the record shows | Why it is stale |
|---|---|---|
| Finance | Lopez-Lira and Tang’s ChatGPT paper was posted on 15 April 2023. Its sixth version is dated 28 October 2025, and arXiv lists no journal reference | It describes models nobody would pick today, and text measures change with the model |
| Machine learning | The AI Scientist’s README, read on 9 October 2026, still gives launch commands for models dated May and October 2024 | A widely cited tool for automated research documents two-year-old models |
| Research automation | Agent Laboratory, posted January 2025, found that o1-preview produced the best outcomes | That ranking is a fact about 2024 |
| Mathematics | OpenAI’s October 2026 results came from an internal model it has not released | Outsiders can check the Lean proofs. They cannot run the search |
| Science at large | In a survey of over 600 scientists, nearly half use AI every day and report a backlog of untested hypotheses | Daily practice is ahead of the published account of it |
| Access | OpenAI’s free academic program drew more than 13,000 applications for 10,000 places | Most researchers who want current tools do not have them |
Rigor does not rescue a paper about last year’s model
A claim about capability is out of date on arrival. “Model X can do this” and “model X fails at that” describe a system the lab has already replaced, and a literature built from such claims describes the past.
Validity work lasts. Look-ahead bias, training leakage and measures that change with the model are as real with the next model as with the last one. Academic researchers are ahead of industry on all three, and those papers are worth publishing slowly.
A careful paper about a model nobody uses is still out of date. Researchers are behind on what the tools can do, and the publishing process keeps them there.
Write papers that survive a model change
- Put the model, its version and the date you ran it in the abstract.
- Build the pipeline so the full paper reruns on a new model in a day, and rerun it before every revision.
- Report results on at least two current models.
- Post the working paper early. The journal version is the archive.
- If you run a department, pay for access to current models and follow releases at the source instead of waiting for the literature to describe them.
Sources
- NBER nber.org
- OpenAI, Sharing AI progress in mathematics openai.com
- OpenAI, ChatGPT for Academic Researchers openai.com
- OpenAI, ChatGPT for Financial Services openai.com
- arXiv, Can ChatGPT Forecast Stock Price Movements? arxiv.org
- arXiv, Agent Laboratory arxiv.org
- arXiv, AI in Science: Early Insights arxiv.org
- The AI Scientist, code repository github.com