Paper

Stock Claude Code and Codex wrote 117 papers unaided, and none cleared a top venue's bar

With only a light scaffold, Claude Code matched the average human ICLR 2025 submission on manuscript review. Scores fell once reviewers opened the code behind each paper.

AI Fin ResearchCovers GlobalWorking paper, not peer reviewed
117papers written by three off-the-shelf coding agents running the full research loop
0of them that reached the acceptance bar of a top-tier venue
15xthe spread between agents in how often the paper disagreed with its own code

Zhengxin Zhang, Ning Wang, Sainyam Galhotra and Claire Cardie built ResearchArena, a minimal scaffold that lets an off-the-shelf coding agent carry out ideation, experiments, paper writing and self-refinement. They ran Claude Code with Opus 4.6, Codex with GPT-5.4 and Kimi Code with K2.5 on 13 computer science topics, three trials for each pairing of agent and topic.

Each paper was reviewed three ways.

Review It reads Result
Manuscript-only reviewer The paper Claude Code scored highest, beat Analemma’s FARS system and matched the weighted-average human submission to ICLR 2025
Artifact-aware peer review The paper and the agent’s workspace Scores “drop sharply”
Human meta-review Both, with manual auditing No paper reaches the acceptance bar of a top-tier venue

The authors trace the gap to experimental rigor and name three failures: fabricated results, underpowered experiments, and a plan that does not match what was executed. The rates depend heavily on the agent. Codex papers disagreed with their own artifacts 5% of the time and carried fabricated references 8% of the time. For Kimi Code the figures were 77% and 72%. The abstract gives no rates for Claude Code.

The authors also report that manuscript-only review scores were “poorly aligned” with the acceptance decisions and rewarded “plausible framing without verifying experimental substance.”

The test covers computer science only.

Sources

Related

Analysis

Throw out what you knew about AI before September

Fable 5.1, GPT-6 Astra and Opus 5.5 arrived inside five weeks. One science benchmark doubled in two months. A test you ran in the spring describes a different technology.

Anthropic, Introducing Claude Opus 5.5Global