Stock Claude Code and Codex wrote 117 papers unaided, and none cleared a top venue's bar
With only a light scaffold, Claude Code matched the average human ICLR 2025 submission on manuscript review. Scores fell once reviewers opened the code behind each paper.
Zhengxin Zhang, Ning Wang, Sainyam Galhotra and Claire Cardie built ResearchArena, a minimal scaffold that lets an off-the-shelf coding agent carry out ideation, experiments, paper writing and self-refinement. They ran Claude Code with Opus 4.6, Codex with GPT-5.4 and Kimi Code with K2.5 on 13 computer science topics, three trials for each pairing of agent and topic.
Each paper was reviewed three ways.
| Review | It reads | Result |
|---|---|---|
| Manuscript-only reviewer | The paper | Claude Code scored highest, beat Analemma’s FARS system and matched the weighted-average human submission to ICLR 2025 |
| Artifact-aware peer review | The paper and the agent’s workspace | Scores “drop sharply” |
| Human meta-review | Both, with manual auditing | No paper reaches the acceptance bar of a top-tier venue |
The authors trace the gap to experimental rigor and name three failures: fabricated results, underpowered experiments, and a plan that does not match what was executed. The rates depend heavily on the agent. Codex papers disagreed with their own artifacts 5% of the time and carried fabricated references 8% of the time. For Kimi Code the figures were 77% and 72%. The abstract gives no rates for Claude Code.
The authors also report that manuscript-only review scores were “poorly aligned” with the acceptance decisions and rewarded “plausible framing without verifying experimental substance.”
The test covers computer science only.
Sources
- arXiv arxiv.org
- Claude Code documentation, model configuration code.claude.com