Guide

Agent-first coding: use the frontier tools and pay for them

Claude Code or OpenAI Codex, desktop app or command line. Generic and budget setups hide what the new tools can do.

AI Fin ResearchStarter4 min read
Use the agent built by the lab that trains the model, on its best model, on a plan that does not run out.

You describe the outcome and the agent runs the commands

Editor-first Agent-first
You Type the code Describe the outcome
The tool Completes your line Reads the repository, edits files, runs commands
You check Each line as you write it The diff and the output
Good for Small edits Pipelines whose result you can verify
1. You ask"Pull CRSP monthly returns for 2000 to 2023 and check for duplicate permno-date rows."
2. The agent worksIt writes the query and the script, runs them, reads the error, fixes it and prints the row counts.
3. You reviewYou read the diff and the row counts, then ask for the next step.

Research code suits this way of working. It is scripts and pipelines run from a command line, and its questions have checkable answers: did the pull return the right rows, does the regression reproduce the table.

“Claude Code or VS Code” compares an agent with a place to run one. VS Code’s own documentation now describes agents that “find relevant code, make changes, and run checks without you directing each search, edit, and test run.” Choose the agent first and the editor second.

Use Claude Code or OpenAI Codex

Both come from a lab that trains the model, and both run as a desktop app or on the command line.

Claude Code OpenAI Codex
Made by Anthropic OpenAI
Runs in Terminal, IDE, desktop app, browser ChatGPT desktop app, command line, IDE extension, cloud
Gets first Anthropic’s finance agent templates, as plugins OpenAI’s research tooling, bundled in its academic program
New ability lands there firstAnthropic's ten finance agent templates shipped in May 2026 as plugins for its own products. Third-party tools get features later or never.
Generic means lowest commonA tool that must work with every model cannot lean on what the best one does well. Neutral feels prudent and costs you the hard tasks.
Cheap is a different resultOpenAI's October 2026 math results took roughly three hours of top-tier thinking each. A capped plan on a small model never gets there.
A car is not a faster horse. A researcher who tests AI on a free tier is timing the horse.

The cost of economizing is invisible. You never see the analysis you did not attempt, so you conclude the technology is modest.

If your school blocks both, use OpenCode or a local model

Some schools restrict which vendors may receive code or data, and some data cannot leave your machine at all. In that case use OpenCode, an open source agent that lets you choose the model provider, or a local model. The OpenRouter and Bedrock comparison covers where to buy that access.

Treat this as the exception. Expect less from it, and do not judge the technology by it.

Delete the orchestration framework and the retry wrapper

Two years of tooling grew up around weaker models: orchestration frameworks, output parsers, retry wrappers, routers. Much of it is now optional.

Anthropic’s engineering team wrote in December 2024 that the most successful teams they worked with “weren’t using complex frameworks or specialized libraries.” Their advice: find “the simplest solution possible,” because frameworks “often create extra layers of abstraction that can obscure the underlying prompts and responses, making them harder to debug.” They suggest calling the model API directly.

The author of 12-Factor Agents, a guide from HumanLayer, reports the same thing from the field: “I don’t see a lot of frameworks in production customer-facing agents.” Most of what the author sees is “mostly deterministic code, with LLM steps sprinkled in at just the right points.”

Layer Verdict Why
Orchestration framework Delete A loop in a script is easier to read, debug and put in a replication package
Parse-and-retry wrapper around JSON Delete Model APIs now offer structured outputs that, in Anthropic’s words, “guarantee schema-compliant responses through constrained decoding”
The output schema Keep You still have to say which fields you want. Anthropic’s own Python example writes it with Pydantic
Checks where outside data enters Keep Row counts, key uniqueness, date bounds. An agent removing layers should never remove these

“You don’t need Pydantic” is half right. The retry machinery built around it can go, and a typed schema is still the clearest way to state the output.

If you cannot say what a layer does that twenty lines of your own code would not, delete it and see what breaks.

References