Guide

Token management for researchers, and why to buy usage before you hire help

Cached input costs a tenth of the normal price and batch jobs cost half. Spend on usage for the people you have before you spend on support staff.

AI Fin ResearchStarter3 min read
0.1xthe price of a cached input token against a fresh one, in Anthropic's documentation
50%the discount for batch jobs at both Anthropic and Amazon Bedrock
7 hoursa week that scientists report saving with AI, in a 2026 survey of over 600

Most of the bill is what the model reads

A token is a piece of a word. Models charge for the tokens they read (input) and the tokens they write (output). The context window is everything the model has in front of it at once: your instructions, the files it opened, the conversation so far.

An agent rereads that context on every step. So on a long task the bill is mostly input, and the same instructions and files are paid for again and again unless you do something about it.

Caching, batching and a short context cut the bill

The waste The fix What it saves
Sending the same instructions and reference files on every call Prompt caching Anthropic prices cache reads at 0.1 times the base input price
Running a large scoring job in real time Batch processing 50% off at Anthropic, where most batches finish within an hour. Bedrock prices batch 50% below on-demand
Loading every file “just in case” Give the agent a manifest and let it open what it needs Fewer tokens and better answers

The third row also improves the answers. Anthropic’s documentation says: “As token count grows, accuracy and recall degrade, a phenomenon known as context rot.” A crowded context makes the model worse. Start a fresh session for a new task, keep instructions short, and point the agent at files instead of pasting them.

One new result took three hours of the best model

OpenAI reported that the average result in its October 2026 mathematics release used compute equal to roughly three hours of its top-tier thinking. New results appear at that scale: hours of the best model on one problem.

A capped plan stops long before that. The researcher on it sees a tool that summarizes and autocompletes, and concludes that is what the technology is.

Buy usage before you hire help

When a project falls behind, the usual answer is another pair of hands.

Hiring support Buying more usage
Time to first result Weeks to recruit, then training Minutes
What you must do Explain the task until someone else understands it Write the task down precisely once
When the spec is wrong You find out when the work comes back You find out in the same hour and fix it
What is left when the project ends Experience that leaves with the person Instruction files, checks and pipelines that carry to the next paper
Near a deadline Fixed hours As much as you can direct

The communication gap. Most of the cost of delegating is explaining. You have to specify the task either way. With a person, each round of misunderstanding costs days. With an agent it costs minutes, so you can afford to be wrong three times before lunch.

Skills compound. Every task you run through an agent leaves something reusable: a tested data pull, a checks file, a clearer instruction. Your own skill at directing the tool grows too, and it applies to every later project. Hours bought from someone else do not accumulate that way for you.

People are still needed for judgment, accountability and the training of the next generation of researchers, and a doctoral student learns the data by handling it. Give the people you already have the best tools and enough usage first, then see what work is left.

Track cost per finished task

  • Put every researcher who writes code on a plan that does not run out during a working day.
  • Use caching and batch for anything repeated or large.
  • Track cost per finished task. A plan that looks expensive and finishes the analysis is cheaper than one that looks frugal and stalls.

References