Automate SEC filings research with a coding agent, from EDGAR to a finished panel
One typed goal makes Claude Code pull SEC filings, build a firm-quarter panel, run the checks and commit the result. Auto mode and /goal keep it working until the checks pass.
Seven steps for a first-time user
Claude Code is a program you install and talk to in a terminal window. You type what you want in plain English. It writes the code, runs it, reads the errors and fixes them. Nothing below needs programming.
-
Install it. Open a terminal (Terminal on a Mac, WSL on Windows) and paste this line. You need a Claude subscription or an Anthropic Console account.
curl -fsSL https://claude.ai/install.sh | bash -
Make a folder for the project.
mkdir -p sec-panel/logs && cd sec-panel git init printf 'data/\nlogs/\n' > .gitignore -
Start Claude Code. Type
claudeand press Enter. You are now typing to the agent, and everything from here goes into its prompt. -
Choose Opus 5.5 and raise the effort. Type each line and press Enter.
opusselects the latest Opus model, which is Opus 5.5./model opus /effort high -
Check that auto mode is on. Look at the bottom of the window for “auto mode on”. If it says something else, press Shift+Tab until it appears.
-
Have it set up the files. Paste this request. It tells the agent to read the plain-text copy of this page and create the three files exactly as they appear below.
Read https://aifinresearch.com/guides/sec-filings-with-an-agent.md and create sec.py, CLAUDE.md and tests/test_panel.py exactly as given there. Put my name and email in the User-Agent header in sec.py: <your name> <your email>. Create config/ciks.csv with a cik column holding 320193, 789019 and 19617 (Apple, Microsoft and JPMorgan Chase), and config/concepts.csv with a concept column holding Assets, Liabilities and NetIncomeLoss. Commit, then stop and show me the files. -
Type the goal and walk away. Copy the
/goalblock from “Type the goal” further down. The agent works until the tests pass.
$ claude start the agent > /model opus Opus 5.5 > /effort high more reasoning on every step > Read https://aifinresearch.com/guides/... and create sec.py ... > /goal data/panel.parquet has one row per cik, concept ...
Terms used on this page
| Word | Plain meaning | How you set it |
|---|---|---|
| Claude Code | The agent: a program that does the task in your project folder | Type claude in a terminal |
| Model | The AI doing the work. Use Opus 5.5 | /model opus |
| Effort | How much the model reasons before each step | /effort high |
| Auto mode | The agent runs commands without asking each time, and a second AI screens every action | Shift+Tab until the status line says “auto mode on” |
/goal |
A finish line. The agent keeps working until a check you named passes | /goal followed by the condition |
CLAUDE.md |
A text file of standing rules that the agent reads at the start of every session | Create it in the project folder |
| Skill | A saved procedure the agent runs the same way every time | A folder under .claude/skills/ |
| MCP tool | A connector that gives the agent an extra tool, such as a database. This pipeline needs none, because the SEC data is reached with plain Python | Type /mcp to see what is connected |
The sections below explain each step and give the files in full.
Auto mode removes the per-tool prompt and /goal removes the per-turn prompt
A plain Claude Code session stops in two places. It asks before each command, and it hands control back to you after each answer. Two features remove both stops, and the documentation says how they fit: “auto mode removes per-tool prompts, and /goal removes per-turn prompts.”
| Auto mode | /goal |
|
|---|---|---|
| Removes | The permission prompt before each command | The wait for you after each turn |
| Who checks the work | “A separate classifier model reviews actions before they run” | “After each turn, a model checks whether the condition holds” |
| Stops when | The turn ends | The condition is met, or the checking model judges it impossible |
| Turn it on | claude --permission-mode auto, or Shift+Tab inside a session |
/goal followed by the condition |
From Claude Code 2.1.283, auto mode is the starting mode for interactive terminal sessions on every plan.
Auto mode went from launch to default in five months
--dangerously-skip-permissions."A researcher who still clicks “yes” on every command is doing by hand what the tool’s default now does for everyone else. The classifier wrongly blocks 0.4% of safe actions in Anthropic’s tests, so it rarely gets in the way.
Shift+Tab switches it on, and plain words set its limits
| You want to | How |
|---|---|
| Switch into or out of auto mode | Press Shift+Tab. The status bar reads “auto mode on” |
| Switch while answering a permission prompt | Choose “Yes, and switch to auto mode” |
| Set a limit for this session | Say it: “don’t push” or “wait until I review before deploying.” The classifier blocks matching actions until you lift the limit |
| Make a limit permanent | Add a deny rule. A spoken limit “can be lost” when a long conversation is compacted |
| See what was blocked | Run /permissions and open the Recently denied tab. Press r to retry an action with your approval |
| Allow one blocked action | Name the action and the detail that makes it risky. “You can force-push” clears nothing. Naming the branch does |
| Know when it hands control back | After 3 blocks in a row or 20 in a session, auto mode pauses and asks you again |
Repeated blocks usually mean the classifier does not know your setup. The documentation’s fix is to tell it which repositories, buckets and domains you trust.
The four SEC endpoints you need
The SEC says its data APIs “do not require any authentication or API keys to access” and are “updated throughout the day, in real time.”
| What you want | Endpoint | What comes back |
|---|---|---|
| Ticker to company key (CIK) | https://www.sec.gov/files/company_tickers.json |
Every ticker with its CIK and name |
| A company’s filing history | https://data.sec.gov/submissions/CIK##########.json |
Form types, dates, accession numbers, document names |
| Everything a company reported in XBRL | https://data.sec.gov/api/xbrl/companyfacts/CIK##########.json |
Every concept, every period, in one call |
| One concept across all companies | https://data.sec.gov/api/xbrl/frames/us-gaap/Assets/USD/CY2023Q4I.json |
One value per filer for that period |
The CIK is ten digits with leading zeros. Stay under 10 requests a second in total and identify yourself: the SEC says it does not allow “unclassified” bots and reserves the right to block addresses that send excessive requests.
Write three files before the agent starts
You fix the request rules, the instructions and the checks. The agent writes the rest.
sec.py makes every request. It caches each response, waits between calls and sends your name and email in the User-Agent header.
# sec.py
import json
import time
from pathlib import Path
import requests
HEADERS = {"User-Agent": "Your Name your.email@university.edu"}
CACHE = Path("data/sec")
def get(url: str) -> dict:
"""Fetch once, cache on disk, stay under the SEC's rate limit."""
path = CACHE / (url.split("//")[1].replace("/", "_"))
if path.exists():
return json.loads(path.read_text())
response = requests.get(url, headers=HEADERS, timeout=30)
response.raise_for_status()
time.sleep(0.15)
CACHE.mkdir(parents=True, exist_ok=True)
path.write_text(response.text)
return response.json()
def company_facts(cik: int) -> dict:
return get(f"https://data.sec.gov/api/xbrl/companyfacts/CIK{cik:010d}.json")
def filings(cik: int) -> list[dict]:
recent = get(f"https://data.sec.gov/submissions/CIK{cik:010d}.json")["filings"]["recent"]
keys = ["form", "filingDate", "accessionNumber", "primaryDocument"]
return [dict(zip(keys, row)) for row in zip(*(recent[k] for k in keys))]
def document_url(cik: int, filing: dict) -> str:
accession = filing["accessionNumber"].replace("-", "")
return f"https://www.sec.gov/Archives/edgar/data/{cik}/{accession}/{filing['primaryDocument']}"
CLAUDE.md holds the rules. Claude Code reads it at the start of every session (the guide to instruction files and skills explains the file).
## SEC data
- Use sec.py for every request. Never call sec.gov directly.
- Under 10 requests a second. Cached files in data/sec/ are never fetched twice.
- The company key is the CIK. Keep it as an integer and pad to ten digits only in URLs.
- Record form type, filing date and accession number for every value you keep.
- Use the filing date, not the period end, for anything that will be matched to returns.
## Git
- Commit after every step that passes its checks. One step, one commit.
- The message says what changed and why, with the row count when data changed.
- Before starting work, read `git log --oneline -20` to see what earlier runs did.
- Never commit data/ or credentials. Never force push, reset or amend.
The last SEC line prevents look-ahead bias. A value is known to the market on the day it is filed, and an agent will not assume that unless you say so.
tests/test_panel.py is the judge. The agent may not edit it, so it cannot pass by lowering the bar.
# tests/test_panel.py
import pandas as pd
panel = pd.read_parquet("data/panel.parquet")
missing = pd.read_csv("data/missing.csv")
ciks = set(pd.read_csv("config/ciks.csv")["cik"])
concepts = set(pd.read_csv("config/concepts.csv")["concept"])
def test_one_row_per_key():
key = ["cik", "concept", "period_start", "period_end"]
assert not panel.duplicated(key).any()
def test_every_firm_is_accounted_for():
assert ciks == set(panel["cik"]) | set(missing["cik"])
def test_only_requested_concepts():
assert set(panel["concept"]) <= concepts
def test_only_annual_and_quarterly_reports():
assert panel["form"].isin(["10-K", "10-Q"]).all()
def test_filed_after_period_end():
assert (panel["filing_date"] >= panel["period_end"]).all()
period_start is empty for balance sheet items, which are measured on one day. Income and cash flow items need it, because a third-quarter filing reports a three-month and a nine-month figure that end on the same date.
Start the session on Opus 5.5 at high effort
Once you know the tool, one line does what steps 3 to 5 did by hand.
claude --model opus --effort high --permission-mode auto
The .gitignore line from step 2 keeps downloaded data and run logs out of the repository, so git status can come back clean.
| Flag | Effect | From the documentation |
|---|---|---|
--model opus |
Selects the latest Opus model, which is Opus 5.5 | The opus alias “uses the latest Opus model for complex reasoning tasks” |
--effort high |
Raises reasoning above the medium default of Opus 5.5 |
high is for “work where verification matters or edge cases are likely” |
--permission-mode auto |
Runs commands without asking, with each one reviewed by a classifier | “Long tasks, reducing prompt fatigue” |
Nobody watches an overnight run, so the model has to catch its own mistakes. That is the reason to raise the effort from Opus 5.5’s default of medium to high.
Type the goal
/goal data/panel.parquet has one row per cik, concept, period_start and period_end
for every firm in config/ciks.csv and every concept in config/concepts.csv, from
10-K and 10-Q filings only, with filing_date, form and accession on each row. Firms
with no XBRL data are listed in data/missing.csv with a reason. Prove it by running
pytest and showing that it exits 0, then show that git status is clean. Use sec.py
for every request. Do not edit config/ or tests/. Stop after 40 turns.
Setting a goal starts work at once. After every turn a second model reads the conversation and returns one of three verdicts: not yet met, met or impossible. That model “doesn’t run commands or read files independently,” so the condition has to name a check whose output the agent prints.
| The documentation asks for | In this goal |
|---|---|
| One measurable end state | One row per firm, concept and period in data/panel.parquet |
| A stated check | pytest exits 0 and git status is clean |
| Constraints that matter | sec.py for every request, config/ and tests/ untouched |
| A bound on the run | “Stop after 40 turns” |
$ claude --model opus --effort high --permission-mode auto > /goal data/panel.parquet has one row per cik, concept, period_start and period_end for every firm in config/ciks.csv ... Prove it by running pytest and showing that it exits 0. Stop after 40 turns. ● I'll read the rules and the tests, then write the pull. ● Read(CLAUDE.md) ● Read(tests/test_panel.py) ● Bash(git log --oneline -20) ⎿ no commits yet ● Write(build_panel.py)
While it runs, /goal with no argument shows the condition, the time elapsed, the turns evaluated, the tokens spent and the checking model’s latest reason. /goal clear stops it.
Let the agent commit its own work
Most researchers commit by hand at the end of the day, if at all. Give that job to the agent. Claude Code’s best practices list git commit among the tools that are safe to allow, tell you to “ask Claude to commit with a descriptive message,” and point Claude at git history when a question is about how code came to be.
git reset --hard and amending a commit the session did not create. The agent can add to the history and cannot rewrite it.● Bash(python build_panel.py) ⎿ wrote data/panel.parquet and data/missing.csv ● Bash(pytest -q) ⎿ 1 failed, 4 passed: test_one_row_per_key ◎ Goal not yet met: pytest did not exit 0. ● Third-quarter filings repeat the nine-month figure with the same end date. I'll drop exact duplicates and keep period_start in the key. ● Edit(build_panel.py) ● Bash(pytest -q) ⎿ 5 passed ● Bash(git add -A && git commit -m "Key the panel on period_start and period_end") ⎿ 3 files changed ● Bash(git status --short) ⎿ nothing to commit ◎ Goal met.
An agent that commits after each passing step also leaves a dated record of how the dataset was built. A referee who asks when a filter was added gets a commit hash.
Leave the session running overnight
Stay in the interactive session. Every command, every verdict from the checking model and every block from the classifier is on screen, and you can step in at any point. If auto mode blocks three actions in a row it pauses and asks you, so the question is waiting when you come back.
Start the session inside tmux on a machine that stays on, such as an office desktop or a server you reach over SSH. The session then survives a closed terminal or a dropped connection.
tmux new -s sec-panel
claude --model opus --effort high --permission-mode auto
Type the goal, press Ctrl+B and then D to detach, and go home. In the morning:
tmux attach -t sec-panel
Type /goal to read the turns taken, the tokens spent and the last verdict, then read the commit log.
To refresh the panel every day, leave the session open and repeat the skill:
/loop 1d /sec-panel
The documentation says /loop runs a prompt or a skill “on repeat while the session stays open,” with intervals in seconds, minutes, hours or days. A loop on a fixed interval ends after seven days, so start it again each week.
The skill it runs:
---
name: sec-panel
description: Build or refresh the firm-quarter panel from SEC XBRL data. Use when asked to update fundamentals, add firms or add concepts to the panel.
---
# SEC panel
1. Read `git log --oneline -20` and the last manifest.
2. Read the firm list from config/ciks.csv and the concept list from config/concepts.csv.
3. For each firm, call company_facts through sec.py.
4. Keep 10-K and 10-Q values only. Where a period was reported more than once, keep the first filing and store the later ones in restatements.parquet.
5. Write data/panel.parquet and data/missing.csv.
6. Run pytest. Report concepts missing for more than 20% of firms.
7. Write manifests/<date>.json with row counts and the date of the run, then commit. Do not change the firm or concept lists.
Three analyses worth automating
Each one takes its own goal and its own test file. For anything where a language model reads the text, run it on more than one model, keep a hand-labeled sample, and test for look-ahead bias.
Bonus: link it to CRSP and Compustat with the wrds library
Downloading SEC data is quick. Matching the SEC’s company key to the identifiers in your returns and accounting data is the slow step. If your school subscribes to WRDS, the wrds Python library gets you there.
Start by looking at what you have. These calls are in the library’s own README.
import wrds
db = wrds.Connection()
db.list_libraries() # what your subscription includes
db.list_tables(library="comp") # tables in one library
db.describe_table(library="comp", table="company")
Compustat’s company table carries the CIK next to its own key, and the CRSP link table connects that key to CRSP’s. Check both table names against your own subscription before you rely on them.
link = db.raw_sql("""
select c.cik, c.gvkey, l.lpermno as permno, l.linkdt, l.linkenddt
from comp.company as c
join crsp.ccmxpf_linktable as l on l.gvkey = c.gvkey
where c.cik is not null
and l.linktype in ('LU', 'LC')
and l.linkprim in ('P', 'C')
""", date_cols=["linkdt", "linkenddt"])
link["cik"] = link["cik"].astype(int)
Merge on cik, then keep rows where the filing date falls inside the link dates. Now every number and every sentence from a filing sits beside the return that followed it.
Keep the licensed side out of the model’s context. The agent writes the query and you run it, as in the guide to agents and licensed data.
References
- EDGAR Application Programming Interfaces, SEC sec.gov
- Developer resources and access limits, SEC sec.gov
- Keep Claude working toward a goal, Claude Code documentation code.claude.com
- Choose a permission mode, Claude Code documentation code.claude.com
- How we built Claude Code auto mode: a safer way to skip permissions, Anthropic (2026) anthropic.com
- What's new, week 32: auto mode becomes the default, Claude Code documentation (2026) code.claude.com
- Claude Code overview, Claude Code documentation code.claude.com
- Model configuration, Claude Code documentation code.claude.com
- Run prompts on a schedule, Claude Code documentation code.claude.com
- Best practices for Claude Code, Claude Code documentation code.claude.com
- Checkpointing, Claude Code documentation code.claude.com
- wrds, the Python client for WRDS pypi.org