If the platform forbids agents, build the dataset yourself from the source
Filings, financial statements and earnings calls were public before any vendor packaged them. An agent can go back to the source, and the tools to do it are free.
A license can stop you from putting a vendor’s files into an AI system. It cannot stop you from collecting the same public facts yourself. That used to be too much work for one researcher. With an agent it takes about a week.
Everything here has a public source except security returns
| Dataset | Public source | What the agent does | What you do not get |
|---|---|---|---|
| Financial statements | SEC XBRL APIs: every reported concept for a company in one call | Builds a firm-quarter panel and documents each field | History from before XBRL reporting, and a vendor’s standardized items |
| Filing text | SEC submissions history and the filings themselves | Downloads, splits into sections and cleans 10-K, 10-Q and 8-K text | Firms that do not file with the SEC |
| Call transcripts | Company webcasts and replays | Transcribes with open speech recognition, splits speakers, tags the Q&A | The past. Replays expire, and S&P’s history goes back to 2004 |
| Macro and banking series | FRED, the BIS, the ECB, the World Bank | Pulls and documents each series | Little. These are already free |
| Factors and anomalies | Open Source Asset Pricing, the French library, Global Factor Data | Downloads and merges | The security-level returns underneath |
| Security prices and returns | Nothing open matches CRSP | This is the one you still license |
The SEC describes its data service plainly: the APIs “do not require any authentication or API keys to access,” they cover filing history and the XBRL data from financial statements, and the files “are updated throughout the day, in real time.” A bulk download of everything is republished every night.
Recent call transcripts take five steps
S&P lists more than 24,000 entities in its transcript product, with history back to 2004. Nobody can rebuild that history. Recent calls for the firms in your sample are within reach.
- Collect the webcast or replay from each company’s investor relations page.
- Transcribe it with an open model. Whisper’s code and weights are under the MIT License and it handles many languages, which matters for firms outside the US.
- Structure the text: speakers, prepared remarks, questions and answers.
- Check a sample by ear and report the error rate in the paper.
- Keep the audio references and a manifest, so anyone can rebuild the set.
The result is yours. No vendor clause governs what an agent may do with it, and you can build measures on it with any model you like.
Open source tools already wrap these sources
These projects wrap public data sources so an agent can query them. Star counts are from GitHub on 9 October 2026.
| Project | What it reaches | License | Stars |
|---|---|---|---|
stefanoamorelli/sec-edgar-mcp |
SEC EDGAR filings | AGPL-3.0 | 368 |
daniel3303/Equibles |
Self-hosted SEC filings and XBRL financials | AGPL-3.0 | 230 |
stefanoamorelli/fred-mcp-server |
FRED economic data | AGPL-3.0 | 124 |
lzinga/us-gov-open-data-mcp |
More than 40 US government data APIs | MIT | 112 |
hanlulong/openecon-data |
FRED, World Bank, IMF and Eurostat indicators | See repository | 84 |
FTShare-Lab/FTShare-MCP |
Chinese market data and factors | MIT | 250 |
Read the warning on other people’s MCP servers before you install any of them. Each one wraps an API you can also call directly, and for a dataset that will sit under a paper, direct is better: fewer moving parts and nothing between you and the source.
The SEC allows 10 requests a second and blocks unidentified bots
The SEC limits each user to 10 requests a second and says it does not allow “unclassified” bots to crawl its site, so identify your requests and stay under the limit. Company webcasts carry their own terms. A vendor’s transcript stays under the vendor’s license even when the call itself was public.
You become the vendor
Coverage gaps, transcription errors, delisted firms and identifier mistakes are now your problem, and a referee will ask about each one. Linking takes longer than downloading: the SEC’s company key has to be matched to the identifiers in your returns data, and nobody gives that match away.
Build what is public, license what is scarce, and state in the paper which is which. A department that can build most of a dataset has a stronger hand when it negotiates for the rest.
Sources
- SEC sec.gov
- SEC, developer resources and access limits sec.gov
- Whisper, open source speech recognition github.com
- S&P Global Marketplace, Machine Readable Transcripts marketplace.spglobal.com
- Open Source Asset Pricing openassetpricing.com