Analysis

Keep paying for call transcripts, but only for the history, the identifiers and the rights

An agent now does the cleaning, splitting and scoring that used to justify a data budget. What it cannot produce is history, identifiers and the right to use the text.

AI Fin ResearchCovers Global

A transcript license is still worth paying for in 2026, for different reasons than before. An agent now does the parsing and the scoring. It cannot produce twenty years of calls, the identifiers that link them to returns, or the right to use the text.

24,000+entities under coverage in S&P's Machine Readable Transcripts dataset, by its own listing
2004how far back that listing says the history goes
0years of that history an agent can recreate for you

An agent replaces two of the seven things a transcript product sells

S&P’s listing describes transcripts of earnings, M&A, guidance, shareholder, conference and special calls, with speaker identifiers that link to its estimates data, and entity recognition that tags which companies are mentioned and where. The announcement of the same data on WRDS cites 9,400+ companies, coverage from 2000 in North America and 2004 globally, and tagging by company, speaker and key development.

What the product gives you Does an agent replace it? Why
Splitting a call into presentation and Q&A, speakers and turns Yes An afternoon of work, and you can check it
Sentiment, uncertainty, topic and tone measures Yes And you should build them yourself, on more than one model, because the numbers change with the model
New calls, as they happen Partly Many calls are webcast, and speech recognition is good. Coverage and accuracy are then your problem
Twenty years of history No The calls happened once. Nobody can re-record 2009
Speaker identifiers linked to analysts and executives No Slow, manual and full of name collisions
Company identifiers that merge with returns and accounting data No This is what makes a text measure usable in a regression
The right to use the text in research No A license, not a file

Vendors were admired for the first two rows. The last four are the product.

The AI clause in your license can forbid all of it

An agent may not read licensed text unless the license allows it. The University of Waterloo says its agreements with publishers “do not allow sharing licensed materials with third parties,” including generative AI services. American University says its contracts “explicitly forbid uploading, processing, or otherwise using the content in AI systems,” and that this applies to all AI tools, including the ones the university itself licenses.

A researcher can hold a transcript license and still be barred from running the text measures that recent papers run. Check before you build.

Buy for long panels, build for recent event windows

Your project Do this
Panel regressions over many years, merged with returns Buy. You need the history and the identifiers
A recent event window for a few hundred firms Building from public webcasts is feasible. Document the error rate
A new text measure on existing licensed transcripts Buy the data, build the measure yourself, and get the AI clause in writing
A dashboard or summary of what the calls said Do not buy anything for this. An agent does it

In any dataset, pay for history, identifiers and rights

Pay for what is slow to build and legally scarce: history, identifiers, linking tables, rights. Stop paying for what an agent does in an afternoon: cleaning, reshaping, summarizing, charting. When a renewal quote arrives, ask the vendor which of the two you are being charged for. If the answer is the interface, you have your negotiating position.

Sources

Related