Seven LLMs scoring the same earnings calls agree with a rank correlation of 0.52
Across thirteen text measures on S&P 500 transcripts, the choice of model changes the size, sign and significance of downstream coefficients.
Hamid Boustanifar and Sasan Mansouri test whether LLM-based text measures survive a change of model. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on thirteen constructs, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk.
Cross-model rank correlations average 0.52. Transcript-level differences that are common across providers account for 34% of total score variation. Disagreement between models does not predict later disagreement among analysts or in the market, which the authors read as a model-specific component and not shared ambiguity in the disclosure.
The choice of model carries into inference: coefficient magnitudes, signs and statistical significance vary substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels still depend on which models are in the ensemble. The authors conclude that LLM-generated variables are model-contingent measurements that need validation across providers.
The sample is S&P 500 firms and one document type, so the result has not been shown for smaller firms, other filings or other languages.
Sources
- arXiv arxiv.org