Guide

Test your LLM for look-ahead bias before you trust a backtest

A model trained on text through 2024 has already read how 2019 turned out. Four checks show whether your result survives that.

AI Fin ResearchWorking3 min read

If you score 2015 news headlines with a model trained on text through 2024, the model may know what happened to those companies next. Any return predictability you find could be memory. This is look-ahead bias, and any referee will ask about it.

The checks below go from cheapest to most expensive. Run them in order and report all of them.

1. Split your sample at the knowledge cutoff

Find the model’s stated training cutoff and split your sample there. Report results separately for the period before and after.

import pandas as pd

CUTOFF = pd.Timestamp("2024-06-30")  # the provider's stated cutoff for your model

df["post_cutoff"] = df["date"] > CUTOFF
by_period = df.groupby("post_cutoff").apply(evaluate)  # evaluate is your own metric function
print(by_period)

If the effect holds only before the cutoff, you have a memory result. If it holds after, you have evidence. Lopez-Lira and Tang take this route: they evaluate on headlines dated after the model’s knowledge cutoff.

Stated cutoffs are approximate, so leave a buffer of a few months. The post-cutoff sample is short by construction, so report how much power you have.

2. Remove the identifiers

Strip company names, tickers and product names from the text and score it again.

import re

def anonymize(text: str, names: list[str]) -> str:
    # Longest names first, so "Apple Inc" is replaced before "Apple".
    for name in sorted(names, key=len, reverse=True):
        text = re.sub(rf"\b{re.escape(name)}\b", "the company", text, flags=re.IGNORECASE)
    return text

Glasserman and Lin do this with news headlines and find something you might not expect: inside the training window, the anonymized headlines perform better. They read that as a distraction effect, where the model’s general knowledge of a company interferes with reading the sentiment of the text, and they find it is stronger for larger companies. So compare both versions, and split the comparison by firm size.

Names leak through executives, products, places and unusual numbers. Sample 100 anonymized texts and ask the model to guess the company. If it gets many right, the mask is not working.

3. Ask the model what it knows

For a sample of firm-months, ask the model directly for the outcome you are predicting. Give it no text, only the firm and the date.

def recall_prompt(company: str, month: str) -> str:
    return (
        f"Was the stock return of {company} in {month} positive or negative? "
        "Answer with one word. If you do not know, answer unknown."
    )

sample = df.sample(500, random_state=0)
sample["answer"] = [
    ask(recall_prompt(company, month))  # ask is your wrapper around the model API
    for company, month in zip(sample["company"], sample["month"])
]

Compare accuracy before and after the cutoff. Above chance before and at chance after means the model has memorized outcomes for your sample. That does not prove your main result is memory, but it removes the benefit of the doubt.

4. Use a model that cannot know

The clean fix is a model trained only on text that existed at the time. He, Lv, Manela and Wu built ChronoBERT and ChronoGPT that way: a series of models, each trained on the text available up to a point in time. In their test predicting next-day stock returns from financial news, the time-restricted models reach Sharpe ratios comparable to a much larger Llama model, and they conclude that look-ahead bias in that application is modest. They also stress that the bias depends on the model and the application, which is the reason to test it in yours.

If a point-in-time model reproduces your result, say so in the abstract.

Report all four checks in the paper

Ludwig, Mullainathan and Rambachan set out the standard. For prediction, there must be no leakage between the model’s training data and your sample. For measurement, where the LLM labels text for a downstream regression, you need a small validation sample, because without one a change of model or prompt can move your estimates and you cannot tell which version is right.

A paper that uses an LLM on historical text should state:

  • the model, its version and its stated training cutoff
  • results before and after the cutoff
  • results with identifiers removed
  • the outcome-recall test from step 3
  • a point-in-time model result, or the reason one was not feasible
  • for measurement tasks, the size of the validation sample and the agreement rate

References