The following is part of a series of posts about 2026 summer intern projects – for more, see “What the interns have wrought, special jumbo 2026 edition”

At Jane Street we spend a lot of effort trying to predict the future, and we’d like to know if LLMs are any good at it. To do this, we might assemble a collection of queries with known answers—say, predicting which team will win the Super Bowl in a given year—and then measure how well an LLM does.

When trying to construct such a benchmark, a problem immediately arises: LLMs are pretrained on a colossal amount of text that contains answers to many of the questions in our dataset. A model may perform well on a historical benchmark by consulting its memorized knowledge of past outcomes, but fail miserably when presented with new scenarios in production.

This summer, Monte Bohde, a summer intern in our ML Research group, had two tasks: first, measure when LLMs are succeeding merely by regurgitating memorized facts from pretraining, and second, find ways to prevent this behavior.

Finding a dataset and model

Obviously our interest here is in predictions that impact financial markets, but the same ideas can be explored (and discussed on our blog) with lower-stakes data sets. A lot of the experiments that Monte ran were on a dataset of cricket match previews—little writeups setting the stage for an upcoming match. He wanted to see if he could find an open model that clearly demonstrated memorization. That is, when you ask the model to predict the outcome of the match given the match preview text, does it do a lot better for matches that took place before its pretraining knowledge cutoff than those that happened after?

Finding such a model was surprisingly difficult. There were quirks in the data, for instance a distributional shift over time that Monte had to tease out: before the cutoff, the previews were mostly about men’s international matches, whereas after the cutoff they became much more varied. Some models performed little better than chance, which made them poor candidates for the analysis. Others couldn’t even consistently format their predictions.

To assess a model’s level of memorization, he asked three questions:

H1: Is the model any good at predicting cricket matches, i.e., is it additive to the baseline post-cutoff?
H2: Is it better pre-cutoff than post-cutoff?
Sharpness: Does it seem more overconfident pre-cutoff? (A qualitative proxy for memorization)

Bigger models in general performed better and were also more memorized, which made them suitable for the study. The question then became how to mitigate the effect of memorization. As is typical of ML research projects, Monte explored a number of ideas, only some of which panned out.

Divergence decoding reduces memorization without affecting out-of-sample performance

Divergence decoding attempts to remove later years of knowledge from frontier models by doing inference-time steering in logits space. Specifically, given two smaller auxiliary models trained only on data pre-2015 and pre-2026 (or some knowledge cutoff of your large model), you steer your frontier model as:

That way you end up with a frontier model that behaves as though it has knowledge of events only before 2015, without having to train such a full-sized model yourself. To a first approximation you’ve wiped its memory for the years spanning your two auxiliary models.

On the cricket dataset, divergence decoding reduced pre-cutoff performance without harming post-cutoff performance, evidence that it mitigates the effect of memorization.

Distillation

It is also in theory possible to distill the steered model back into the large model and obtain weights for a frontier model that lacks knowledge of recent years. Monte tried to do this distillation, but the resulting model consistently underperformed both the inference-time divergence-decoding model and the original.

Monte’s hypothesis for what went wrong was the data mix during distillation. He used a subset of CommonCrawl tokens from 2019-2024, but in order for the distillation to actually work, your model needs to see sufficiently many task-related tokens. Cricket is somewhat obscure and probably not very well represented in this set.

Prompt engineering and synthetic rewrites

Is there a way to re-cast your data set in a form such that the model is less likely to reach for memorized knowledge? Monte tried a few tacks. One was to rewrite the cricket match previews “in basketball terms.” Maybe then the model would reason about what was in the preview text itself, and what that implied about win probabilities, instead of relying on specific cricket match outcomes.

That didn’t work, but what did work, surprisingly, was to rewrite the previews to be in “podcast form” – that is, to appear to be transcripts from a podcast discussing the upcoming match. In this rewrite Monte also removed player quotes, which seemed to be low signal. With this slightly different presentation of the same data, the model performed better both pre- and post-cutoff but had a larger improvement post-, implying reduced memorization.

Natural language autoencoder

Another method is to try to directly catch the model when it is regurgitating facts. To try this, Monte trained a Natural Language Autoencoder (NLA). An NLA contains an encoder model, which maps a model’s activations to a string of text, and a decoder model, which maps that text back to activations. The encoder and decoder are jointly trained via reinforcement learning to accurately reconstruct activations from the produced text.

Monte found that while the NLA decodings didn’t directly explain what the model was thinking about, the sentiment of the NLA decoding did contain a lot of information. For example, the relative count of how many times each team was mentioned in the cricket task was highly correlated (corr ~0.93) to the models’ final prediction:

Unfortunately this effect was not different pre/post cutoff, so it was not a useful signal to identify/remove memorization.

Models were also prone to hallucination in the NLA decoded text. For example, on the cricket task, the model frequently mentioned the team that it thought would win and then filled in the other team with a random placeholder, usually India.

Diff decoding

At one point, Monte tried fine-tuning the base model on the narrow task of predicting cricket match outcomes from the match previews, in the hopes of inducing it to memorize more often, while maintaining its performance post-cutoff. That had mostly lukewarm results, but an interesting finding was that you can use an autoencoder to probe the difference between the two models. In particular you can decode the “diff embedding,” which is just the FT model embedding minus the base model embedding. On the cricket dataset this decodes almost exactly into a description of the FT objective, i.e. the diff decode explicitly mentions making probability based predictions about cricket:

In other words, fine-tuning introduces a “task vector” for predicting probabilities along with a piece that’s correlated to the linear head for what the probability is.

Sequence effects

Monte created a synthetic reasoning dataset to investigate NLA decodings in a high-signal setting. The rows were of the form:

He found that if he applied the NLA unembedding matrix to intermediate embeddings, he saw that the model made up its mind around the middle layer and then in the last few layers started to think about formatting tokens. This suggested that he could remove those last layers. (This turned out to be correct, and removing the unnecessary layers sped up training by 1.5x.)

In general, he found that the NLA readouts and linear head predictions varied across the sequence dimension—and in particular the model didn’t seem to make up its mind until reading the very last token (!).

Conclusions

Like many ML research projects, Monte’s work here was exploratory. It didn’t result in a de-memorizer, but did help identify the most promising avenues for future research. In particular the success of divergence decoding was suggestive, as was some of the structure that he found using NLA.

Looking forward to next summer…

If you’re interested in doing work like this, consider applying! You can find more details here: Jane Street Internships. Applications for our 2027 Summer ML Research internship are now open!