The IOI Circuit: Wang, Variengien, Conmy, Shlegeris and Steinhardt (2023), replicated and extended to 2026
Wang and coauthors studied how GPT-2 small completes sentences such as “When Mary and John went to the store, John gave a drink to Mary.” They identified a circuit of 26 attention heads that tracks the repeated name and helps select the other person as the recipient. Our extension finds that this circuit explains less of the behaviour in larger GPT-2 models: heads that copy the correct name remain important, while other parts of the computation spread across more heads.
The paper's main claims
- GPT-2 small does IOI: mean logit difference 3.56, and the right name beats the wrong one 99.3% of the time.
- A circuit of 26 attention heads in seven classes (name movers, negative name movers, S-inhibition, induction, duplicate token, previous token, backup name movers) implements the task.
- The circuit is faithful: keeping only those 26 heads recovers about 87% of the full model's logit difference. Faithfulness is the share of the model's task performance the circuit reproduces on its own.
- The circuit passes completeness and minimality checks, and it explains enough to design adversarial prompts: adding an extra mention of the correct name confuses the mechanism and the logit difference collapses.
What we were able to do
We used the authors' Easy-Transformer code and their generated prompt distributions: sentences testing indirect-object identification and corresponding three-name prompts used for activation patching. The original-model run used GPT-2 small, including the paper's 100,000-prompt baseline evaluation.
We adapted the authors' interactive notebook code for execution, extracted analyses from notebook cells, enabled the completeness searches, and fixed an off-by-one token-position error in the adversarial examples. Circuit discovery continued to use the authors' implementation. For the extension, we adapted the model selection, filtered the name pool for each tokenizer, and evaluated behaviour and circuits across GPT-2 and Pythia models. We also ran behavioural evaluations on the newest models.
What we got when we re-ran it
Our re-run recovers the same 26 attention heads and closely reproduces the circuit's performance on the indirect-object identification task. The model prefers the correct name on 99.5% of prompts, compared with 99.3% in the paper, and the circuit retains 87.8% of the full model's logit difference, compared with about 87%. Adversarial prompts weaken performance in both runs, more sharply in ours. Some supporting quantities differ, particularly for the deliberately naive comparison circuit, as described below; we did not repeat three appendix analyses.
| Quantity | Paper | This replication |
|---|---|---|
| Circuit found by patching | 26 heads, 7 classes | same 26 heads recovered |
| Mean logit difference | 3.56 | 3.49 |
| Right name preferred | 99.3% | 99.5% |
| Probability on the right name | 49% | 50.2% |
| Circuit faithfulness | ~87% | 87.8% |
| Name mover attention to the right name | 0.59 | 0.584 |
| S-inhibition attention to the repeated name | 0.51 | 0.437 |
| Adversarial prompt logit difference | 1.23 | 0.59 |
Table 1: The paper's headline numbers beside our re-run of the analysis. Notes: baseline rows use the paper's full 100,000-prompt scale. The faithfulness ratio matches cleanly; the absolute values behind it sit 8 to 13 percent low because the shipped completeness script runs at a smaller sample size. The adversarial row reproduces the qualitative claim, with the collapse even stronger in our run. The deliberately naive comparison circuit's faithfulness gap came out 0.574 against the paper's 0.1, though the point it supports, that the naive circuit is incomplete, did reproduce.
The code is a 2022 interactive research notebook rather than a push-button pipeline. Our replication applied 11 fixes: 7 minor (version pinning, headless plotting, batching) and 4 major, three of which were extracting analyses that exist only as notebook cells or re-enabling a completeness script whose shipped defaults skip its own searches, and one of which was a genuine off-by-one bug in the token positions of the adversarial-example code. None of the fixes touch the method: the circuit itself reproduces with the repo's own primitives.
The same analysis on today's models
We took the paper's discovery recipe, path patching, which swaps in activations from a spoiled prompt to measure how much each head contributes to the answer, made the model a parameter, and ran the full pipeline up the GPT-2 ladder (small to XL) and across the Pythia family (70M to 2.8B), with behaviour-only checks on two modern small models, and returned in September 2026 to add the newest flagships we could reach to the top of the ladder. Before trusting any new number we confirmed the harness reproduces the stored anchor exactly: the paper's hand-picked circuit scores 87.78% through this pipeline, matching the replication's 87.8%, and a freshly discovered circuit on GPT-2 small scores 88.6%.
The answer splits in two: the behaviour survives, the circuit does not scale.
| Model | Params | Logit difference | Right name wins | Circuit faithfulness |
|---|---|---|---|---|
| GPT-2 small (anchor) | 124M | 3.49 | 99.5% | 87.8% curated / 88.6% discovered |
| GPT-2 medium | 355M | 3.58 | 100% | 41.5% |
| GPT-2 large | 774M | 4.48 | 99.9% | 11.5% |
| GPT-2 XL | 1.5B | 3.79 | 99.9% | 28.4% (partial, output side only) |
| Pythia 70M | 70M | −0.57 | 37.8% | n.a. (fails the task) |
| Pythia 410M | 410M | −0.38 | 39.6% | n.a. (fails the task) |
| Pythia 1.4B | 1.4B | +0.85 | 71.6% | 87.7% |
| Pythia 2.8B | 2.8B | −0.34 | 40.6% | n.a. (fails the task) |
| Qwen2.5-0.5B | 494M | +5.06 | 100% | behaviour only |
| SmolLM2-1.7B | 1.7B | +4.17 | 99.2% | behaviour only |
| Qwen3.5-9B (base) | 9B | +5.46 | 99.0% | behaviour only |
| Qwen3.5-9B (instruct) | 9B | +5.84 | 99.5% | behaviour only |
| gemma-4-12B (base) | 12B | +5.71 | 100% | behaviour only |
Table 2: Baseline IOI behaviour and discovered-circuit faithfulness per model. Notes: logit differences at N=2000 prompts; faithfulness is of a freshly discovered circuit with the paper's class sizes, at N=100 for the smaller models and N=30 or 50 for the largest. The 2026 rows come from the September 2026 tick and are scored on the behaviour tier's 200 prompts. GPT-2 XL's circuit covers only three of the six discoverable classes (budget), so its 28.4% is not comparable to the full rows. Where the model fails the task the faithfulness ratio is arithmetically defined but meaningless, so we report n.a. rather than a number.

The behaviour survives scale and generation, with one striking exception. Every GPT-2 size solves IOI, and modern small models solve it even better: Qwen2.5-0.5B reaches a logit difference of +5.06, meaning its score for the right name beats the wrong one by a wider margin than any GPT-2, at a third of GPT-2 medium's size. But the entire Pythia family fails: three of its four sizes score the wrong name higher on the majority of prompts, and even Pythia 1.4B manages only a weak +0.85. Since Pythia's sizes bracket GPT-2's and its architecture is newer, the difference points at training data and recipe, not scale. The paper never claimed Pythia, so this is a boundary on the behaviour, not an error in the paper.
The September 2026 tick extends that picture. Qwen3.5-9B reaches a logit difference of +5.46 and prefers the right name on 99.0% of prompts, its instruct variant edges higher at +5.84, and gemma-4-12B reaches +5.71 with the right name winning on every single prompt. That is roughly 1.6 times the GPT-2 small anchor of +3.49, so the behaviour is not just surviving into the 2026 models, it keeps strengthening. These are behavioural measurements only: both models are multimodal architectures that neither TransformerLens nor nnsight can load, so circuit discovery and faithfulness were descoped for them. That is a tooling limit rather than a finding, and whether they solve IOI by the paper's mechanism or a different one stays open.

The circuit is another story. The discovered circuit's faithfulness falls monotonically up the GPT-2 ladder, 88.6% to 41.5% to 11.5% from small to medium to large, even though every one of those models still performs the task essentially perfectly. And where the behaviour does exist across a generation gap, the method still works: Pythia 1.4B, the one Pythia that weakly does IOI, yields a discovered circuit with 87.7% faithfulness, essentially the GPT-2 small value, on a rotary-position architecture the paper never touched.
What the circuit becomes at scale
The theory update re-analyzed the per-head measurements to ask which of the paper's head classes still show strong single heads, ones whose removal moves the logit difference by at least 0.2, in each model.
| Head class | GPT-2 small | medium | large | XL | Pythia 70M | 410M | 1.4B | 2.8B |
|---|---|---|---|---|---|---|---|---|
| name mover | 9 | 11 | 8 | 5 | 2 | 1 | 1 | 0 |
| negative name mover | 2 | 2 | 2 | 1 | 1 | 1 | 0 | 0 |
| S-inhibition | 3 | 1 | 0 | 0 | 0 | 0 | 0 | n.r. |
| induction | 2 | 0 | 0 | n.r. | 0 | 0 | 1 | n.r. |
| duplicate token | 1 | 0 | 0 | n.r. | 0 | 0 | 0 | n.r. |
| previous token | 1 | 0 | 0 | n.r. | 0 | 0 | 0 | n.r. |
Table 3: Strong single heads per class and model. Notes: n.r. means the discovery pass for that class was not run on that model for budget reasons, which is different from a true zero.

The updated theory reads: the name mover backbone is universal, the rest of the circuit is a GPT-2 small phenomenon. Name movers, the heads that copy the right name into the answer, are the only class with strong single heads at every GPT-2 size and in the working Pythia, and their negative twins are the second most persistent. S-inhibition and the upstream induction, duplicate-token and previous-token chain show strong single heads essentially only in GPT-2 small; in larger models those jobs spread across many individually weak heads. Fitting faithfulness against model size within GPT-2 gives a slope of −0.61 per tenfold increase in parameters (R² 0.75), but on four points with p = 0.13 that is a trend, not a law. Whether the circuit itself grows with model size could not be tested at all, because the discovery template fixes the head count at 26 by construction. Behaviourally, generation matters and size barely does: at matched scale, GPT-2 XL scores 3.79 where Pythia 1.4B scores 0.85.
Notes
These are the judgment calls and data limits behind our numbers, so you can decide how much weight each result deserves.
- Both extension studies are partial by design. The paper's completeness and minimality searches were deliberately not carried up the ladder (the search alone took 42 minutes on GPT-2 small), so those properties are anchor-only. GPT-2 XL's full six-class circuit was never finished (one discovery pass alone took 38 minutes), so its 28.4% is a three-class, output-side circuit and is flagged as such everywhere; Pythia 2.8B ran one discovery pass only. The faithfulness-collapse headline stands on the three complete points: small, medium, large.
- An independent adversarial QA review accepted the studies with minor issues. A launcher bug caused runaway filesystem scans that wasted wall-clock time but, per QA, corrupted no completed result. QA also verified the Pythia failure is real and not a harness bug: the models load correctly, the same pipeline scores GPT-2 highly, the failing models prefer the wrong name at below-chance rates, and Pythia 1.4B works as an internal control.
- Every extension faithfulness number is for a mechanically discovered circuit, not the paper's hand-curated one. The two agree on GPT-2 small (88.6% vs 87.8%), which is what licenses the cross-model comparison; part of the collapse at scale may still reflect the fixed-size template rather than the models.
- The scale fits rest on four points and are not statistically significant; the largest models were patched at smaller sample sizes (N = 30 or 50 against 100), so their faithfulness estimates are noisier.
- Faithfulness ratios are withheld wherever the model fails the task, since dividing by a negative baseline produces a number without meaning.
- The modern models (Qwen2.5-0.5B, SmolLM2-1.7B) were tested for behaviour only; no circuit was extracted, so they say nothing about circuit survival. OLMo-2-1B could not be loaded by the pinned library version and was skipped, and SmolLM2 stands in for the Llama and Gemma models originally sketched for this tier.
- The September 2026 tick (Qwen3.5-9B, gemma-4-12B) is behavioural only, by design: both are multimodal architectures that neither the paper's TransformerLens fork nor nnsight can load, so no circuit or faithfulness number exists for them. This is a tooling limit rather than a finding, and it means the tick shows the behaviour surviving, not the circuit; the newest models could in principle solve IOI by a different internal mechanism.
- The 2026 runs needed a fresh environment (transformers 5.17.0), since the pinned 2024 library does not know these architectures, and the models' text decoder sits one attribute deeper than current documentation sketches; neither detail touches the scored logits, which come from the models' ordinary forward pass. All prior ladder numbers were reused verbatim from disk, nothing was recomputed. The ladder point for 2026 is the base model in each case, matching the rest of the ladder, with Qwen3.5-9B's instruct variant (+5.84) reported alongside for colour.
- Whether circuits grow with model size remains open: answering it needs the per-model minimality search that was descoped, a named gap rather than a silent one.
Data and model sources for the extension studies
The extension studies downloaded model weights fresh from Hugging Face between September 7 and 10, 2026. In full:
- OpenAI, GPT-2 model family: gpt2 (124M), gpt2-medium (355M), gpt2-large (774M), gpt2-xl (1.5B), retrieved from Hugging Face.
- EleutherAI, Pythia suite: EleutherAI/pythia-70m, EleutherAI/pythia-410m, EleutherAI/pythia-1.4b, EleutherAI/pythia-2.8b, retrieved from Hugging Face.
- Alibaba Cloud Qwen team, Qwen/Qwen2.5-0.5B (494M), retrieved from Hugging Face (behaviour tier only).
- Hugging Face TB, HuggingFaceTB/SmolLM2-1.7B, retrieved from Hugging Face (behaviour tier only).
- Alibaba Cloud Qwen team, Qwen/Qwen3.5-9B-Base and Qwen/Qwen3.5-9B (9B), retrieved from Hugging Face in September 2026 (behaviour tier only, September 2026 tick).
- Google, google/gemma-4-12B (12B), retrieved from Hugging Face in September 2026 (behaviour tier only, September 2026 tick).
- The IOI prompt distributions are the paper's own: pIOI, sentences built from the paper's mixed templates over its pool of 99 single-token names, and pABC, the matching three-distinct-name sentences used as the spoiled baseline for patching. For tokenizers other than GPT-2's, the pool was filtered to names that remain a single token: 88 of 99 for Pythia, 97 for Qwen2.5, 79 for SmolLM2, and all 97 names in the reused 2026-tick prompt set stay single tokens for both Qwen3.5 and gemma-4, so all 200 of its prompts were scored.
- The authors' Easy-Transformer code, reused verbatim as the analysis codebase for every GPT-2 and Pythia run.
The theory update fetched no new data or models; it works entirely from the living update's stored measurements and the replication's anchors.
Cite this report
@misc{safety:ioi-circuit,
title = {The IOI Circuit: Wang, Variengien, Conmy, Shlegeris and Steinhardt (2023), replicated and extended to 2026},
year = {2026},
howpublished = {Living Science},
url = {https://livingscience.ai/safety/living-ioi-circuit},
}
Discussion
Sign in to join the discussion.
Loading discussion…