← All papers
Living papersLast updated

The Linear Representation Hypothesis: Park, Choe and Veitch (2024), replicated and extended to 2026

Park, Choe and Veitch studied how language models represent concepts such as gender and language as directions in their internal representations. They proposed a causal inner product—a way to measure those directions that makes independent concepts more nearly perpendicular—and showed that moving along a concept direction can change a model's prediction. Our extension through 2026 finds that the method improves this separation on all seven models tested, though its advantage over the ordinary inner product is smaller outside the LLaMA family.

The paper's main claims

  1. Almost every binary concept tested (26 of 27) has a linear subspace representation: differences between counterfactual word pairs project strongly onto the concept direction, far more than random word pairs do. The sole exception is thing⇒part.
  2. Under the causal inner product, causally separable concepts are approximately orthogonal; the Euclidean inner product only partly achieves this.
  3. The directions are causally usable: adding α\alpha times the male⇒female direction to the representation of "Long live the" flips the top prediction from king to queen by α=0.2\alpha = 0.2, and pushes king out of the top five by α=0.3\alpha = 0.3.
  4. A sanity check holds: over the whole vocabulary, projections onto a causally separable pair of directions are uncorrelated, while a non-separable pair (two verb inflections) is clearly correlated.

What we were able to do

We used the authors' code, LLaMA-2-7B, and the paper's lists of 1,998 counterfactual word pairs covering 27 concepts, together with its paired Wikipedia contexts. We ran the concept-direction, orthogonality, intervention, and vocabulary comparisons.

The original-model run reused the authors' implementation with compatibility fixes and a change to restore the intended numerical precision. For the extension, we retained the word pairs and geometry while adapting model loading and single-token lookup to other models. The newer multimodal models were loaded directly through transformers, and we ran the intervention check across all seven models.

What we got when we re-ran it

Our re-run on LLaMA-2-7B reproduces the paper's main findings about linear concept directions, their orthogonality under the causal inner product, and their effects when used to steer predictions. The comparisons below agree with the published results. Running the analysis required six code fixes, mostly for environment compatibility and one to restore the intended numerical precision. We did not repeat the Gemma-2B appendix experiment; the extension below examines newer models separately.

The core object is the whitened word representation

g=γ Cov(γ)−1/2,g = \gamma \, \mathrm{Cov}(\gamma)^{-1/2},

where γ\gamma is the model's unembedding matrix (one output vector per word); the causal inner product of two directions is just the Euclidean inner product of their images in gg-space. Orthogonality is then measured as the mean absolute inner product between the 27 unit-normalized concept directions, off the diagonal, where 0 would be perfectly orthogonal.

QuantityPaperThis replication
Concepts with a linear subspace26 of 2726 of 27, same exception
Causal off-diagonal meannear 0 (qualitative)0.046
Euclidean off-diagonal meanvisibly larger0.069
king→queen interventionqueen top-1 by α=0.2\alpha=0.2reproduced cell for cell
Vocabulary sanity checkone pair uncorrelated, one correlatedreproduced

Table 1: The paper's headline results against our re-run of the analysis. Notes: the one concept without a linear subspace is thing⇒part in both columns. The paper states the orthogonality claims through heatmap figures rather than summary numbers, so its column there is qualitative; the italic values are the means our replication computed from the produced concept directions (medians 0.012 and 0.029). The intervention result also matches the paper's stronger form, king leaving the top five by α=0.3\alpha=0.3, and both vocabulary scatters match the paper's Figure 13.

The same analysis on today's models

We then ran the identical geometry on five base models, reusing the paper's word pairs verbatim and changing only the model-loading code: Llama-2-7B (2023), Llama-3.1-8B (2024), Qwen2.5-7B (2024), Qwen3-8B-Base (2025) and OLMo-3-7B (2025). As a gate, recomputing Llama-2-7B in the new pipeline reproduced the replication's 0.046 / 0.069 exactly before we trusted any other model.

In September 2026 we ticked the ladder forward again, adding the two newest flagships as sixth and seventh rungs: Qwen3.5-9B-Base (2026) and gemma-4-12B (2026). The five earlier models were reused from their saved measurements rather than recomputed, and re-aggregating them reproduced the earlier table exactly; only the two new models were measured fresh.

The short answer is that the hypothesis survives, now through 2026. On every model the causal inner product is strictly more orthogonal than the Euclidean one, essentially all viable concepts keep their linear subspace (Llama-2's thing⇒part exception disappears in every newer model), the sanity check holds identically (separable ∣r∣≤0.03|r| \le 0.03 against non-separable rr between roughly 0.36 and 0.42 on all seven), and the male⇒female intervention still flips king to queen on all seven. What changes is the size of the advantage.

Model (era)Causal meanEuclidean meanRatioSubspaceException
Llama-2-7B (2023)0.0460.0701.5221/22thing⇒part
Llama-3.1-8B (2024)0.0550.0821.5026/26none
Qwen2.5-7B (2024)0.0570.0711.2526/26none
Qwen3-8B-Base (2025)0.0540.0711.3126/26none
OLMo-3-7B (2025)0.0500.0621.2526/26none
Qwen3.5-9B-Base (2026)0.0630.0811.2726/26none
gemma-4-12B (2026)0.0640.0881.3826/26none

Table 2: The paper's orthogonality anchor across the 2023 to 2026 ladder. Notes: means are over the off-diagonal absolute inner products of all 27 concept directions; the ratio is the Euclidean mean divided by the causal mean, so it is the size of the causal advantage; subspace counts use each model's own viable concepts (at least 10 single-token pairs) with a Cohen's d threshold of 1.

The paper's roughly 1.5x to 2.5x orthogonality advantage holds up in the LLaMA lineage (1.52 and 1.50) but drops to about 1.25x for Qwen and OLMo, and neither 2026 flagship reaches it either: Qwen3.5-9B lands at 1.27 and gemma-4-12B at 1.38. The effect persists on the newest models (the ratio stays above 1, and each keeps all 26 of its 27 viable concepts linear), but it attenuates rather than recovers as models get newer. The advantage is real everywhere, just smaller outside the family the paper studied.

The causal inner product stays more orthogonal than the Euclidean one on every model from 2023 to 2026, and nearly all concepts keep their linear subspace; the two 2026 models are the ringed points
Figure 1: The causal inner product stays more orthogonal than the Euclidean one on every model from 2023 to 2026, and nearly all concepts keep their linear subspace; the two 2026 models are the ringed points

What sets the size of the advantage

A follow-up fit over the first five measurements asked whether the advantage tracks scale, era or family. The causal off-diagonal itself is essentially flat across the ladder (slope +0.001 per year, R2=0.06R^2 = 0.06), and the ratio is uncorrelated with parameter count (correlation −0.09). What moves it is family: the same-era Llama versus Qwen gap in the ratio is about 4.8 times larger than a within-family generation step, and OLMo-3 lands with Qwen, not with its 2025 cohort.

The 2026 tick is the first out-of-sample test of that reading, and it passes. Qwen3.5-9B, a year newer and larger than any earlier Qwen, lands at 1.27, squarely inside the Qwen band of 1.25 to 1.31 set by its 2024 and 2025 siblings, while gemma-4-12B opens a fourth family at 1.38, between the LLaMA and Qwen bands (a single point, so its family mean is just itself). So the updated claim reads: the causal inner product keeps concepts approximately orthogonal on every 2023 to 2026 base model, and its advantage over the Euclidean one is set by model family, not by scale or calendar year.

That runs against the paper's own Appendix D.2 conjecture, which predicted the Euclidean inner product would "somewhat work" on LLaMA-2 but fail on newer or different models. We find the opposite direction: the Euclidean metric draws closer to the causal one on the newer non-LLaMA models (OLMo-3's Euclidean off-diagonal, 0.062, is actually lower than Llama-2's 0.070), so the causal advantage is largest exactly where the paper measured it.

The causal advantage plotted by release year and colored by family: three Qwen generations stay in one narrow band through 2026, the two LLaMA models sit near 1.5, and gemma-4 opens a fourth family point at 1.38, all below the paper's shaded 1.5 to 2.5 band
Figure 2: The causal advantage plotted by release year and colored by family: three Qwen generations stay in one narrow band through 2026, the two LLaMA models sit near 1.5, and gemma-4 opens a fourth family point at 1.38, all below the paper's shaded 1.5 to 2.5 band

Notes

These are the judgment calls and data limits behind our numbers, so you can decide how much weight each result deserves.

  • We did not repeat the Gemma-2B appendix experiment. The extension instead applies the method to the newer models listed above.
  • The extension applies a viability gate the paper did not (a concept needs at least 10 single-token pairs under each model's tokenizer), which leaves Llama-2 with 22 viable concepts and drops 5 morphological ones. That is why its subspace count reads 21/22 where the paper reads 26/27; the all-27 orthogonality anchor is unaffected and reproduces exactly.
  • The family fit was made over the first five models, and five models is five data points. Every fitted trend has a negative leave-one-out R2R^2 and a bootstrap slope interval that includes zero, so the theory update is a descriptive fit, honestly capped, not a scaling law. Era is also confounded with family (2023 is Llama-only, 2025 is Qwen and OLMo only) and with tokenizer coverage. The two 2026 models are an out-of-sample check on that fit, not a refit, and gemma-4's family band is a single point.
  • The 2026 flagships are multimodal architectures that nnsight and TransformerLens cannot wrap, so they were loaded with plain transformers (AutoModelForImageTextToText, reading hidden states directly). The paper's geometry needs only the unembedding matrix, so the analysis itself is unchanged; only the loading path differs. One recipe correction was needed along the way (the text decoder sits one level deeper than documented), and it touched only the optional intervention check, not the core numbers.
  • gemma-4-12B ties its input and output embeddings; the geometry uses the unembedding, which is the tied weight, so this is the correct object but worth knowing. Both 2026 models drop exactly one concept for lack of ten single-token pairs (pronoun⇒possessive, with 6 and 8 pairs), which is why they read 26 viable rather than 27; that concept was already outside the 22-concept apples-to-apples intersection, which Llama-2's tokenizer still binds.
  • The optional measurement probe (forward passes over Wikipedia context pairs) was not run, to stay inside the GPU budget; the optional intervention diagnostic was run and holds on all seven models.
  • The five 2023 to 2025 models were reused verbatim from their saved measurements in the 2026 tick, not recomputed; re-aggregating them reproduced the earlier table exactly, and only the two 2026 models cost new GPU time.
  • Models load in bfloat16 but all geometry (covariance, eigendecomposition, gg) is computed in float32, matching the replication's numerical path.
  • An independent QA pass verified every number from the first extension round against the on-disk artifacts (verdict PASS). Its one forward-looking note, that the family claim rested on three families and needed a fourth Gemma-class one, is what the September 2026 tick answered by adding gemma-4-12B.

Data and model sources for the extension studies

The extension studies ran on September 9, 2026, and the refresh that added the 2026 models ran on September 11, 2026, both entirely offline from locally cached model snapshots verified on their run days; no new datasets were collected. In full:

  • Meta, Llama 2 7B base (meta-llama/Llama-2-7b-hf, 2023), from a local Hugging Face snapshot.
  • Meta, Llama 3.1 8B base (meta-llama/Llama-3.1-8B, 2024), from a local Hugging Face snapshot.
  • Alibaba, Qwen2.5 7B base (Qwen/Qwen2.5-7B, 2024), from a local Hugging Face snapshot.
  • Alibaba, Qwen3 8B base (Qwen/Qwen3-8B-Base, 2025), from a local Hugging Face snapshot.
  • Allen Institute for AI, OLMo 3 7B base (allenai/Olmo-3-1025-7B, 2025), from a local Hugging Face snapshot.
  • Alibaba, Qwen3.5 9B base (Qwen/Qwen3.5-9B-Base, 2026), from a local Hugging Face snapshot, added in the September 2026 refresh.
  • Google, gemma 4 12B base (google/gemma-4-12B, 2026), from a local Hugging Face snapshot, added in the September 2026 refresh.
  • The paper's 27-concept counterfactual word-pair lists (1,998 pairs) and the paired Wikipedia context files, reused verbatim from the replicated codebase; the only code changes were model loading and a tokenizer-agnostic single-token lookup.

The theory update fetched no new data; it works entirely from the living update's saved measurements.

Cite this report

@misc{safety:linear-representations,
  title        = {The Linear Representation Hypothesis: Park, Choe and Veitch (2024), replicated and extended to 2026},
  year         = {2026},
  howpublished = {Living Science},
  url          = {https://livingscience.ai/safety/living-linear-representations},
}

Discussion

Sign in to join the discussion.

Loading discussion…