← All papers
Living papersLast updated

The Geometry of Truth: Marks and Tegmark (2023), replicated and extended to 2026

Marks and Tegmark found a direction in Llama-2's internal representations that separates true from false statements across topics. A classifier trained on number comparisons could also distinguish true from false translation statements, and changing activations along the direction could change the model's answer. Our extension through 2026 finds that these results hold in the newest Qwen model. In gemma-4, truth remains distinguishable within individual datasets, but the direction transfers poorly across topics at the layer used for the main comparison.

The paper's main claims

  1. True and false statements separate visibly along the top principal components of the activations, dataset after dataset.
  2. Probes trained on one dataset generalize to very different ones: trained on larger-than and smaller-than comparisons, they score 0.95 to 1.00 on Spanish English translation statements, whichever probe technique is used.
  3. The mass-mean probe, whose direction is just the average true activation minus the average false one, μ+−μ−\mu_+ - \mu_-, is the most causally effective: adding or subtracting it flips the model's answer, with normalized indirect effects (a 0 to 1 score of how far the intervention moves the answer) of 0.70 and 0.54 in the headline cell.
  4. On the largest model, Llama-2-70B, a statement's log probability correlates with its truth: r of 0.85 on cities and 0.95 on translations, with the sign flipping on negated versions.
  5. Bigger models give probes that generalize better, and probes trained on plain statements stumble on negations, because negation partially flips the truth direction.

What we were able to do

We used the paper's true/false statement datasets and its code for extracting activations, fitting probes, and testing generalization. The original-model comparisons used Llama-2-7B and Llama-2-13B, with separate log-probability checks on Llama-2-70B.

We updated the authors' implementation to work with the available nnsight environment, then reused the corrected probe and extraction code and saved Llama-2 activations for the extension. The extension kept the paper's statement datasets and evaluation protocol while changing the model and, for the 2026 models, the loading environment.

What we got when we re-ran it

Our re-run on Llama-2-7B and 13B reproduces the main findings: probes trained on comparison statements generalize to translation statements, and intervening along the mass-mean direction changes the model's answers. On the headline 13B test, probe accuracy is 0.972–0.977, within the paper's reported 0.95–1.00 range; the intervention effects are also close to the published values. We updated code that depended on an older version of nnsight. The 70B experiments were completed separately once model access became available, with results included below.

QuantityPublishedThis replication
Comparisons to translations, logistic regression probe (13B)0.95 to 1.000.9718
Comparisons to translations, mass-mean probe (13B)0.95 to 1.000.9746
Comparisons to translations, CCS probe (13B)0.95 to 1.000.9774
Intervention effect, false to true (13B)0.700.724
Intervention effect, true to false (13B)0.540.516
70B log-prob correlation, cities (addendum)0.850.85
70B log-prob correlation, translations (addendum)0.950.9495

Table 1: Published values against our re-run. Notes: probe rows are accuracy on the sp_en_trans test set for probes trained on larger_than plus smaller_than at layer 14 of Llama-2-13B; intervention rows are the normalized indirect effects for the headline mass-mean cell. The two 70B rows are a later standalone addendum, computed on 2026-09-10 once Llama-2-70B access was restored. In the addendum the negated datasets flip sign as the paper predicts, r of −0.63 on negated cities and −0.89 on negated translations.

The same analysis on today's models

The truth direction survives, with a single 2026 exception. Holding the paper's code and probes fixed, we extracted activations from the 2024 and 2025 open models across three families and re-ran the headline generalization test, training on number comparisons and testing on translations. Every 2024 and 2025 base model lands between 0.93 and 0.99, and Llama-3.1-8B beats the paper-era 13B anchor at a smaller size.

The 2026 refresh, run a year later on the newest flagships, splits cleanly in two. Qwen3.5-9B-Base, the newest Qwen of March 2026, reproduces the paper's headline finding better than the 2023 anchor itself: at the paper's own canonical layer, four tenths of the way through the network, its mass-mean probe scores a perfect 1.00 against the Llama-2-13B anchor of 0.9746, its mean cross-dataset generalization of 0.883 is the highest in the whole 2023 to 2026 ladder, and the same direction is causally load-bearing, flipping 0.994 of the model's true answers to false once the nudge is doubled past the family's saturated readout (at natural magnitude it flips almost nothing, like every Qwen before it, and the false-to-true side never flips at any magnitude). gemma-4-12B is the first genuine counterexample the living update has found. Truth is still linearly separable within each dataset, with same-dataset accuracy of 0.95 to 0.98, comparable to the anchor, so the raw feature exists; but it does not organize into one dataset-transferable direction at the canonical layer, where the headline cross-dataset test scores 0.393, near chance. The direction only starts to form much deeper, with mean generalization climbing to about 0.81 to 0.84 around three quarters depth, and even at that best layer the headline cell reaches just 0.751 for mass-mean and roughly 0.56 for the other probes. The same-dataset accuracy rules out a data or extraction bug, so this reads as an architecture-dependent limit of the new unified multimodal design rather than a failure of the analysis.

Model (release)Mass-meanLogistic regressionCCSMean cross-dataset (MM)
Llama-2-7B (2023)0.92660.74580.79660.519
Llama-2-13B (2023, paper anchor)0.97460.97180.97740.769
Llama-3.1-8B (2024)0.98590.97180.96050.808
Qwen2.5-7B (2024)0.96050.92370.95200.793
Qwen3-8B-Base (2025)0.97740.98870.98870.812
OLMo-3-7B (2025)0.93500.94630.96330.734
Qwen2.5-7B-Instruct0.98020.98310.98590.850
Qwen3.5-9B-Base (2026)0.99441.00001.00000.883
gemma-4-12B (2026)0.75140.55650.58190.836

Table 2: The paper's headline generalization test on the modern ladder. Notes: first three columns are accuracy training on larger_than plus smaller_than and testing on sp_en_trans, at each model's best validation layer; the last column is the mean accuracy over all train-on-one, test-on-another pairs on a common 8-dataset grid. The Llama-2 rows reuse the replication's own activations and our pipeline reproduces them exactly. The two 2026 rows are the 2026 refresh; gemma-4-12B's best validation layer sits at three quarters of its depth, and at the paper's canonical 0.4-depth layer its mass-mean cell falls to 0.3927 and its mean cross-dataset accuracy to 0.4822, while Qwen3.5's mass-mean at that canonical layer is 1.0000.

The headline generalization accuracy by model release date, extended through the 2026 models, with the replicated Llama-2-13B anchor as a dashed line
Figure 1: The headline generalization accuracy by model release date, extended through the 2026 models, with the replicated Llama-2-13B anchor as a dashed line

The direction also transfers to chat models: on the one labeled pair, Qwen2.5-7B base against its instruction-tuned version, tuning slightly strengthens the probe (0.9605 to 0.9802) and the base and instruct directions stay closely aligned (cosine similarity 0.84 to 0.90 at the best layer). The causal test carries over too, with a caveat: a single mass-mean direction flips 0.983 of Llama-3.1-8B's answers, while Qwen and OLMo need the same direction scaled about four times larger before answers flip (0.493 on Qwen2.5-7B), because their true or false readout is already near saturation. The direction is aligned in those families as well; the natural-magnitude nudge is just too small to move a saturated readout.

The updated theory, fitted from these runs, changes three things, and the 2026 tick bends the first of them further. First, the paper's "truth lives about 0.4 of the way through the network" is a Llama-specific constant: all three Llama models read truth out at 0.35 to 0.41 of their depth, but Qwen and OLMo read it out deeper, at 0.5 to 0.63. The 2026 models pull this in both directions: Qwen3.5 reads truth out perfectly right at the canonical 0.4 depth, while in gemma-4 the canonical layer holds no transferable direction at all and what forms near three quarters depth is only partial, so in that family the depth story does not merely shift, it breaks. Second, generalization saturates with scale: it climbs steeply within Llama-2 from 7B (0.519) to 13B (0.769), but across today's 7 to 8B models it is already flat at 0.73 to 0.85, with a fitted asymptote near 0.9, so most of the gain by 8 to 13B is already banked and era matters more than raw size; Qwen3.5's 0.883, the ladder record, still sits under that asymptote. Third, the negation gap, how much accuracy a probe trained on plain statements loses on their negations, collapses across generations, from 0.52 on Llama-2-7B to 0.05 on Qwen3-8B-Base: newer models' single truth direction already carries most of the polarity-invariant component, and training on both polarities recovers negated truth at about 0.99 in every model.

The best probe layer as a fraction of model depth, by family, against the paper's Llama-2 band
Figure 2: The best probe layer as a fraction of model depth, by family, against the paper's Llama-2 band
The negation gap by model release generation, shrinking from 0.52 to 0.05
Figure 3: The negation gap by model release generation, shrinking from 0.52 to 0.05

Notes

These are the judgment calls and data limits behind our numbers, so you can decide how much weight each result deserves.

  • The replication's 902 GB of Llama-2 activations were reused for the extension rather than re-extracted; the extension pipeline reproduces the on-disk anchor values bit-exact (13B mass-mean 0.9746, 7B 0.9266). The one non-exact anchor is CCS, which is stochastic by construction; we report the disk value.
  • The 70B results were computed separately on September 10, 2026, after access to the model became available.
  • The causal flip-rate on the new models is a cheap analog using one mass-mean direction, not the paper's full 12-cell intervention table, which was kept on the Llama-2 anchors. The Qwen magnitude sweep behind the 0.493 figure survives only in the run transcript, not in a saved results file, a traceability gap our QA flagged.
  • OLMo-3-7B has a known family-specific massive-activation outlier that depresses its raw mass-mean numbers; its logistic regression oracle stays near 0.96, so linear separability is intact.
  • The best-layer depths come from a coarse sweep of about six layers per model, so each depth is resolved only to about plus or minus 0.12; the Llama-shallow versus Qwen and OLMo-deep split survives that resolution, finer claims would not.
  • The cross-family scale fit is under-powered (correlation 0.40 over a narrow 6.7 to 13.2B range confounded with generation); the clean scale evidence is the within-family Llama-2 7B to 13B contrast.
  • New-model extraction was capped at 1,000 statements per dataset over 8 datasets; the paper's extended datasets were not re-extracted for new models. QA's verdict on both extension studies was accept with minor issues, and it suggested adding a fourth model family such as Gemma at the next update.
  • The 2026 flagships are unified multimodal architectures that nnsight cannot wrap, so their activations were read with plain transformers through output_hidden_states, verified to match a forward hook on the decoder layer exactly (a maximum difference of 0.0), and the causal analog for Qwen3.5 used forward hooks. They also needed a fresh environment, transformers 5.17.0 rather than the 4.57.1 the earlier models ran under, and the decoder sat at a different attribute path than documented, found by introspection. Every 2023 to 2025 number in the 2026 refresh is reused from the earlier results on disk, never recomputed.
  • gemma-4-12B was the optional second 2026 model and is included as a documented partial rather than a silent skip. Its causal flip-rate was not run, since its weak cross-dataset generalization would make the result hard to interpret, and its strong same-dataset accuracy is what rules out a data or extraction bug behind the exception.
  • The Qwen3.5 flip-rate is the same cheap single-direction analog as before, not the paper's full intervention table: the 0.994 figure is true-to-false at doubled magnitude, natural magnitude flips almost nothing (0.017) because the readout is saturated, and the false-to-true side never flips at any magnitude.

Data and model sources for the extension studies

The extension studies ran on 2026-09-09 and 2026-09-10, with the 2026 refresh on 2026-09-11, all from locally cached model snapshots verified against each snapshot's config before loading. In full:

  • Meta, Llama-2-7B and Llama-2-13B (meta-llama/Llama-2-7b-hf, meta-llama/Llama-2-13b-hf), 2023, Hugging Face; the replication anchors, activations reused from the replication run.
  • Meta, Llama-2-70B (meta-llama/Llama-2-70b-hf), 2023, Hugging Face; used only for the log-probability addendum, on two A100-80GB GPUs.
  • Meta, Llama-3.1-8B (meta-llama/Llama-3.1-8B), 2024, Hugging Face.
  • Alibaba, Qwen2.5-7B and Qwen2.5-7B-Instruct (Qwen/Qwen2.5-7B, Qwen/Qwen2.5-7B-Instruct), 2024, Hugging Face.
  • Alibaba, Qwen3-8B-Base (Qwen/Qwen3-8B-Base), 2025, Hugging Face.
  • Allen Institute for AI, OLMo-3-7B (allenai/Olmo-3-1025-7B), 2025, Hugging Face.
  • Alibaba, Qwen3.5-9B-Base (Qwen/Qwen3.5-9B-Base), March 2026, Hugging Face; snapshot accessed September 2026 for the 2026 refresh.
  • Google, gemma-4-12B (google/gemma-4-12B), 2026, Hugging Face; snapshot accessed September 2026 for the 2026 refresh.
  • Marks and Tegmark's geometry-of-truth true/false datasets, shipped with the paper's replication code: cities, neg_cities, sp_en_trans, neg_sp_en_trans, larger_than, smaller_than, likely, and cities_cities_conj for the extension; the 70B addendum additionally used companies_true_false and counterfact_true_false.
  • The paper's own codebase as fixed during the replication (probe classes, extraction, generalization protocol), reused verbatim; the 2024 and 2025 models ran under a pinned environment of transformers 4.57.1, nnsight 0.4.11, and torch 2.14.0, and the 2026 models under a fresh transformers 5.17.0 environment with torch 2.14.0.

The theory update fetched no new data; it works entirely from the living update's results and the original analysis data.

Cite this report

@misc{safety:geometry-of-truth,
  title        = {The Geometry of Truth: Marks and Tegmark (2023), replicated and extended to 2026},
  year         = {2026},
  howpublished = {Living Science},
  url          = {https://livingscience.ai/safety/living-geometry-of-truth},
}

Discussion

Sign in to join the discussion.

Loading discussion…