Concept Space Has Real Geometry — and Known Limits
Six research programs show that the “concept space” LLMs project into has real geometric structure — and two of them identify the limits where it breaks down. Those limits are exactly where external memory lives.
In August I wrote that your AI already has a concept space — an internal geometric structure where related concepts cluster together, and where hallucinations are collisions in a lossy projection. That post was speculative. It proposed concept space as a useful framing without anchoring it to published evidence.
Since then I built a 2,400-belief knowledge base from the LLM internals literature using External Epistemic Memory (EEM) — structured claims with justification chains, truth values, and retraction cascades, stored outside the model in a database. The research says something stronger than what I claimed: concept space has measurable geometric structure — directions, polytopes, orthogonal subspaces — documented across six independent research programs spanning four years. Two of those programs also identify where the structure breaks down. Both findings matter for understanding why external memory works.
The Evidence
Huh et al., ICML 2024 — “The Platonic Representation Hypothesis.” Different neural networks trained on different data modalities converge toward the same representation of reality as they scale. Not the same embeddings — the same distance relationships. The kernel K(x_i, x_j) = f(x_i) dot f(x_j) captures how a model measures similarity between datapoints, and these kernels converge across architectures, objectives, and modalities. The name is deliberate: models are prisoners in Plato’s cave, seeing different shadows, but the shadows converge because the reality casting them is shared.
Park et al., ICML 2024 — “The Linear Representation Hypothesis.” Concepts are encoded as directions in LLM activation space. “Queen minus king” is not a party trick — it reflects genuine linear structure. Park formalizes this with a causal inner product that correctly separates independent concepts (gender vs. language) as orthogonal vectors, and finds that 26 of 27 tested concepts in LLaMA-2 7B are linearly encoded in the unembedding space. The one failure (thing-to-part) marks the boundary of where concept space geometry breaks down.
Park et al., ICLR 2025 — “Geometry of Categorical and Hierarchical Concepts.” Beyond binary directions, categorical concepts (colors, animals, professions) are represented as polytopes — convex hulls in representation space. Hierarchical concepts (animal > mammal > dog) occupy orthogonal subspaces. This is not a flat coordinate system. Concept space has geometric structure: polytopes for categories, orthogonality for hierarchies, simplices for natural kinds.
Elhage et al. / Bricken et al., Anthropic 2022-2023 — Superposition and Sparse Autoencoders. Models pack more concepts than they have dimensions through superposition — overlapping features in compressed space. Sparse autoencoders decompose these into interpretable features. Bricken et al. (“Towards Monosemanticity,” 2023) showed that SAEs trained on independently initialized copies of the same architecture recover largely the same features — SAE features across models are more similar to each other than to raw neurons within either model. This suggests the features reflect the data more than the initialization, though as Usama & Chang later showed, shared architecture itself accounts for some measured convergence.
Koepke et al., 2026 — “Back Into the Cave.” A partial falsification of the strong PRH thesis. At coarse granularity (semantic category level), vision and language models show stable cross-modal alignment. But at fine granularity (individual items), alignment drops as gallery density increases. The one-to-one image-caption setup in Huh et al. breaks down in realistic many-to-many settings. Most damaging: the reported trend of stronger language models increasingly aligning with vision does not appear to hold for newer models — the scale trend flattens or fails to replicate. Models share the large-scale topology of concept space but diverge on the details, and the divergence may not shrink with scale.
Usama & Chang, 2026 — “Convergence Without Understanding.” Another challenge to the strong thesis. Sixteen models (1.5B to 72B parameters) tested on 800 reasoning problems. Representational convergence and reasoning convergence dissociate — models can converge in how they represent concepts while diverging in how they reason about them. More troubling: randomly initialized models show higher CKA similarity (0.864) than trained models (0.612), meaning a substantial component of measured convergence comes from shared architectural inductive biases rather than learning. Convergence is highest precisely where reasoning fails, which the authors say inverts the PRH prediction.
What This Means
The first four programs build the case for geometric structure (with caveats — Park’s LRH tested one model family, and the concept sets are curated). The last two identify where the structure breaks down. Both halves matter:
-
Concept space has measurable geometry. Directions for binary concepts. Polytopes for categories. Orthogonal subspaces for hierarchies. These are empirical findings — measurable, formal, and falsifiable — though tested primarily on one model family (LLaMA-2 7B) with a curated concept set.
-
Convergence is in distance, not in coordinates. Models do not learn the same embedding vectors. They learn the same similarity structure. Two models agree on what is near what, even when they encode it differently.
-
Models are lossy projections. Each model compresses concept space into finite dimensions. Superposition provides a plausible geometric mechanism for hallucination — distinct concepts colliding in the compressed representation — though the connection to production hallucination rates has not been measured.
-
The convergence has resolution limits. Coarse structure (semantic categories) converges. Fine structure (individual items) diverges. The scale trend may not hold for newer models. And some measured convergence reflects shared architectural biases rather than learned structure.
-
Models share the map but not the navigation. Usama & Chang found that LLMs converge strongly in representation (CKA ~0.87 before the model commits to an answer) but diverge sharply in reasoning (CKA ~0.27 after). The thing models don’t share is how they get from representations to conclusions. External memory that carries explicit derivation chains — “X depends on Y, which was derived from Z” — supplies exactly the missing piece.
The Cave and the Map
The Platonic Representation Hypothesis gives the concept space idea its most vivid framing: models are prisoners in Plato’s cave, and concept space is the reality outside.
This is a useful frame for understanding External Epistemic Memory, but the connection needs to be stated carefully. EEM beliefs are text tokens in a context window — they get projected through the model’s embedding layer like everything else. They do not bypass projection.
What they do is supply fine-grained structure that the model’s weights lost. A belief with its justification chain encodes explicit relationships — “X depends on Y, which was derived from Z” — that the model’s finite-dimensional compression dropped. The belief is still a shadow on the wall, but it is a shadow cast by a more detailed object.
This framing fits the empirical results:
Why EEM transfers across models. Sonnet with beliefs matches Opus (agents-python ablation, 55 questions x 4 conditions x 4 models). Haiku with beliefs reaches 94% where Opus alone scores 98% (expert-service dual-path eval). If models share coarse-grained similarity structure (as PRH shows), then semantically dense text activates overlapping regions across model families. The beliefs are not model-independent in any deep sense — any natural-language string is model-independent — but they carry more structure per token than raw source documents.
Why cross-vendor transfer is weaker. Beliefs generated by Opus transfer more effectively to Sonnet (+34.5pp) than to Gemini Flash (+15pp) (cross-model transfer experiment). Different training regimes produce different projection geometries. The beliefs are being projected, and the projections differ across vendors — which is exactly what you would expect if concept space convergence is coarse-grained rather than universal.
Where EEM helps most — and an open question. Our architectural ablation showed +12-14pp on architectural questions but minimal gain on factual recall (beliefs ablation, 2,200 invocations). One reading: factual recall is knowledge where models already converge, while architectural reasoning requires synthesizing multiple observations into cross-cutting judgments — the kind of fine-grained structure that projections lose. But the mapping between Koepke’s coarse/fine distinction and our factual/architectural distinction is ambiguous. Koepke’s fine-grained level is individual items, which could map to specific facts rather than architectural reasoning. This is a discrepancy worth investigating, not a clean confirmation.
Hallucinations may be geometric. Superposition provides a plausible mechanism: models pack more concepts than they have dimensions, and distinct concepts can collide in the compressed representation. EEM beliefs could disambiguate collision points by providing explicit structure that the projection collapses. But this connection is a hypothesis — superposition interference is well-characterized in toy models, not yet linked to hallucination rates in production LLMs.
What Changes
The previous concept space post was written from the inside out — starting from our experience building EEMs and reasoning about why they work. This revision comes from the outside in — starting from published representation learning research and discovering that four programs describe the structure we hypothesized, and two programs identify its limits.
Three things change:
Concept space is falsifiable. It is not just “a useful way to think about beliefs.” Park’s linear representation tests, PRH’s kernel alignment measurements, and SAE feature universality are empirical claims that could be wrong. The measurements so far support the geometry: twenty-six of twenty-seven tested concepts are linearly encoded in LLaMA-2 7B, and SAE features recovered from independently initialized models are more similar to each other than to the raw neurons of either model. But these are findings on specific model families with curated concept sets, not universal proofs. And the scale trend that PRH predicted — kernel alignment increasing with model size — does not appear to hold for newer models (Koepke 2026).
The limits are the product opportunity. Koepke and Usama & Chang are not refinements of the strong convergence thesis — they are partial falsifications. The convergence is real at coarse granularity but breaks down at fine granularity. Some measured convergence reflects architectural biases rather than learned structure. These are genuine challenges to PRH. But for EEM, the limits are the point. Usama & Chang’s dissociation is the most direct: LLMs share the map (representational convergence, CKA ~0.87) but not the navigation (reasoning divergence, CKA ~0.27). If what models lack is not representations but the reasoning chains that connect representations to conclusions, then EEM — which stores explicit derivation chains with justifications — supplies exactly what’s missing. The product opportunity is not in the space where models converge, but in the reasoning gap where they diverge.
The Library is manifold cartography, aspirationally. When we said “each domain EEM is a chart covering a region of concept space,” that was a metaphor borrowed from differential geometry. Concept space now has measured geometric structure (polytopes, orthogonal subspaces, universal features), which makes the metaphor more grounded. But a formal atlas requires charts with smooth, compatible transition maps on overlapping regions, and we have not established that two domain EEMs overlap coherently, let alone smoothly. The Library is heading toward an atlas. It is not one yet.
The Papers
For the full belief network behind this post — 2,400+ beliefs extracted from 50+ papers with justification chains, retraction records, and cross-reference links — browse wiki.llmeem.ai/llm-eem.
| Paper | Year | Key Finding |
|---|---|---|
| Huh et al., “The Platonic Representation Hypothesis” | ICML 2024 | Models converge in kernel (distance structure) across modalities as they scale |
| Park et al., “The Linear Representation Hypothesis” | ICML 2024 | Concepts are directions in activation space; causal inner product formalizes separability |
| Park et al., “Geometry of Categorical Concepts” | ICLR 2025 | Categories are polytopes; hierarchies are orthogonal subspaces |
| Elhage et al., “Toy Models of Superposition” | Anthropic 2022 | Models compress more features than dimensions; superposition is geometric |
| Bricken et al., “Towards Monosemanticity” | Anthropic 2023 | SAEs on independently initialized copies of one architecture recover the same features |
| Koepke et al., “Back Into the Cave” | 2026 | Cross-modal convergence is coarse-grained; fine-grained alignment drops; scale trend does not hold for newer models |
| Usama & Chang, “Convergence Without Understanding” | 2026 | Representational and reasoning convergence dissociate; random-init models show higher CKA than trained models |
Ben Thomasson builds External Epistemic Memory systems. The research wiki is at wiki.llmeem.ai. The tools are open source at reasonsforge.com.