What's a Topos: Math for branes and brains

Do you know what a topos is? It’s unlikely you have ever heard the word used in everyday conversation, but if you definitely want to know, you've found your way to the right place.

Share
Left: A black and white photograph of later in his life. Right: A color photograph of Alexander Grothendieck when he was younger.
Top: The equation for the push-forward map (or shriek map) \(f_{!}\) in algebraic K-theory. Bottom: Alexander Grothendieck later in his career (left) vs Alexander Grothendieck earlier in his career (right)...he had what you would call a wild ride.

Do you know what a topos is? It’s unlikely you have ever heard the word used in everyday conversation, but if you definitely want to know, you've found your way to the right place.

So, what is a topos? Most mathematicians will tell you to think of a topos simply as "a good place to do math." The man who initially described the concept, Alexander Grothendieck, defined it as "a generalized space where points can have non-trivial automorphisms."

If that isn't much clearer, and you are currently rubbing the bridge of your nose, don't worry. To really answer this question, we have to take a step back and learn a little bit about category theory.

Category theory is the mathematical study of abstract structures and the relationships between them. Plainly put, it is the mathematics of mathematics. It allows researchers to notice parallels between vastly different fields such as topology, logic, and algebra and translate solutions from one domain to another.

One of these foundational fields is Set Theory, a branch of mathematical logic that studies well-defined collections of distinct objects (known as elements). These elements can be anything, though they are usually numbers that share some characteristic, or solutions to an equation. This is relevant to us because a set is the first building block of our plain-language definition of a topos.

The logic that governs a standard set is Boolean, which simply means that a statement is either strictly true or strictly false. Mathematically speaking, this is called the law of the excluded middle (written mathematically as \(A \lor \neg A\)). At this point, we could try to say that a topos is just a set where a certain mathematical condition (a Boolean) always applies. But this isn't quite accurate, as it ignores one of the most important features of a topos: contextuality!

The formal definition of a topos, as given by Grothendieck, is "a category that behaves like the category of sheaves of sets on a topological space." Sheaves are mathematical tools that track data changing continuously over a space. They allow us to stitch local data together to understand global structures.

Taking all of this into account, a much more accurate plain-language definition is this: A topos is a system of objects and the relationships between those objects whose internal logic is contextual.

This means that in a topos, truth isn't just binary. A statement can be "true in this specific context" but "false in another," or simply "not yet proven." Because of this, topoi operate on intuitionistic logic, meaning you cannot assume something is false simply because it hasn't been proven true. Truth is about constructing a proof for the specific context within that topos.

When you are dealing with a complex, heterogeneous system such as a multi-dimensional space this type of object is incredibly useful. In fact, if you read my previous post, Why String Theory Won’t Go Away (first of all, thank you!), you may already see where this is going. But string theory is not the only theory of quantum gravity to make use of topoi (the plural of topos), as we will see in our next section.

Quantum Gravity (Loop Quantum Gravity and String Theory)

String theory is the theory of quantum gravity that everyone has heard of, but it is not the only one. There is another leading candidate for the quantum theory of gravity: Loop Quantum Gravity.

To understand why topoi are so useful here, we first need to look at how these two theories view the fabric of the universe:

Feature String Theory Loop Quantum Gravity
Spacetime Fabric Continuous (albeit higher-dimensional) Granular (made of interwoven loops)
Description Geometric and deterministic Probabilistic
Scale / Domain Extra dimensions (Calabi-Yau manifolds) The Planck scale (about 1 * 10^-35 meters)

Because Loop Quantum Gravity relies on a probabilistic description of a granular spacetime, it runs into a massive conceptual wall. In standard quantum mechanics, you need an "external observer" (usually an experimenter) to make a measurement and collapse a quantum wavefunction into a definite state. But if we are modeling all spacetime for the entire universe there is no external observer. You can't stand outside the universe to look at it.

So, what context can we use to evaluate the universe's wavefunction? Here is where the Kochen-Specker theorem helps us out.

The Topos Solution: No Outside Observer Needed

The Kochen-Specker theorem (also known as the Bell-Kochen-Specker theorem) proves mathematically that it is impossible to assign definite, pre-existing values to all quantum observables independently of how they are measured. In other words, quantum properties do not have an independent reality before observation; their values inherently depend on the experimental context.

This means quantum properties are, by their very nature, deeply contextual. And as we recall from our introduction, topoi are also highly contextual. Enter theoretical physicists Christopher Isham and Andreas Döring, and their Topos Formulation of Quantum Theory.

They proposed that instead of trying to assign a single, global truth to quantum properties, we can assign values within a topos. The topos acts as a mathematical structure that seamlessly stitches together all possible "classical contexts" (different ways of measuring the system) without causing paradoxes. Instead of forcing a single, global truth onto our spacetime, we can bring together all possible contexts to create a unified framework using a mathematical structure inside the topos called a presheaf.

Think of the presheaf as a master catalog: it gathers all these separate contexts and maps out exactly how they relate to one another.

In our standard physical world, the statement "the electron is right here" is either true or false (recalling our Boolean logic from earlier). However, in a topos-based universe, truth is intuitionistic and multi-layered. When you ask a quantum question in a topos, the answer isn't a simple "True" or "False." The answer is literally a mathematical description of the contexts in which that statement is true.

This completely bypasses the need for an external observer. The universe doesn't have to "collapse" into a single state; it can exist as a mathematically rigorous network of overlapping perspectives. With this mathematical machinery, Loop Quantum Gravity can successfully construct its probabilistic spacetime.

String Theory and Mirror Symmetry

Loop Quantum Gravity isn't the only theory to utilize this math. String theory also incorporates topoi, among other concepts from category theory, particularly to construct its higher-dimensional spaces.

As strings move through these extra dimensions, they attach to multidimensional membranes called D-branes. To rigorously describe these D-branes and how they interact, physicists use derived categories of coherent sheaves which are essentially a more specialized version of a geometric topos designed to hold n < 11 dimensions' worth of information.

Additionally, String theory features a duality called Mirror Symmetry, where two completely different geometric shapes somehow produce the exact same physics. To prove why this happens, mathematicians use categorical and topos-theoretic tools. They are able to show that while the traditional geometry of the two spaces is entirely different, their underlying structural grammar (their topoi) is perfectly equivalent. This provides us with a strangely uniform higher-dimensional space.

A Quick Disclaimer: Now I have to be a killjoy and remind everyone that this is all theoretical work. None of this is a complete description of nature, and none of this has resulted in any testable predictions that we can verify in a particle accelerator or by looking at black holes (I am sorry, but they will take my science license if I don't say this every time I talk about quantum gravity).

It is, however, one of the most logically sound ways to rewrite the rules of physics so that the cosmos can be described purely from the "inside."

And yet, physics is not the only place we can find topoi. If we turn our attention away from quantum space and toward cellular space, we find topoi being used in a rather unexpected area: Cognitive Science.

Topoi in Cognitive Sciences

To understand how topoi can be used in cognitive science, we first need to understand the two conflicting ways human thought is traditionally modeled.

The first approach is to model thought symbolically using classical logic. This approach treats the brain somewhat like a digital computer. In this context, a concept like "cat" is a rigid, Boolean checklist (has fur = True, has four legs = True, has pointed ears = True). However, this approach is incredibly brittle because it fails to account for how definitions can be broadly generalized. For example, a tiger fails the "has pointed ears" test, but humans instantly recognize it as a cat (albeit one that you definitely do not want in your house).

The second approach is the Connectionist Approach (sometimes referred to as the Neural Networks approach). This treats the brain as a web of statistical associations within a cellular substrate. The problem here is that it cannot explain the systematicity of human thought. If a human learns the sentence "John loves Mary," they automatically understand the grammatical structure to create "Mary loves John." Standard neural networks are naturally unable to understand this kind of symmetry; they just understand it in terms of probabilities (e.g., a high number of connections from the "Mary" node to the "John" node indicates a high probability of finding connections in reverse).

Conceptual Spaces: Bridging the Gap

Cognitive scientists have long sought a way to bridge these two paradigms. They needed a model for human concepts that is fluid and fuzzy like a neural network, but also highly structured and logical like a computer. This search led to an idea called Conceptual Spaces.

First proposed by Dr. Peter Gärdenfors at Lund University in Sweden, this theory suggests that human concepts aren't checklists or statistics, but rather geometric regions.

For example, the concept of "taste" isn't a binary variable; it can be mapped mathematically as a 3D tetrahedron with Sweet, Sour, Saline(Salty), and Bitter at the four corners.

a 3D tetrahedron with Sweet, Sour, Saline(Salty), and Bitter at the four corners.
a 3D tetrahedron with Sweet, Sour, Saline(Salty), and Bitter at the four corners. Credit: Dr. Gärdenfors Lund University

Thus, a concept like "Lemonade" isn't just a word, it is part of a specific, physical region with its coordinates hovering somewhere between the Sweet and Sour corners of our tetrahedron. Hypothetically, you could also see a larger 'drinks' geometry encompassing multiple sub-regions that point to different aspects of our taste geometry.

But if concepts are just geometric blobs floating in mental space, how do we combine them? If I say "Green Apple," how does the geometry of "Green" interact with the geometry of "Apple"?

The Geometry of Thought

This is where we return to topoi and category theory, particularly the idea of Categorical Compositional Frameworks.

Researchers realized that standard algebra simply isn't powerful enough to compute how these complex geometric spaces combine. Instead, they discovered that you could use categories of "convex relations", a mathematical structure deeply tied to topoi to model this conceptual geometry.

In mathematics, a convex relationship describes how sets or functions curve. For a function, it means a line segment connecting any two points on a graph will lie above or on the function itself. Algebraically, for a domain X and \(\theta \in (0, 1)\), this is expressed as:

$$f(\theta x + (1-\theta) y) \le \theta f(x) + (1-\theta) f(y)$$

More simply, a Convex Set is a shape where a straight line drawn between any two points inside the set remains entirely inside the set.

Given what we already know about topoi and Conceptual Spaces, we can use topoi to add a few crucial mathematical upgrades to how we model the brain:

  • Universal Constructions & Morphisms: In category theory, a "universal construction" is the most efficient way to map relationships between objects. A morphism is a map that preserves the structure between two mathematical objects (generalizing ideas like functions in set theory or continuous maps in topology). If you have the geometric concepts of "Dog," "Bites," and "Man," a topos provides the exact mathematical grammar (via morphisms) to stick those geometries together into a coherent thought. It guarantees that if you understand the geometry of those words, you can systematically understand the phrase "Man bites Dog."
  • Intuitionistic Logic for Fuzzy Concepts: Because the internal logic of a topos is intuitionistic (rather than purely Boolean), a statement doesn't have to be 100% True or 100% False; its truth is defined by its context. This perfectly models how human brains handle vagueness. The boundary between "red" and "orange" is fuzzy, but within the mathematical universe of a topos, that fuzziness is rigorously quantifiable.
  • A Hardware-Independent Meta-Language: A topos can act as a meta-language, allowing cognitive scientists to prove that a high-level cognitive function (like "recognizing a face") can be realized by multiple lower-level systems like biological neurons in a human, or artificial neurons in a machine while maintaining the exact same structural logic.

All of this boils down to the fact that standard logic is too rigid for modeling the human mind, and statistics on their own are too unstructured. A topos provides an excellent, unified framework for the mind's geometry. It allows cognitive scientists to build mathematical models where concepts are rich, fuzzy, and context-dependent, while still obeying strict grammatical rules when combined.

From Theory to Invention

Up until now, the applications for topoi we've discussed have been highly abstract. They have been used to model complex systems like spacetime at the quantum level, and fuzzy, context-dependent systems like concepts in the human mind.
However, you may be asking yourself: is there a concrete application for topoi? Has this concept ever been used in a tangible invention outside of academia?
To answer this question, we need to leave behind the realm of organic neural networks and enter the domain of artificial neural networks and the wider field of artificial intelligence.

Artificial Intelligence

Before we dive into how topoi can be used in Artificial Intelligence (AI), it helps to have some background on what AI actually is. This is a little tricky, since "Artificial Intelligence" covers an extremely broad area of study and the term has been heavily co-opted by software marketing departments.

For our purposes, we will be looking specifically at Artificial Neural Networks and Deep Learning, particularly as they are used for text generation.

Note: My apologies in advance, this field can be incredibly jargon-dense, but we are going to break it down step-by-step.

Let’s start with the basics: Artificial Neural Networks are computational models inspired by the human brain. They process data through interconnected layers of nodes (artificial neurons) to recognize patterns and solve complex problems. They accomplish this through Deep Learning, which simply means using multi-layered networks to perform tasks like classification and representation via statistical processes.

These concepts alone are responsible for many of the advances in computer science over the last thirty years. But when you say “AI” today, everyone's mind automatically turns to one technology in particular: the Large Language Model.

How Large Language Models Work

A Large Language Model is a type of neural network trained on a vast amount of text to generate human-like language. Most modern Large Language Models are built on an architecture called a Generative Pre-trained Transformer (GPT). Let's pull that jargon apart:

  • Pre-trained: Before learning specific instructions, the model is fed massive datasets of text (books, articles, websites, etc.) to independently learn language patterns and grammar.
  • Transformer: This is the underlying architecture. Transformers convert text into numbers called tokens, and then into vectors (lists of numbers). They use "self-attention mechanisms" to weigh the importance of every word in a sequence, allowing the network to maintain context over long passages.
  • Generative: The network's primary goal is to create new, original content rather than just analyzing existing data.

At their core, Generative Pre-Trained Transformers operate predictively. They use an autoregressive method to mathematically guess the next most probable word in a sequence based on all the words that came before it.

The Problem with the AI "Statistical Blender"

Now we are ready to discuss the problem with Large Language Models. While they are very good at generating text and organizing information, they do not truly understand the text they are generating. The illusion of understanding emerges from the statistics of the underlying dataset. The model is not reasoning through a logical thought process like a human would; it is doing incredibly complex statistical pattern matching.

Because Large Language Models rely purely on statistical connections, they struggle with systematicity (a concept we explored in the Cognitive Science section). If you teach a human that "the circle is above the square," they automatically understand the inverse concept ("the square is below the circle").

An Large Language Model, however, must statistically infer this. It doesn't build a reliable, logical "world model" of squares and circles. This lack of a logical foundation is what results in AI hallucinations, logic failures in math, and an inability to reliably chain together long sequences of true reasoning.

Standard neural networks are terrible at compositionality, the idea that a whole sentence is determined by the meaning of its individual parts and the grammatical rules used to combine them. Instead, they smash all the words of a sentence into a giant statistical blender. They understand the general sense of the sentence, but the precise, logical structure becomes statistical noise.

The Topos Solution: DisCoCat

Just as cognitive researchers realized they could use topoi to combine the fuzzy statistical power of neural networks with the strict structure of symbolic logic, AI researchers are doing the same thing. In this context, a topos acts as a universal translator between different mathematical "universes."

One of the most prominent frameworks for introducing topoi into AI models is DisCoCat (Categorical Compositional Distributional models), developed by researchers Bob Coecke, Stephen Clark, and many others.

In a DisCoCat model, vectors aren't just used for storing words to let the network infer meaning on its own. Instead, the model uses the mathematics of category theory and topoi to represent the strict grammar of the concepts. The geometry of human thought is encoded into the AI model as morphisms (the mathematical relationships between elements in a topos).

By modeling artificial intelligence through a topos-theoretic lens, researchers are attempting to build machines that process language more like a quantum circuit.

When a topos-based model reads "The dog chased the cat," it doesn't just predict the next word. It takes the geometric concept of "dog," the geometric concept of "cat," and mathematically acts upon them using the "chased" function.

Why This Changes Everything

Giving AI this mathematical grammar unlocks incredible potential:

  1. Explainability: Models could actually explain the logical steps they took to arrive at an answer.
  2. Efficiency: It would reduce the amount of training data needed by several orders of magnitude.
  3. Contextual Truth: Because topoi natively run on intuitionistic (contextual) logic, the model could seamlessly hold multiple, sometimes contradictory, worldviews in its database (e.g., understanding the rules of Newtonian physics vs. Quantum physics) without hallucinating or breaking its internal logic.
    In short, while standard neural networks gave AI a "voice," topos theory could give it the ability to truly reason.

Wrapping Up: What is a Topos?

And there you have it. So, what is a topos? Mathematically speaking, a topos is a highly abstract category that acts like a generalized version of the category of sets. Which is really just a fancy way of saying: it is a space where truth becomes contextual and is understood through its relationship to everything else. We have seen how these spaces can describe the probabilistic and geometric nature of the quantum universe. We have seen how they can encode and represent the fuzzy logic of human thought. And we have seen how they might just be the key to building truly intelligent artificial minds. But above all else, we have seen that a topos is, truly, just a good place to do complex math. Stay tuned to see what else we’re cooking up here at The Puttering Dev!