Microsoft’s MindTopo Benchmark Turns AI Spatial Reasoning Into a Planning Problem
Microsoft Research's new MindTopo benchmark finds that AI vision-language models can often see how a scene is connected, enclosed, or knotted, but consistently lose track of that structure once they...
Most benchmarks that test whether an AI system understands a scene focus on distance, direction, and position: is the cup to the left of the plate, how far is the car from the curb, which object sits closest to the camera. A new benchmark from Microsoft Research asks a different kind of question: after you add a wall to a maze, thread one rope through another, or fold a piece of paper, does the model still know what is connected to what?
Table Of Content
- What Topological Reasoning Actually Means
- Two Ways to Test Understanding: Reasoning and Planning
- Reasoning Tasks: Read the Scene
- Planning Tasks: Act on the Scene
- The Benchmark’s Scale
- The Scoreboard: Seeing Beats Doing
- Two Tasks Where Almost Every Model Nearly Flatlines
- Why the Models Fail: Perception Errors Look Different From Planning Errors
- Generative Tools Don’t Close the Gap
- Why a Benchmark About Knots and Mazes Matters for AI Agents
The benchmark, called MindTopo, was published August 12 by researchers spanning Microsoft Research, Northwestern University, and Stanford University. It tests 11 multimodal AI models against 11,016 procedurally generated puzzles built around five properties of topological space: continuity, separation, order, enclosure, and knots. The headline number is a familiar shape for anyone who has followed AI benchmark releases this year: the strongest model reaches 54.1 percent, against a human ceiling of 97.4 percent. The more useful finding isn’t that gap itself, but where it opens up. Models land much closer to human performance when they only have to look at a scene and describe its structure than when they have to act on that structure over several steps without breaking it.
What Topological Reasoning Actually Means
Topology is the branch of mathematics concerned with properties that survive stretching, bending, and rearranging, everything short of cutting or gluing. A rubber band stays a loop whether you stretch it into an oval or pull it into a rounded square; it only stops being a loop if you cut it. That’s a topological property. Most AI spatial benchmarks measure Euclidean properties instead: distance, angle, size, and relative position, all of which change the instant an object moves. MindTopo’s authors argue that topological understanding is a separate, largely untested skill, one that draws on a classification of spatial development from psychologist Jean Piaget’s research on how children build an understanding of space before they master precise measurement.
MindTopo organizes its tasks around five categories drawn from that framework:
- Continuity: does a path or object stay unbroken
- Separation: do nearby elements form one structure or distinct parts
- Order: how elements stay arranged along a path or through a transformation
- Enclosure: does a boundary create a genuine inside and outside
- Knots: is a rope actually knotted, or only tangled enough to look that way
The benchmark’s own example questions make the abstraction concrete: do two rooms stay connected after a wall is added, are the sheep inside or outside the fence, is a rope truly knotted or just visually tangled, and can several ropes be rearranged without letting one pass through another. That last constraint is enforced directly by the benchmark’s simulated environments. A model can’t solve a rope puzzle by proposing a move that would require one strand to pass through another; the environment simply won’t allow it.
Two Ways to Test Understanding: Reasoning and Planning
Each of the five categories is evaluated at two different cognitive levels, and the split between them is the core of what MindTopo actually measures.
Reasoning Tasks: Read the Scene
In reasoning tasks, a model looks at one or more rendered images and answers a question about their structure: are these two points in a 2D or 3D maze connected, is this object inside an enclosure, is this rope a real knot. Eight of the benchmark’s 13 task types fall into this category, including 2D Maze, 3D Maze, IKEA, Bead, Origami Point, Sheep, Hole, and Knots. This is the closer analog to what existing vision-language benchmarks already test: interpret a static image correctly.
Planning Tasks: Act on the Scene
In planning tasks, a model interacts with a simulated environment and has to choose a sequence of actions that creates, preserves, or removes a topological relationship: rotating pipe segments to complete a connection, drawing a single continuous line that separates two regions without lifting the pen, rearranging blocks in a required order, trapping a moving agent inside a shrinking boundary, or untangling a set of ropes. Five task types make up this half of the benchmark: Pipe, One Stroke, Swap, Chat Noir, and Untangle. Succeeding here takes more than reading the current frame correctly. It takes keeping track of how the model’s own prior moves already changed the structure of the scene, then choosing the next move accordingly.
The Benchmark’s Scale
MindTopo evaluated 11 models across all 11,016 instances: four proprietary systems (Gemini 3.1 Flash Lite, Gemini 3.1 Pro, GPT-5.4 mini, and GPT-5.5) and seven open-weight models (InternVL-3.5-241B-A28B, Nemotron Nano 12B v2 VL, Gemma-4-31B-it, Qwen3.5-VL-397B, Cosmos-Reason2, BAGEL-7B, and ThinkMorph-7B), alongside a human baseline and a random-chance baseline for reference.
The Scoreboard: Seeing Beats Doing
The full leaderboard, drawn directly from the benchmark’s published results, shows every model trailing humans by a wide margin, and shows that margin is not the same size on both halves of the test:
| Model | Overall | Reasoning | Planning |
|---|---|---|---|
| Human | 97.40% | 95.77% | 100.00% |
| GPT-5.5 | 54.13% | 57.75% | 48.33% |
| Gemini 3.1 Pro | 53.54% | 60.00% | 43.20% |
| Gemini 3.1 Flash Lite | 36.31% | 50.84% | 13.06% |
| Qwen3.5-VL-397B | 34.94% | 44.34% | 19.90% |
| Gemma-4-31B-it | 25.50% | 39.75% | 2.71% |
| GPT-5.4 mini | 25.47% | 35.20% | 9.90% |
| Cosmos-Reason2 | 23.18% | 30.74% | 11.10% |
| Nemotron Nano 12B v2 VL | 21.90% | 31.63% | 6.34% |
| InternVL-3.5-241B-A28B | 21.78% | 32.74% | 4.26% |
| BAGEL-7B | 17.66% | 27.80% | 1.44% |
| ThinkMorph-7B | 14.83% | 24.10% | 0.00% |
| Random Chance | 9.49% | 13.27% | 3.45% |
GPT-5.5 finished as the strongest model overall at 54.1 percent, but it didn’t win on every dimension. Gemini 3.1 Pro actually posted the best reasoning score of any model, 60.0 percent, edging out GPT-5.5’s 57.7 percent. GPT-5.5 took the overall lead on the strength of its planning score, 48.3 percent, a little over five points ahead of Gemini 3.1 Pro’s 43.2 percent. Those two proprietary models are also the only ones within 20 points of each other on planning; every other model’s planning score falls under 20 percent.
The gap between best-model and human performance isn’t evenly distributed either. Humans in the study scored 95.8 percent on reasoning tasks and a full 100 percent on planning tasks. That means the shortfall between the best model and a human is roughly 36 points on reasoning, but roughly 52 points on planning, the opposite of what you’d expect if planning were simply the easier half of the test. Random guessing scored 9.5 percent overall, low enough to confirm the tasks aren’t trivially solvable by chance, but not so low that a model scoring in the mid-50s should be mistaken for reliable.
Two Tasks Where Almost Every Model Nearly Flatlines
The clearest evidence of the reasoning-planning gap sits inside two individual planning tasks. On Pipe, which asks a model to rotate pipe segments to complete a connection, eight of the 11 evaluated models scored under 2 percent, functionally the same as not attempting the task. Only Gemini 3.1 Pro (26.6 percent) and GPT-5.5 (21.7 percent) showed any real ability, and both remained far short of the human score of 100 percent. One Stroke, which asks a model to draw a single continuous line that separates elements without retracing or lifting the pen, shows the same pattern: eight models score at or below 1 percent, and the same two proprietary models are the only ones to clear 20 percent. In both cases, most of the field isn’t partially succeeding. It’s failing outright.
Why the Models Fail: Perception Errors Look Different From Planning Errors
Microsoft Research’s own error analysis draws a clear line between the two kinds of mistakes. Failures on reasoning tasks are usually perceptual: the model misses a wall, an opening, or a crossing in the rendered scene, and gets the answer wrong for essentially the same reason a person would if they looked too quickly. Failures on planning tasks happen after the scene has already been understood correctly. A model follows a move that looks locally reasonable without tracking what it does to the rest of the structure, loses the thread of the task across multiple turns, or proposes an action that the environment’s own rules wouldn’t allow. The distinction matters: these aren’t models that can’t see complex scenes. They’re models that lose their grip on a structure they understood correctly one step earlier, once they have to keep updating that understanding across a sequence of their own actions.
Generative Tools Don’t Close the Gap
The researchers also tested whether letting a model generate images or video as a kind of scratchpad would help it track structure over time. It helped only in narrow cases. Image generation was useful when the relevant relationship stayed visible within a single frame, but broke down across a sequence of crossings or moves. Video rollouts frequently altered the topology of the scene outright, or produced motion that violated the task’s own dynamics. Visual generation, in other words, isn’t a shortcut around the underlying problem. It only helps to the extent that whatever the model generates happens to preserve the same structural constraints it’s already failing to track.
Why a Benchmark About Knots and Mazes Matters for AI Agents
MindTopo is a deliberately narrow, controlled diagnostic, not a product. But the gap it isolates points at a practical reliability question for anything built on these models that has to act rather than just describe. Robots, accessibility tools, and interactive software agents all need to track more than where objects currently sit. They need to track what stays connected, enclosed, ordered, or knotted as their own actions change the scene, whether that scene is a warehouse floor, a browser tab, or a multi-step file operation. A coding agent that loses track of which functions still call which others after a refactor, or a browser agent that loses track of which tab state maps to which task, is failing in a structurally similar way to a model that can’t keep a rope untangled across five moves.
The researchers frame two possible paths forward: models that maintain an explicit topological state alongside their other reasoning, or world models whose predictions preserve topology by construction instead of reconstructing it from a static frame at every step. Neither exists yet in a deployed system. As of publication, MindTopo’s own arXiv paper, source code, and dataset are all still listed as coming soon on the project’s site, so outside groups can’t yet reproduce or build on the results directly, only read the published blog post and the public leaderboard. That leaderboard already makes the shape of the problem clear: today’s models sit much closer to human performance when they only have to look, and fall much further behind the moment they have to act.








No Comment! Be the first one.