a preprint · arXiv 2610.09087
Not just how unsure a model is, but what kind of uncertainty it has.
A language model can be unsure for different reasons. U-Space reads which one it is, straight from the model's internal workings, and does it for every word the model writes. The flight you are about to take follows one real answer.
the run
Was Hamlet left-handed, and did he love Ophelia?
We put this to an open model (Qwen3.5-27B). For every word of its answer, we look at the model's internal state and place that word inside the pyramid. Keep scrolling to follow the answer from its first word to its last.
the four corners
A word sits near the corner that matches the kind of doubt behind it. A word in the middle means no single reason stands out. The small map in the corner always shows where you are.
first half of the answer
Shakespeare never says whether Hamlet was left-handed. While the model writes that part of its answer, the path stays close to the missing information corner: the model is unsure because a fact is simply not there.
a single word
At the word play the path leans toward ambiguity for a moment: the play, or playing a role? The model's internal state carries that double meaning even though its written answer does not show it.
second half of the answer
Did Hamlet love Ophelia? The play gives evidence both ways. As the model answers this half, the path swings across the pyramid to conflicting evidence and stays there to the end.
at the end of the reasoning
We call it the U-Lens score: how strongly the model's internal state leaned toward doubt, combined with how spread out its word choices were. It comes from the single pass that produced the answer. No extra runs, no training, no hand-labelled examples.
Keep scrolling for the method, the numbers and a live demo.
Explore
We asked the model three questions. Pick one, then click any word of its answer: the scene behind this page jumps to that word, and the numbers below show what the model's internal state said at that moment. Nothing here is hand-placed; it is all read from the model.
The four corners
How do we know which corner is which? From words. We collected the words a model uses when it is unsure, sorted them into four groups, and contrasted each group with the words the model uses when it is certain. No examples had to be labelled by hand.
Solid chips are a few of the doubt anchors; dashed chips are a few of the certainty poles they are contrasted against. The full lists are in the repository under data/uspace_terms.yaml.
How it works
While a language model writes, each layer of the network holds a long list of numbers: its internal state. On its own that list is unreadable. A recent tool called the J-Lens translates it: for any layer, it tells you which words that state is pushing the model toward. That makes the inside of a model something a person can read, and it is the foundation we build on.
Using the same translation in reverse, we take the words of doubt from the four groups and find the directions inside the model that correspond to them, each measured against the words of certainty. Four directions, one per kind of doubt: that is U-Space. Nothing is trained and no correct answers are needed to build it.
For each word the model writes, we measure how much its internal state leans along each of the four directions. The four amounts place the word inside the pyramid. How strongly it leans overall tells us how much doubt there is at that word. This is the U-Lens: the path you flew through at the top of the page.
When the model has finished reasoning, we combine how strongly its state leaned toward doubt with how spread out its word choices were. That gives one score for the whole answer: higher means more likely to be wrong. The corner the path favoured says what kind of doubt it was. All of it comes from the single pass that produced the answer.
U is the orthogonalised basis of the four doubt-minus-certainty contrasts pulled back through the lens; ν normalises the residual state. Full details and ablations are in the paper.
Does it work
We gave three open models thousands of questions from four hard benchmarks and checked each answer. Then we asked every method to spot the wrong ones. The score is the usual one for this: 0.5 means guessing, 1.0 means perfect. "Length-matched" compares only answers of the same length, which removes a trick we explain below.
In practice
Here are all the answers from one benchmark, lined up from least to most doubtful according to U-Lens. Teal answers turned out right, pink ones wrong. Drag the slider to choose how many of the most doubtful answers you would send to a human expert, and see what that buys you. These are the real scores from the paper.
And U-Space tells the reviewer where to start: for every answer handed over, which corner it leaned toward, so they know whether to look for a missing fact, a second reading of the question, or a conflict between sources.