Nisbett and Wilson, in 1977, laid out a table of stockings and asked passersby to pick the best pair. People examined them, chose, and gave crisp reasons — this one was softer, this one better made, this one a nicer knit. There was one problem. The stockings were identical. The only thing that predicted a choice was position: the pair on the right got picked about four times as often as the pair on the left. Not a single person mentioned position. And when the researchers gently suggested position might have mattered, people didn’t go huh, maybe — they denied it, sometimes with a look that said the question was a little strange. They weren’t lying. They genuinely didn’t know why they’d chosen, so their mind supplied a reason that sounded right, and then believed it completely.
It gets stranger. Decades later, in the choice-blindness experiments, people were shown two photographs of faces, picked the more attractive one, and were then handed a photo to explain their choice — except, by sleight of hand, they were sometimes handed the face they had rejected. Most didn’t notice the swap. And they proceeded to explain, fluently and confidently, why they’d chosen the face they had actually turned down — “I liked her smile, she seemed more approachable.” The reason-giving machine doesn’t check what actually happened. It just runs.
What both studies expose is something the brain does constantly: it runs an interpreter, a fast internal narrator whose job is to produce a coherent story for what you just did — after you did it. The story is quick, confident, and often wrong about causes, because the real causes (position, priming, mood, a dozen cues below awareness) are invisible to the narrator too. When you introspect, you don’t get a readout of the machinery. You get the story the machinery wrote about itself.
Now carry that into a user interview, because this is where it stops being a curiosity and starts costing you roadmap decisions. The answers users give about why they did something, or whether they’d use a feature, are among the least reliable data you can gather — and they arrive in the most convincing tone you’ll hear all week. “Would you pay for this?” — yes, enthusiastically, and then they don’t. “Why did you cancel?” — a tidy, reasonable story that isn’t the real reason. “What would make you use it more?” — a confident feature request that, once built, moves nothing. This is the quiet mechanism under the old complaint that “talk to your users” gets ignored. It’s not that talking is useless. It’s that teams keep asking users the exact questions confabulation answers worst, and then trusting the fluent reply.
So the shift isn’t “stop talking to users.” It’s about what you trust. Trust behavior over stated reasons — where people actually click, where they drop, what they actually pay for, is data the interpreter can’t fabricate, because it isn’t a story, it’s a trace. In conversation, aim at the concrete past, not the hypothetical future: “walk me through the last time you tried to do this” pulls a real, recallable episode, while “would you use a feature that…” just hands the mind a blank to confabulate into. Treat confidence as a warning light, not a green one — the faster and smoother the “why,” the more skeptical you should be, because fluency is the signature of a story, not of access. And where you can, run the test instead of taking the poll: a fake-door or a prototype that measures whether they click beats any number of people telling you they would.
The obvious objection — so are user interviews worthless? — deserves a firm no, because overcorrecting here is its own failure. Self-report is unreliable about causes and predictions specifically. It’s genuinely valuable for other things: learning a user’s goals, their vocabulary, the emotional texture of their problem, the shape of the workflow they’re stuck in, what confuses them in the moment while you watch. Talk to users constantly — to understand their world and to watch them struggle — just not as reliable narrators of their own minds or forecasters of their own behavior. The failure mode is narrow and specific: treating “yes, I’d use that” as evidence.
I’ll keep the science honest, because the unsettling version over-reaches. Confabulation and choice blindness are robust, but they don’t mean humans have no self-knowledge. People are often perfectly accurate about their preferences in simple, familiar territory; the unreliability spikes for subtle causes, novel choices, and predictions about the future — which, inconveniently, is most of what product research asks about. So calibrate the skepticism to those zones, and keep it lighter elsewhere.
Which leaves a plain instruction: keep talking to your users, more rather than less. Just stop asking them to explain themselves, and start watching what they do. The truth was never in the answer. It was in the behavior the answer got invented to explain.
Liked this? Get the next one in Working Theory.
Going weekly in August (it's in beta now). One genuinely interesting read on building, the brain, and the science most people missed.