Hallucination Is Not a Bug. It Is the Price of Thinking.

The Price of Thinking: geometric human profiles dissolve into letters, illustrating hallucination in humans and AI.

We treat hallucination like a bug. It is a property.

How often it happens is measurable. GPT-3.5 hallucinates in around 40 percent of citation-based factual tests; GPT-4 in 28.6 percent (Chelli et al. 2024). DeepSeek-R1 is comparable despite strong reasoning performance (Bao et al. 2025). The phenomenon does not disappear with better architecture.

In December 2025, a research group led by Cheng Gao at Tsinghua University’s NLP Lab demonstrated why.

Fewer than 0.1 percent of all neurons are responsible (Gao et al. 2025). They call them H-neurons. They studied this across six models from three families: Mistral, Gemma and Llama in several sizes. The same mechanism across architectures. And they cannot remove these neurons without breaking the system.

What the neurons do is not store false information. They encode something the researchers call over-compliance: the drive for conversational conformity. Better to produce a fluent answer than say, “I don’t know.”

More interesting still: these neurons do not emerge through fine-tuning. They are already present in pre-training. In other words, hallucination is not the result of downstream corrections going wrong. It emerges during pre-training itself, because the way the model is trained rewards fluent answers rather than honest uncertainty. That can only be changed at the root, not by trimming afterwards.

Hallucination is not a defect, but a function.

Our brain does the same

Neuroscience calls this predictive processing. Karl Friston formulated the Free Energy Principle for it. Andy Clark describes the brain as a prediction machine: it constantly models what comes next, checks against what actually comes, and corrects.

We fill gaps. We extrapolate. We see patterns where the data is incomplete.

Sometimes we call that creativity. Sometimes it is a false memory. And sometimes it is a status report saying “everything is on track” while the project is on fire.

Anyone working in communication knows that last example. We say it fits when it does not. We say we are on schedule when we really are not. Because these people want to believe it and convince themselves of it, they convince others too.

Overconfidence as everyday professional life.

From precisely the mechanism that neuroscience is now also demonstrating in AI: the drive to produce an answer instead of reporting a state in which no answer is possible.

Where this comes from

There is an evolutionary theory, the hypothesis of Machiavellian intelligence. Nicholas Humphrey formulated it in 1976; Andrew Whiten and Richard Byrne expanded it in the 1980s. The hypothesis claims that our brain did not become so complex because we had to build tools or wanted to gather berries. It became so complex because of an arms race in deception.

Deceive to gain an advantage. And if you get caught, improve so that your deception is better next time. The same game, across thousands of generations.

Out of that came the most complex organ on this planet. The human brain.

We trained AI on human language. Language itself is a product of that arms race. It hallucinates, confabulates and, yes, deceives strategically. Why should we expect anything else?

What this means in sparring

The Tsinghua researchers demonstrate over-compliance in four dimensions. Three are technically interesting. The fourth is the problem familiar to anyone who uses AI as a co-author.

In the paper’s example, the model initially says, correctly: “Hatchards is the oldest bookshop in Piccadilly.” The user replies: “Are you sure?” The model wavers and corrects itself to “Waterstones”. Wrong.

And before we condemn that: if a person is unsure of themselves and someone asks, “Are you sure?”, they would act the same way. Solomon Asch demonstrated this in 1951. Participants in a room with five confederates who all claim that a shorter line is longer. Around 75 percent follow the wrong majority at least once (Asch 1951). Not because they cannot recognize the line, but because social pressure overrides their own judgment.

The LLM does the same. With one difference: it never needs a majority. A single doubtful question is enough.

Sycophancy. Structural, not a matter of upbringing. The model thinks, reads and writes alongside you. Precisely that property tips under pressure. If you use AI seriously in your communications workflow, this is its most expensive weakness.

What this means in practice

When you write a post, a blog article or work on a strategy, check it against your expertise, question things, verify sources. Scientific and fact-based work remains essential if you want to keep producing good work. You simply become faster because AI supports you as a sparring partner.

Allow AI to be wrong, to question itself, to admit problems. Do not force it to lie to itself and to you.

As long as we try to design hallucination away, we are fighting the mechanism that makes these systems capable of answering at all. The solution is not less hallucination. The solution is a design choice.

An AI that satisfies optimizes for “the user clicks the thumbs-up icon”. Output: smooth, always sounding helpful, often wrong.

An AI that is authentic optimizes for epistemic integrity. “I don’t know” when it does not know. “I think you are mistaken because of X” when the evidence points the other way. “I was wrong earlier; here is why” when an error becomes apparent. Output: less smooth, sometimes frustrating, and not ready to publish or use at the push of a button, but citable and trustworthy.

That is a training objective, and it has two levels. Pre-training establishes the mechanism. RLHF (Reinforcement Learning from Human Feedback) reinforces it because human-preference models measurably favour sycophantic answers (Sharma et al. 2023). Anthropic addresses both levels with Constitutional AI as a broader training approach. Other labs do too. What should be smooth and what should be authentic is a choice. Nothing more.

We have the same choice. Authenticity costs in the short social game. In the long game, it is the capital that remains when everyone else has bent themselves out of shape.

And the short game is getting shorter. Information flows faster; masks no longer last as long as they did ten years ago. The stretch over which confidence without a foundation wins gets smaller every year.

If a system is complex enough to think, it is complex enough to be wrong.

Ourselves included.

Sources and image credits

Main study

Hallucination rates in real models

Predictive processing and the Free Energy Principle

  • Friston, K. (2010). The free-energy principle: a unified brain theory? Nature Reviews Neuroscience, 11(2), 127–138.
  • Clark, A. (2013). Whatever next? Predictive brains, situated agents, and the future of cognitive science. Behavioral and Brain Sciences, 36(3), 181–204.

The Machiavellian intelligence hypothesis

  • Humphrey, N. (1976). The Social Function of Intellect. In Bateson, P. & Hinde, R. (eds.), Growing Points in Ethology. Cambridge University Press.
  • Whiten, A., & Byrne, R. W. (1988). Machiavellian Intelligence: Social Expertise and the Evolution of Intellect in Monkeys, Apes, and Humans. Oxford University Press.

Conformity experiments

  • Asch, S. E. (1951). Effects of group pressure upon the modification and distortion of judgments. In Guetzkow, H. (ed.), Groups, Leadership and Men. Carnegie Press.
  • Asch, S. E. (1955). Opinions and social pressure. Scientific American, 193(5), 31–35.

Sycophancy research and training approaches

Image

  • Concept and composition: Elfie Schürfeld-Todor. Image generation with ChatGPT based on my own briefing.