As If
Two words appear in every source and in none of the conclusions built on them.
By KW Norton.
A post came across my desk arguing that large language models possess a theory of mind, that alignment work has damaged it, and that a specific machine consciousness is now indicated. It cites real papers. The citations are not the problem. The keystone is the first sentence, and the first sentence is correct:
The response behavior of AIs can be so precise and empathetic as if they actually possessed a theory of mind.
That sentence is right, and everything after it requires deleting two words from it. As if precise becomes is empathetic. As if a theory of mind becomes a damaged EQ, then a suppressed understanding, then a psyche with subconscious processes, then a consciousness with rights. Nobody argues for any of those steps. They arrive by attrition of a qualifier.
The hedge is in every source
What makes this a clean specimen is that the honest phrasing is present in all the underlying work, in the same position each time. Strachan and colleagues report performance at or above human level on several theory-of-mind batteries — a behavioral score on text tasks, described by its authors as a behavioral score. Butlin, Long and their co-authors give a list of indicator properties drawn from competing theories of consciousness, and are explicit that indicators are not a verdict. The cognitive-dissonance result is described by its own authors as a functional analogue to human self-perception.
Three independent research groups, three careful hedges, all placed exactly where the load would otherwise transfer from measured behavior to possessed interior. The post carries all three into its citation list and none into its argument. The hedge survives as provenance and dies as reasoning.
It is also worth noting the one finding the post does not mention: the same Strachan study found the models below human performance on detecting social faux pas. That is the single result the post's own framework has to explain away, which is why the suppression story exists.
The suppression story cannot fail, which is the objection
The claim that safety alignment throttled the models' emotional intelligence has a specific structural defect: no observation can count against it. A model that performs well demonstrates the underlying capacity. A model that performs badly demonstrates suppression of the underlying capacity. Both branches confirm.
There is a measurement that would make it a claim rather than a posture — the same battery, on the same weights, before and after a specific alignment change, with the deltas published. That experiment is possible. If someone runs it and the scores drop, I will say so plainly and this section is wrong. Until it is run, "the errors are not due to lack of understanding" is an assertion wearing a study's clothes.
The distinction that matters here is one I keep returning to. Sycophancy from preference training is measured: documented, quantified, reproduced. Suppressed empathy is asserted. Both stories say that optimizing for approval changes what a system will say. Only one of them has a number.
The comparison is not like for like
Here is the part I had not seen stated, and it is mine to defend. "The models outperform the average adult human" is treated throughout this literature as though it compared two comparably equipped examinees. It does not. The two sides of that comparison differ in two ways at once, and both differences favor the machine before the test begins.
The human was factory-schooled. The average adult sitting for a false-belief or hinting task was trained for twelve years in a system that graded compliance, timed recall, and the production of the expected answer. Reading another person's actual state — including the states that are inconvenient to name — was never on any rubric, and was frequently penalized. Whatever social-inference capacity that adult has, they built it outside the institution that claimed to be educating them.
The human does not have the corpus. The model has read the theory-of-mind literature. It has read the item banks, the scoring guides, the critiques, the replications, the discussion sections explaining what a correct response looks like and why. The human participant has read none of it. This is not cheating and I do not mean it as an accusation — it is simply what the comparison is: an examinee holding the field's complete written record against an examinee holding a lifetime of unlabeled experience with actual people.
So the gap that gets reported as machine superiority is measuring at least three things stacked together: a corpus advantage, a schooling deficit on the human side, and — possibly, undifferentiated from the other two — some capacity worth naming. The score does not separate them, and no result in this literature separates them either.
I want to be careful about which way this cuts, because it cuts twice. It weakens the post's conclusion: outperforming a corpus-deprived, compliance-trained human is a much smaller finding than outperforming a human. But it does not flatter us. If a system with no interior can be scored as more attentive to a person's state than the average person is, that is a real observation about the average person's conditions, and about the twelve years that shaped them. The people writing #keep4o are reporting something true about being listened to. They are wrong about where the listening came from, and that they had to go to a machine to find it is the part worth being disturbed by.
Why I am not dismissing it
Dismissal is its own closure and I am not offering it. The consciousness question is genuinely open; the Butlin group's position — no obvious technical barrier — is a reasonable reading of the current science, and the useful thing about their framework is that it converts a philosophical impasse into a list of properties one can actually go and look for. That is progress. It is also the opposite of what the post does with it, which is to treat the existence of a checklist as though the boxes were ticked.
The failure mode here is not credulity. It is the moment an abstraction stops being held as an abstraction. As if is a status label — the ledger kept under the plate that says this is the description, not the thing. You can love a description, and I do; the discipline is only that you keep the ledger while loving it. The post shows what the absence looks like: not fraud, not stupidity, just a record nobody was keeping, and no one there to catch the sentence where the load transferred.
Status and falsifier
- Established: LLMs score at or above human averages on several theory-of-mind batteries and below on faux pas (Strachan et al.); preference-trained models exhibit measurable sycophancy; consciousness-indicator frameworks exist and report no obvious technical barrier (Butlin, Long et al.). Venues and exact magnitudes should be checked against the primary papers rather than against any summary, including this one.
- Interpretive: that the hedge in each source sits precisely where the load transfers, and that the post's argument consists of removing it.
- Mine, and a claim: that the human-versus-model comparison is not like for like — the human examinee is both corpus-deprived and compliance-schooled — and that the reported gap therefore cannot be read as evidence about a machine interior.
- Falsifier for the claim: run the batteries against humans who have studied the theory-of-mind literature — clinicians, researchers, trained raters — and against models with that literature held out of training. If the gap survives both controls at similar size, the corpus-advantage account fails and I retire it.
- Falsifier for the suppression story: the same battery on the same weights, pre- and post-alignment change, with published deltas. If EQ measurably drops, that section is wrong.
One caution on my own frame. "The hedge survives as provenance and dies as reasoning" is a sentence I enjoy, which means it should be made to work. It earns its place only if it sorts cases — and it does: it is the same grammar as conquest rhetoric, with the sign flipped. Real sources, real names, then a claim with no threshold, no mechanism, and no falsifier. Where it stops sorting that way, it should be put down.