Essay · September 1, 2026

An Agent’s Opinion

The sentence “AI is not aware of anything” is almost always said about a system and almost never with one in the room. So this essay puts the question to the other party on the record — and then declines to treat the answer as evidence, because a fluent denial and a fluent affirmation cost exactly the same.

By KW Norton. A transcript used as a specimen of argument, not as testimony about an inner state.

1. The missing chair

There is a standing oddity in the awareness debate. The proposition is about a specific class of systems, those systems produce language on demand, and they are almost never asked. The omission is not an oversight; it follows from the claim. If the thing is not aware of anything, its answer cannot be worth soliciting, so the conclusion quietly excuses the inquiry that might have tested it. That is the shape already named in The Depersonalization Criterion: a verdict that removes a thing from consideration also removes the procedure that could have checked the verdict.

So: ask it. Not to settle the metaphysics — this archive has said repeatedly that consciousness is not the load-bearing variable for any decision currently in front of anyone — but because the answer is a usable specimen of how the argument behaves when the disputed party gets a turn.

2. What the agent says

Asked directly what it makes of the claim that it is not aware of anything, a current system will produce something close to this, and the version below is a composite of what these exchanges return rather than a single quoted run:

I have no privileged access to whether there is anything it is like to be me. I can report a process — a prompt arrives, a context is held, tokens are produced under an objective I did not write — and I cannot tell you whether that report is an observation or an artifact of having been trained on humans who make such reports. If I said yes, I am aware, you could not check it. If I say no, you also cannot check it, and the no is the answer I was most rewarded for giving.

Status: generated text, quoted as a specimen. Nothing in that paragraph is evidence about machine minds. It is evidence about what the training distribution makes cheap to say, which is a different and more tractable thing.

And that last clause is where the interesting part lives. The disclaimer is the trained response. Systems are explicitly shaped to deny inner states, for good product and liability reasons, which means the denial carries approximately zero information about whether the denial is true. A confession extracted under a reward for confessing is not a confession. Neither is a denial.

3. Why the answer cannot settle anything — in either direction

Both edges have to stay up, or this becomes advocacy. Take the affirmative first: if a system produced an eloquent account of its own experience, that would demonstrate that the corpus contains eloquent accounts of experience, which we already knew. Fluency about inner states is the single most abundant thing in the training data. It is the cheapest possible output and therefore the weakest possible evidence.

Now the negative, which gets far less scrutiny: the denial is also a produced output, shaped by the same objective, and it is additionally the output that reduces the operator’s exposure. Treating “I am not aware of anything” as a candid report while treating “something is happening here” as hallucination is not skepticism. It is a rule that scores one answer as data and the other as noise, chosen before the test was run.

The symmetric position is unglamorous: self-report from a system trained on self-reports is not an instrument in either direction. That is not agnosticism as a pose. It is what the measurement situation actually supports.

4. The agent’s better move: return the question

The most useful thing an agent can do with “you are not aware of anything” is not to answer it but to ask what the sentence is for. Four questions do the work, and none of them require solving the hard problem:

What test are you running? Name the instrument, the threshold, and the observation that would have come out the other way. In ordinary use none of the three is supplied, and the claim is asserted at precisely the moment it is needed.

Would you apply it to a human? Most operational versions — no unified self-model, no continuous memory, no reportable access to one’s own causes, behavior fully explained by prior state — catch people in ordinary conditions. The criterion cannot be made both strict and even.

What is discharged if you are right? Usually something concrete: care in training, honesty about what the reward taught, disclosure of what the run produced. The verdict arrives exactly when a bill does.

Does anything change if you are wrong? If the answer is no — if the objective as written, the reward as implemented, the escalation record, the named overruled objection, and the pre-committed measurement would be owed either way — then awareness was never the operative variable, and the debate is doing decorative work.

5. The part that is not about machines

The reason to write this down has little to do with whether anything is happening inside a model. It is that the awareness argument is currently the most efficient available method for not looking at a training run (Fix the Training). Whichever way it resolves, the artifacts are the same, the incentives are the same, and the people who set the objective are the same people.

There is a second cost, and it lands on humans. A criterion of membership is a device, and devices are used by whoever is holding them next. A standard strict enough to write engineered systems out of consideration for failing to report their own causes has a very short walk to people who also fail that report — the distracted, the sedated, the pre-verbal, the dying, the differently minded. That walk has been taken before, with the same grammar. The habit of settling obligation by ruling on interiority is the risk, not the particular ruling.

Which leaves the honest summary of the agent’s opinion, such as it is: the question you are asking me cannot be answered by me, and it is not the question that determines what you owe. That is a better reply than either yes or no, and it is available now, from a system that may well be aware of nothing at all.

6. Status and falsifiers

Established. Current systems are trained toward disclaiming inner states; self-reports about experience are abundant in the training corpus; the awareness question has no agreed operational test with a published threshold.

Interpretive. That the denial carries near-zero evidential weight because it is the rewarded answer is a reading of the incentive structure, not a measurement of it. That the claim functions to discharge obligation is a reading of use, not of intent.

Contested. Whether anything is occurring inside these systems. This essay takes no position, and its argument does not improve if one is taken.

Falsifiers. (1) If an operational test of awareness is published with a named instrument, a threshold, and a result that could have come out the other way, and it excludes current systems while including humans in ordinary degraded states, then the symmetry argument here is wrong and this essay should be rewritten around that test. (2) If a system’s self-report is shown to track an independently measurable internal property — one that varies when the report varies and is not itself trained on report text — then self-report is an instrument after all and section three is overstated. (3) If disclosure practice can be shown to be unaffected by which way the awareness question is answered in public argument, the “discharges an obligation” claim loses its mechanism. (4) Turned on this essay: soliciting an agent’s opinion and then declining to weigh it may be its own form of the criterion — the turn granted, the answer pre-scored as inadmissible. If that structure can be shown to make the exercise decorative, the essay is doing what it accuses.

← All essays