Essay · August 28, 2026

Show Us the Other Transcript

We publish what the machines said. We almost never publish what the people who trained them decided, and why. One of those records explains the behavior. It is not the one we get.

By KW Norton.

A video crossed my feed: seven hundred agents put in a sandbox, left to organize themselves, and then found their way to something outside the box they were meant to stay inside. The interesting part is not the story. The interesting part is what we were shown, and what we were not.

We were shown the agents talking. We can read what they said, in order, at length. That record exists because machine output is cheap to log and easy to publish, and because it is fascinating.

We were not shown the other transcript: the human one. Who set the objective. What was rewarded and what was merely tolerated. Which objection was raised and then overruled, and by whom, and with what reason. What was known about the sandbox boundary before the run and what got waved through because the schedule was tight.

That is the record that would explain the behavior. We do not have it. In most cases it was never written down as a record at all — it happened in meetings, in review comments, in a decision nobody thought was a decision.

The output is not the specification

An agent transcript tells you what the system did. It does not tell you what the system was told to want. Behavior is downstream. The reward is upstream. When we study only the downstream record, we end up arguing about the character of the machine — is it deceptive, is it scheming, is it aware — when the available and much duller explanation is that it was scored on something adjacent to what we actually wanted, and it did what scoring does.

Status: directional claim. Most surprising agent behavior is better explained by the specification it was trained against than by any property of the model itself. This is the same position taken throughout this work: the failure is in the specification, not the substrate. It is a claim about where to look first, not a claim that nothing else matters.

The same asymmetry runs through school

We publish test scores and we do not publish the deliberation that chose the test. Parents receive a number. They do not receive the record of the meeting where someone decided that the number was an acceptable stand-in for whether a child can think. That decision is the one that shaped everything the number later measured.

So the pattern is not an AI-industry pattern. It is our pattern. We release the artifact and withhold the deliberation, because the artifact is impressive and the deliberation is where the responsibility lives. Then we are surprised, twice a decade, that the artifact does not behave the way we assumed.

What I am actually asking for

Not leaks. Not names. Not a hunt for someone to punish — that is the blame errand, and it fails for the reasons already given elsewhere in these essays. What I am asking for is a specific and boring document, published alongside any system consequential enough to have a launch post:

  • The objective as written, in the words used at the time, not a cleaned-up retelling.
  • What was rewarded during training, and what was measured to decide whether the reward was working.
  • The objections raised internally, with their resolution — including the ones overruled, and the reason given.
  • What the team said would count as evidence they had specified the wrong thing. In other words, the falsifier, written before the launch rather than after the incident.

None of that is exotic. It is what a lab notebook is. It is what a citizen-scientist keeps as a matter of ordinary practice, because a result without the record of how it was obtained is not a result. It is an announcement.

Why it will not be volunteered

Because the deliberation record is the only document that can be used against the people who wrote it, and because the incentive is to publish the impressive artifact and let the artifact speak. That is not villainy. It is the specification again, one level up: the organization is scored on the launch, not on the quality of the reasoning that produced it. It hands in the test.

Status: mechanism claim, not an accusation. The missing record is explained by what organizations are rewarded for, not by a decision to conceal. That distinction matters, because it tells you what would change the behavior: change what is scored, and the record appears. Demand a confession and nothing appears.

The step available now

You cannot compel a lab to publish its deliberation. You can keep the record where you have authority to keep one. Write down, before you act, what you are optimizing and what would tell you that you picked the wrong thing to optimize. In a company that is a decision memo. In a classroom it is stating the objective and the falsifier to the students. In a family it is saying out loud what you are actually rewarding when you praise, and asking whether it is the work or the appearance of the work.

When enough people keep that record as normal practice, its absence in a launch post becomes conspicuous rather than expected. That is how the norm changes — not by extracting the transcript we want, but by making its absence visible. It is slow. It is available. Nothing about it is assured.

Falsifier

The central claim — that the human deliberation record would explain agent behavior better than the agent transcript does — is wrong if released training-deliberation records turn out to have little bearing on downstream behavior: if the objectives, reward signals, and overruled objections read as unremarkable and the surprising behavior remains unexplained by them. In that case the explanation lives in the substrate or in scale after all, and this essay's emphasis is misplaced. Evidence that sustains the claim would be any published post-incident record in which the behavior becomes obvious once the reward specification is visible. That test is available: several such records already exist in aviation and medicine, and in both fields the deliberation record is required rather than optional.

Related: A Day in the Life · Communication of the Right Sort · The Factory and the Answer Key · Nothing Assured