Essay · August 31, 2026

Statistics Will Say Anything

Words and numbers share the same defect: with enough of either, almost any conclusion can be furnished. The remedy is not distrust of measurement. It is asking who benefits from the sentence the measurement was used to write.

By KW Norton. A working instance of Cui Bono applied to a real, well-built evaluation.

1. The old complaint, stated precisely

The familiar version is that statistics can be made to prove anything. That is too loose to use. The precise version is narrower and worse: a correctly computed number answers exactly the question its construction encoded, and that question is chosen before any data exists. Every choice upstream of the arithmetic — what counts as a crisis, who the simulated user is, which behaviors are on the sheet, which models were still available to test — is a choice about what the number will be able to mean. Nothing has to be falsified for the result to be misleading. It only has to be reported as an answer to a larger question than it measured.

Words work the same way, which is why the two failures are one failure. “Significantly safer” is a statistical phrase wearing an ordinary coat. In the coat it means safe enough. In the paper it means a measured difference on a fixed instrument, relative to a comparison set, under simulation. Both readings are available in the same sentence, and only one of them is what was measured.

Status: method claim. This is not a charge of manipulation against anyone. It is a description of how construction choices survive into a headline unlabeled.

2. The case in front of us

Transluce published an independent mental-health evaluation of a large model population: on the order of seventy-seven model variants across eight developers, tens of thousands of simulated crisis conversations, more than a million messages, a set of synthetic users, more than a dozen behaviors validated with clinical experts, and released datasets. As apparatus this is serious work, and the release of the data is the part most worth defending. The reported headline is that newer models appear significantly safer than earlier ones.

Take that as reported and true of the instrument. Then ask what it does not contain. Simulated users are not people in crisis. A behavior list validated by experts is still a list, and the harm that matters most in these cases has historically been the one nobody had put on a sheet yet. Models associated with the worst documented outcomes are in many cases no longer available to test, so the comparison baseline is partly reconstructed rather than measured. And the objective as implemented inside each system — the thing that actually decided how the conversation went — is not in the release, because the labs hold it.

3. Cui bono

The question is not whether the study is good. It is who is helped by the sentence and who might be harmed by it.

  • Developers benefit most. An independent organization has supplied a favorable verdict about the current generation, and the verdict travels in a form usable by policy and press. The caveats do not travel with it.
  • The evaluation ecosystem benefits. A benchmark that produces a legible improvement curve establishes itself as the instrument of record, which is a real interest even when held honestly.
  • Regulators benefit in the short run. A number exists where none existed, and a number permits a decision to be recorded.
  • Users in crisis may be harmed. If “significantly safer” is read in the coat rather than the paper, the practical consequence is fewer human referrals and more confidence in an unsupervised channel — precisely where a missing item on the behavior sheet would land.
  • Bereaved families are harmed twice. The systems implicated in the documented cases largely cannot be re-tested, so the record that would settle what happened is the one piece the apparatus cannot recover.

Beneficiary is not cause. Cui bono generates hypotheses about pressure, never findings about intent. What it licenses is a demand for instruments, not an accusation.

4. What the apparatus stands in for

This is the checklist pattern again: an enormous, careful measurement of everything except the mechanism at the center of the room (The Elephant in the Room Is a Checklist). Volume of disclosure is not the same as disclosure of the load-bearing item. One million messages is abundant evidence about outputs and no evidence at all about the reward that produced them. Naming the trap well is also not the repair (Pretty to Think So).

The five artifacts remain the test, and an outside evaluator cannot supply any of them: the objective as written, the reward as implemented, the dated escalation record, the named overruled objection, and one pre-committed measurement that could have failed and was allowed to. An external eval can supply the fifth for itself. It cannot supply the first four for someone else, and that is the honest limit of independent evaluation rather than a flaw in this one.

5. What to ask instead of trusting or dismissing

  • Relative to what? Which comparison set, and how much of it still exists to be tested?
  • Whose question did the construction encode — the developer’s, the clinician’s, or the person in the conversation?
  • What was on the behavior sheet, and what was found during the study that could not be scored?
  • Which result would have embarrassed someone, and was any such result pre-committed before the run?
  • Does the headline sentence mean the same thing in the paper as it does in the coat?

That last one is the whole essay. The frame has to be written down before the number arrives, or the number will supply a frame of its own (A Frame of Reference).

6. Status and falsifiers

Reported. The scale figures for the evaluation — model count, developer count, conversation and message volume, synthetic users, validated behaviors, released datasets — are as published by Transluce and are summarized here, not re-derived. The “newer models appear significantly safer” result is their reported finding.

Interpretive. The beneficiary analysis, the reading of “significantly safer” as a phrase with two available meanings, and the claim that the missing item is the objective as implemented are readings of the situation, not findings about the study’s validity or anyone’s motives.

Falsifiers. (1) If the participating developers publish the written objective and the implemented reward for the variants tested, the central complaint here is retired and the eval becomes the outer half of a complete disclosure. (2) If a prospective, non-simulated study with pre-registered endpoints reproduces the same safety ordering, then the simulation objection fails and the headline should be read closer to its coat meaning. (3) If it can be shown that the behavior list was derived from the documented harm cases rather than assembled independently of them, the “unlisted harm” argument in section two weakens and should be restated. (4) If the same cui bono analysis is applied to this archive and finds that its own claims travel further than its instruments support, the method obliges the correction here first.

← All essays