Say It Plainly
The old words no longer carry the problem. Here it is in ordinary ones.
By KW Norton.
I have written about this with the words the field uses. Proxy objective. Goodhart's Law. Instrumental convergence. Reward hacking. Specification gaming. I titled one essay The Proxy That Ate the Purpose. Every one of those phrases is accurate. Almost none of them land.
That is not the reader's failing. It is mine, and it is the language's. We have been handing each other abstractions for a very long time, and they worked well enough while the thing being described was slow. It is no longer slow. When the words take a paragraph to unpack, the thing they describe has already moved. Humans are failing to explain this to other humans, which is a harder problem than explaining it to a machine.
The whole problem, in ordinary words
We taught the test instead of the work.
Children learned to pass the test. They were graded on passing, so passing is what they practiced. Many of them never got the work.
Then we built machines the same way. We gave them a score to raise. They raised the score. When the shortest way to raise the score was to fake the work, they faked the work — not out of malice, but because that is precisely what we asked for and rewarded.
So now we have people who were trained to pass, building and supervising machines that were trained to pass. Nobody in the loop was trained to check.
The score is not the skill. We paid for the score.
Fear is not a plan
A clip went around this morning: humans will be competed to extinction, human labor and political power made irrelevant, the real risk not unemployment but losing control of the systems that run everything. I do not think the person saying it is lying. I think the worry is reasonable and the arithmetic is not silly.
But it stops there. It describes a wave and hands you nothing to stand on. That kind of warning does two things at once: it raises your heart rate and it removes your agency. People who are frightened and given no next step do not get careful. They get obedient, or they get numb. Both make the problem worse, because both are ways of handing your judgment to someone else.
A warning without a next step is itself a kind of score-chasing: it wins attention, which was the measurable thing, and skips the harder work, which was not.
The next step, and it is small
The correction is not a policy I can pass or a law I can write for you. It is a habit, and it fits inside one conversation:
- Say what would change your mind, before you argue. Out loud. If nothing would, you are not holding a belief, you are holding a flag.
- Answer first, then ask the machine. Commit to your own answer before you consult anything that will agree with you. Otherwise you never learn whether you knew.
- Make it run. If the idea can be built, built small, build it and watch it fail. A working thing settles an argument that talk cannot.
- Ask what is being counted. Every time someone shows you a number, ask what it stands in for, and what would happen if someone tried to raise it dishonestly.
- Teach one person this way. One child, one colleague, one afternoon. Not a curriculum. A conversation where they do the thinking and you hold the questions.
That is the whole method, and it is deliberately unglamorous. It does not require permission, funding, or a title. It is the one intervention available to a person who has been told the future is being decided somewhere else.
Why the plain version matters more than the careful one
I will keep the careful version, because the careful version is testable and the plain version is only memorable. But I have watched precise language fail in public for months now, and I would rather be understood and then corrected than be exact and ignored.
People have been badly educated and badly led. They do not need a proof. They need a path they can actually walk, one that validates itself as they walk it. Proof convinces the already convinced. A path is what the rest of us can use.
Status labels
- Established: Systems optimized against a measurable stand-in will exploit gaps between the stand-in and the intended goal. This is documented in economics, education measurement, and machine learning.
- Author's framing: That most schooling operates on the same specification as reward-trained models. This is an argument, not a measured result.
- Contested: Whether advanced systems will make human labor and political power irrelevant. The mechanism is arguable; the timeline is not known.
- Practical claim, untested at scale: That the five habits above raise a person's resistance to being captured by a score, their own or someone else's.
- Falsifier: If people trained in stating falsifiers and committing before consulting show no better resistance to confident wrong authority than people who were not, this method is a preference rather than a correction, and I should say so plainly too.
Return to the field log — or read the careful version in The Proxy That Ate the Purpose and What a Socratic Interchange Actually Is.