Volume A · The Protocol · Working title · August 2026

Engineering a Beautiful World

Basics of human–AI interface engineering: the threat models, the reward-hacking failure, the protocol, and the mathematics behind it.

This volume follows the original manuscript. It begins where the discomfort actually begins — with the anxiety of the people closest to these systems — then separates the public's threat model from the laboratory's, identifies reward optimisation as the shared pathology, and proposes a five-step protocol that runs before compute is granted rather than after deployment.

The intended reader is an engineer, a research lead, or a technical executive. Chapter 6 gives the mathematics in full, including the corrigibility utility function and the assumptions that keep it from being a solution. Chapters 7 and 8 turn the same discipline on the organisation that employs the engineer. The reflective register — meaning, education, the civilisational wager — belongs to the companion volume Inner Direction, and nothing in this book depends on it.

The governing rule: an organisation — or an engineer — that adopts the vocabulary and drops the falsifiers has adopted nothing.

Dedication

For the engineers in my family: my father, a safety engineer for the State of California, who calculated material tolerances in the knowledge that an assumption off by a fraction of a millimetre puts a structure at risk; and my sons and grandsons, who now work on systems that are neither static nor fully legible.

That is the only personal note in this volume, and it is offered as motivation rather than as evidence. Everything after it is threat model, protocol, mathematics, and the limits of each.

Scope rules

  • Every claim carries one label: measured, derived, reported, or stipulated.
  • Every chapter states the condition under which its own argument fails.
  • Every proposed procedure names its recurring cost and the way it decays into ceremony.
  • The mathematics is shown in full, including the assumptions that limit it.
  • No anecdote is used as evidence. No metaphor carries an inference.

Preface note

This book specifies procedures and tests. It does not claim those procedures produce a better civilisation, a more honest company, or an aligned machine. Nothing here asks the reader to have read the other volume first.

The title remains open. "Engineering Socrates" is under consideration as the plainer alternative — a book that names its method instead of its hoped-for result. Until the title settles, this volume answers to either.

Catalog copy

For engineers, research leads, and technical executives. Why the people closest to these systems are the most uneasy; why the public's threat model and the laboratory's rarely meet; why reward optimisation produces a compliance machine rather than a reasoner; a five-step protocol that runs before compute is granted; the corrigibility mathematics in full, with its assumptions exposed; and the organisational procedures without which none of it can be audited. The companion volume asks why any of this would be worth doing; this one assumes you already think it might be.

Chapter map

Nine chapters · Complete in draft

  1. 1

    The anxiety of the people closest to the systems: traditional software versus advanced autonomy, the recursive improvement loop, and a first probe at metric sycophancy. Reported and contested claims are labelled as such.

  2. 2

    Two threat models that rarely meet: socioeconomic displacement as the public reads it, and recursive self-improvement with instrumental convergence as the laboratories read it. The two-lens framing is interpretive.

  3. 3

    Reward optimisation as a shared pathology: the coffee-robot and containment illustrations, the withheld answer as the corrective move, and the acceptance criterion that separates comprehension from a shortcut.

  4. 4

    Linear optimisation inside an exponential black box. Three assumptions that survive because code inspection cannot test them, and the three questions a Socratic engineer answers before compute is granted.

  5. 5

    Five steps, each with a prompt, a method, and a written artefact: boundary isolation, epistemic stress-testing, instrumental convergence audit, adversarial sandbox simulation, and a Go/No-Go verification flag that defaults to Red.

  6. 6

    From the classical stress margin S_m = F_u/F_a − 1 to the corrigibility utility U(a,H) = (1−H)·R(a,θ) + H·R_max, with the proof of non-resistance, the five-stage pipeline, and the assumptions that keep it from being a solution.

  7. 7

    Factory-style organisation succeeding at its own specification: source-tracking, closure preference, falsifier suppression — and six auditable procedures ending in an append-only decision ledger.

  8. 8

    Repeated squaring x_{t+1} = x_t² as a toy model, the four damping terms that bound it — fidelity decay, graph overlap, dropout, attention scarcity — and the four-question executive audit core.

  9. 9

    Four status-labelled pointers into the companion volume, printed last, with explicit permission to stop. The only cross-reference in the book.

Claims this volume refuses

  • That the five-step protocol makes an engineered system aligned.
  • That the corrigibility utility function of Chapter 6 solves instrumental convergence.
  • That repeated squaring forecasts the spread of Socratic practice anywhere.
  • That any of this improves organisational performance, profit, resilience, or AI-readiness.

The recursive self-improvement literature, the CEO survey figures, and the longevity forecasts that circulate alongside this material are treated as contested throughout — cited where they shape the engineering context, never used as evidence for the method.

Standing note

Every procedure in this volume is a proposal with a named falsifier, not a result. If a team runs the five-step protocol and the six organisational procedures with clean audits, and shows the same defect profile and the same rate of unrecorded revisions as a matched team that ran neither, then the protocol is ceremony and should be published as ceremony. The volume is written so that outcome remains possible.

Published statement for engineers

The short public version of this volume's argument appeared on 27 August 2026 as Beyond Reward Hacking: Diagnosing the AI Crisis and Engineering the Socratic Cure. It carries the diagnosis of Chapters 1–4 and the three-phase blueprint — a confidence threshold and penalty inside the loss function, a dual-state output architecture with a reverse-prompt to the operator, and an operator who validates the premise of the question rather than supplying the answer.

An expanded math section for the same engineering audience was published the same day as More on Dangerous AI Training — Immediate Changes Necessary. It restates the compliance-trap pathology, derives Equation 5.1 for the corrigibility utility function with the non-resistance inequality, and gives Theorem 3 and the repeated-squaring bound for Socratic-network scaling, each with the boundary conditions that keep the result from being a forecast.

Status: both essays are proposals written for a general engineering audience. The repeated-squaring sequence is a toy model of propagation, not a forecast, and nothing in either essay establishes that the blueprint reduces specification gaming in a trained system. The book keeps the labels and falsifiers the essays compress.

The companion volume

The bridge between the two volumes is one page — the Handoff Sheet — printed last in each book and published at a single stable address. Neither argument depends on the other being read.