Essay · HAIIE & Method · July 31, 2026
The Human-Centered Alibi
The humane case for AI is sound — which is why the most accomplished users of it are the people it was written to restrain. An audit for telling the earned version from the borrowed one.
By KW Norton.
Two companion essays on this site argue that the humane option and the durable-margin option point the same direction — The Profit Case for Humane AI and The Quantum Comprehension Premium. This third one takes the obvious objection seriously: that argument is now standard corporate vocabulary, and the firms most fluent in it are often the ones consolidating fastest. An argument that cannot be failed cannot be used to hold anyone to anything. So the work here is to give the humane claim a set of teeth — operational questions with dated, documentary answers that a well-resourced narrator cannot produce on demand.
A condensed, printable version of the audit below appears in the Due Diligence Field Guide.
01
The argument is sound, which is exactly the problem
Why the humane case is the most useful sentence a consolidator can borrow
There is a genuine finding underneath the consultancy language: enterprise AI programmes fail more often for organisational reasons than for model reasons. Adoption stalls when people cannot trace a recommendation, when the system does not know internal vocabulary, when nobody can say which human is accountable for the output. That is not marketing. It matches what anyone who has watched a deployment from the inside already knows.
The danger is that a correct diagnosis is also a portable one. "Humans in the loop" is the single most transferable sentence in the current market, because it can be attached to almost any programme without changing the programme. It absolves, it reassures the board, it satisfies a procurement checkbox, and it costs nothing to say. The techno-fabulists are not bad at this argument. They are the best at it, because they have the most to gain from being trusted while consolidating.
So the task is not to argue against human-centred AI. It is to make the phrase falsifiable — to write down what a firm would have to be doing differently if the claim were true, so that the claim stops functioning as an alibi.
- Established — Enterprise AI adoption failures cluster around trust, context, and workflow integration rather than raw model capability; this is documented across vendor and consultancy reporting.
- Licensed inference — That an unfalsifiable humane claim will be adopted preferentially by the firms whose behaviour it least describes, because it is cheapest for them to assert.
- Asserted — That publishable operational tests — the audit below — materially reduce the alibi's usefulness. This is the essay's testable claim.
02
Reading the four-part leadership framework literally
Model the behaviour, redesign the workflow, flatten the structure, preserve the judgement
The current leadership playbook, as circulated in early 2026, has four moves: senior leaders should personally use the tools rather than delegate the vision; roles should be redesigned as routine work shifts to agents; management layers should be reduced because agents widen a manager's span of control; and human leaders should retain responsibility for strategy, ethics, and hard calls.
Three of the four are defensible on their own terms. Leaders who use the instrument understand its failure modes. Role redesign is honest about the fact that the work changes. Retained human accountability is the only version of this that survives contact with a regulator.
The fourth is a load-bearing wall dressed as a productivity note. "Flatten structures" means fewer people between a decision and its execution — which is also fewer people positioned to say no. Span of control is not a neutral quantity: middle management, for all its cost, is where an organisation historically stored its friction, its institutional memory, and its dissent. Remove the layer and you have not only removed overhead; you have removed the mechanism by which the third principle — preserved human judgement — was actually exercised. Principle four quietly defunds principle three.
Note what happens to the sequence when it is stated plainly: the humane principles justify the adoption, the structural principle harvests it, and the accountability principle is left with no organ to act through. That is not hypocrisy. It is an unexamined interaction between two recommendations that were written by different concerns.
| Stated move | The defensible reading | The dominance reading |
|---|---|---|
| Leaders model AI behaviour | Decision-makers learn the instrument's actual error surface instead of its brochure | Enthusiasm at the top becomes a loyalty test below; scepticism reads as low digital fluency |
| Redesign roles for human-plus-agent work | Honest re-specification of what a job now is, with training attached | Reduction announced as redesign; the learning budget never lands |
| Flatten structures via wider spans of control | Fewer redundant approval hops on low-stakes decisions | Removal of the layer that carried refusal, escalation, and institutional memory |
| Preserve human judgement on ethics and hard calls | Named accountability that survives audit and litigation | Accountability asserted at the top and delegated nowhere, so no one can exercise it in time |
| Outcome-based commercial models | Vendor carries measurable risk instead of billing for effort | Outcome defined by the party who also measures it |
03
Contextualisation is not the same as comprehension
Grounding a model in your vocabulary does not ground it in your physics
The vendor version of the humane strategy usually resolves to four principles: contextualisation, collaboration, transparency, and continuous learning. Each is a real engineering requirement. None of them is comprehension in the sense used in the companion essays.
Contextualisation means the system knows your product names, your workflows, and which internal expert has solved the problem before. That is retrieval quality. It reduces irrelevance. It does not tell you whether the underlying claim about the world is true, whether the benchmark survives a tuned classical baseline, or whether the biological target is a distribution rather than a structure.
This is where a knowledge-graph vendor and a frontier-lab founder can use the same word to mean two different sizes of thing — and where a buyer can be fully satisfied on human-centredness while remaining short on all three comprehensions. A firm can have excellent internal expert routing and still be mis-pricing its roadmap by two years.
Keep the two audits separate: contextualisation asks whether the answer is relevant here; comprehension asks whether the answer is true anywhere.
- Established — Retrieval grounded in organisation-specific knowledge measurably improves relevance and adoption of enterprise assistants.
- Working claim — Relevance gains are routinely reported as accuracy or reliability gains, which they are not, because the evaluation set is internal.
- Conjecture — That the conflation of relevance with truth is a leading indicator of the comprehension deficits priced in the companion essay.
04
Six questions that cost a fabulist something to answer
An audit, not a values statement
The test of a humane AI strategy is whether any of its commitments is expensive to fake. The following six are. Each has a documentary answer, a date, and a party who can be held to it.
First: name the last decision a human overrode, with the date and the person. A firm that has preserved judgement has a log of refusals. A firm that has only claimed it has none, because refusal leaves paperwork and alibis do not.
Second: after flattening, who now carries escalation? If the answer is a channel rather than a named role with time in their week to use it, the layer was removed and the function was not rehoused.
Third: show the training spend per redesigned role, and how it moved year over year. Role redesign without a learning line is headcount reduction with better vocabulary.
Fourth: who defines the outcome in the outcome-based contract, and who measures it? If those are the same party, the model has transferred risk in the slide and retained it in the ledger.
Fifth: what does the system refuse to answer, and where is that list published? A model with no published refusal surface is a model whose limits were never characterised — which is the operational signature of sycophancy, priced elsewhere on this site.
Sixth: which claim in the last public roadmap has since been retired, and on what evidence? A firm that has never retired a claim is not being consistent; it is not measuring.
| Audit question | Cheap-to-fake answer | Expensive-to-fake answer |
|---|---|---|
| Last human override | "Humans are always in the loop." | A dated log entry, a named decision-maker, and the reason recorded before the outcome was known |
| Escalation after flattening | "We have an escalation process." | A named role, its span, and the hours per week protected for review |
| Training per redesigned role | "We invest heavily in upskilling." | Spend per role, year over year, alongside the headcount change in the same function |
| Outcome definition and measurement | "We are aligned on outcomes." | Separation of the defining party from the measuring party, in the contract |
| Published refusal surface | "The model has guardrails." | A public list of question classes the system declines, with the failure modes that produced it |
| Retired roadmap claims | "Our roadmap is on track." | A changelog of withdrawn claims with the evidence that withdrew them |
05
Why the audit is in the firm's own interest
The humane claim only earns a premium if it can be distinguished from its counterfeit
There is a market reason to want this audit, and it is not reputational. When a claim is free to assert, it stops carrying information, and the firms that actually did the work lose their ability to charge for it. Cheap humane language is a commons problem: every unearned assertion lowers the value of every earned one, until buyers discount the whole category and price on demo quality instead.
A firm that has genuinely rehoused escalation, funded role redesign, published a refusal surface, and retired a roadmap claim has assets that a fabulist cannot manufacture in a quarter. Its interest is in a screen sharp enough to separate it from the imitation. The audit is not a constraint on the honest operator; it is the only mechanism by which the honest operator gets paid the difference.
Which reverses the usual posture. The people best served by making "human-centred" falsifiable are not the critics. They are the firms that meant it.
- Licensed inference — Unverifiable quality claims compress price differentiation within a category; this is standard adverse-selection reasoning applied to AI procurement.
- Asserted — That the six-question audit is sharp enough to separate earned from asserted humane strategy in practice. Untested; it needs buyers to run it.
06
What this essay is not claiming
No named antagonists, no inevitability
This is not a claim that any specific firm, consultancy, or vendor is acting in bad faith. The documents cited are ordinary professional output and the diagnoses in them are largely correct. The argument is about an interaction between recommendations, and about the asymmetry that lets a correct sentence be borrowed by a programme it does not describe.
Nor is it a claim that flattening is always harmful, that agents cannot widen a span of control usefully, or that consolidation is the destined outcome of AI adoption. Nothing here is ordained. What is claimed is narrower and checkable: remove a layer without rehousing its refusal function and you have reduced the organisation's capacity to say no, whatever the strategy deck says about human judgement.
If the audit questions above are answered well by firms that flattened aggressively, the concern in this essay is misplaced and should be withdrawn.
07
The double bottom line, and the arithmetic it leaves out
Davos 2025: a correct instinct written in a currency nobody has to report
The World Economic Forum published the cleanest version of this argument in January 2025: leaders should adopt a "double bottom line" that weighs societal and individual empowerment alongside profit, because AI "could either uplift humanity to new heights or exacerbate old inequities." The piece is honest about the pressure it is written against — a majority of CEOs name AI a top priority and roughly seventy percent are already investing at scale, with profitability, efficiency, and performance the stated drivers. It cites the IMF figure that forty percent of global jobs are exposed to AI, and answers it with reskilling, public-private partnership, and regulation on bias, privacy, and transparency.
Every one of those recommendations is defensible. The structural problem is that only one of the two bottom lines has a reporting standard. Profit is audited quarterly, in a unit that everyone agrees on, by a party who is not the company. The second line — well-being, empowerment, community resilience — has no comparable unit, no cadence, and usually no external measurer. When one term in a two-term objective is measured continuously and the other is measured rhetorically, the optimiser does not split the difference. It follows the gradient it can actually see. That is the same failure the companion essays price in reward-function terms: an unmeasured objective is not a weak constraint, it is an absent one.
The forward-looking employment numbers make the point sharper. The essay reaches for a projection of 97 million new AI-related jobs — a figure whose horizon has now passed, which is exactly the kind of claim the sixth audit question exists to retire. The IMF exposure figure is a measured estimate of jobs at risk; the job-creation figure was a forecast. Netting one against the other, as strategy decks routinely do, is not analysis. It is an accounting move that lets a displacement cost be offset by an unaudited receivable.
None of this makes the double bottom line wrong. It makes it incomplete in a specific, fixable way: give the second line a unit, a cadence, and an outside measurer, or expect it to lose every time it competes with the first. The concrete instances the piece cites — a regulatory sandbox, an open-weights foundation, a skills partnership with a target and a date — are the parts of the argument that already have those properties, which is why they are the parts worth copying.
| Bottom line | Unit | Cadence | Who measures |
|---|---|---|---|
| Financial | Currency, standardised and comparable | Quarterly, mandatory | External auditor, with liability |
| Human / societal (as usually written) | None specified | Annual report narrative, discretionary | The company, in its own words |
| Human / societal (auditable version) | Training spend per redesigned role; override count; incidents caught pre-release; refusal surface published | Same cadence as financial reporting | A party that does not also set the target |
- Established — The IMF's exposure estimate (~40% of global employment exposed to AI, higher in advanced economies) and the CEO-investment surveys cited in the WEF piece are published, dated figures.
- Licensed inference — In a two-term objective where one term is externally audited and the other is self-reported, effort concentrates on the audited term. Standard measurement-distortion reasoning, not a claim of bad faith.
- Asserted — That giving the social term a unit, a cadence, and an independent measurer is sufficient to make it compete. Untested at firm scale.
- Speculative — Any specific net-jobs figure for the AI transition, in either direction. Forecasts of this class have not earned the confidence with which they are quoted.
What would show this wrong
- If firms that reduced management layers show equal or better rates of human override, escalation, and post-deployment incident catching relative to matched firms that did not, the load-bearing objection in Section 02 fails and should be withdrawn.
- If "human-centred AI" language turns out to correlate with measurable governance practice — override logs, published refusal surfaces, funded role redesign — then the alibi thesis is wrong and the phrase is already carrying information.
- If buyers running the six-question audit cannot distinguish outcomes between firms that pass and firms that fail it within two to three years, the audit is not sharp enough and should be replaced rather than defended.
- If contextualisation gains do in fact predict out-of-sample accuracy on external benchmarks, the separation drawn in Section 03 between relevance and truth collapses for enterprise systems.
- If outcome-based commercial models with a single party defining and measuring the outcome produce buyer results indistinguishable from independently measured ones, the conflict-of-interest objection is unfounded.
- If organisations that publish refusal surfaces show no reduction in sycophantic failure modes compared with those that do not, the fifth audit question is decorative and should be dropped.
- If firms reporting a double bottom line without an external measurer for the social term nonetheless match, on documented governance practice, firms that publish audited social metrics, the measurement objection in Section 07 fails.
Sources
- Starmind — Building a Human-Centered AI Strategy: Why People Make AI Work (Nov 13, 2025) — Vendor articulation of the four principles read in Section 03: contextualisation, collaboration, transparency, continuous learning. Useful precisely because it is candid about adoption failure being organisational.
- McKinsey & Company — Building leaders in the age of AI (Jan 12, 2026) — Source of the leadership moves read literally in Section 02: leaders modelling AI use, role redesign, and retained human accountability.
- McKinsey & Company — Are your people ready for AI at scale? (Mar 2, 2026) — Workforce-readiness framing; the training-spend question in Section 04 is aimed at the gap between this and implementation.
- Business Insider — McKinsey's new AI leadership playbook: flatten teams and move faster (Apr 3, 2026) — Trade reporting on the flattening recommendation and wider spans of control. The specific move examined in Section 02.
- World Economic Forum — Leading with purpose: why human-centric strategies are vital in the AI era (Harrison Lung, Jan 2025) — Source of the "double bottom line" framing, the CEO-investment and IMF exposure figures, and the reskilling/regulation recommendations read in Section 07.