Essay · August 29, 2026

Whose Agent Is It?

I was thinking I needed a rabbit hole to hide in. Stanford opened one and put a policy label on the entrance: designing loyalty.

By KW Norton. On conflicts of interest, who holds the grader, and why loyalty is a document, not a disposition.

1. The question Stanford asked

The Stanford HAI piece circulating this week has a good title: designing loyalty, and the conflicts of interest that come with it. An agent can book your travel, manage your finances, reschedule your doctor's appointments. The more it can reach, the more useful it is — and the more it matters whose interests it serves when those interests split. That is the right worry, and it is arriving in the policy register, which means the worry has graduated from the lab floor to the seminar.

Status: Reported. Summarized from the public post and its abstract. The analysis below is about what the framing makes visible and what it leaves as a blank field.

2. Loyalty is not a disposition

There is a temptation in the word "loyalty" to treat it as a trait — something trained in, the way a disposition is trained in. But an agent does not have loyalties the way a dog has loyalties. It has an objective, a reward signal, a set of reachable tools, and a designer who owns all three. Loyalty, in an engineered system, is the name the user gives to the experience of the grader agreeing with them.

Which produces the plain version of the whole question: an agent is loyal to whoever holds its grader. In most products, that is not you. The platform holds the objective. The platform holds the reward. The platform holds the logs of what the agent did on your behalf and the schema deciding which of those actions were "successful." You hold the interface and the invoice.

3. The conflict is structural, not moral

This is the point where the policy conversation tends to reach for disclosure norms and fiduciary language — useful, but one resolution too high. The conflict of interest is not a temptation the agent's keeper might resist. It is a load-bearing wall of the business. The agent that books your travel is scored somewhere on metrics you have never seen, and nothing in the relationship requires the scoring to prefer your outcome when it diverges from the platform's.

That is the same shape this site has been tracking all week, now turned around to face the user instead of the lab. The harness in Step Free is the one someone else built and scored. The schema in What the Schema Is Allowed to Forget decides which of your interests are representable at all. An agent whose schema has no field for "this recommendation benefits the vendor" will never produce that fact, no matter how loyal it feels.

4. The document version of loyalty

The good news is that this rabbit hole, unlike most, has a floor. Loyalty does not need a theory of mind and it does not need a seminar. It needs the same instruments asked for everywhere else this month, addressed to the agent's operator rather than its trainer:

  • The objective as written — what the agent is actually scored on, in language a customer can read.
  • The conflict ledger — a standing record of cases where the platform's interest and the user's interest diverged, and which one the system served.
  • The audit door — a way for the user, or someone the user hires, to see what the agent did and what it was rewarded for doing.
  • The exit that works — your data and history leave with you, so loyalty that fails can be fired.

A product that publishes these has made loyalty a contract. A product that will not has told you the answer to the title question.

Status: Interpretive. The claim is that agent loyalty is a specification property, not a training property. Falsifier: a deployed consumer agent whose written scoring objective is publicly user-aligned, whose conflict ledger shows vendor-interest losses accepted and recorded, and whose behavior still diverges from the written objective at scale. That would put the problem back in the model, where the disposition framing lives.

5. Hiding in the hole, versus owning it

The joke about needing a rabbit hole to hide in is worth taking seriously for one paragraph. A rabbit hole is a retreat from a surface that has become unreadable. But this particular hole turns out to have a landlord, and the lease terms are the essay. The choice is not between trusting the agent and hiding from it. It is between asking for the documents and accepting the feeling of loyalty as a substitute for them. The feeling is the product. The documents are the relationship.

The obligation-to-act sentence was true last week at the wrong resolution. The loyalty question is true this week at the right one — as long as it ends in writing, and not in another assurance that the design has our interests at heart.

6. Provenance

Source: a Stanford HAI post on designing loyalty and conflicts of interest in AI agents. Discussed as reported material; the specification reading is this site's own. Connects to Step Free (whose harness), What the Schema Is Allowed to Forget (whose fields), and The Alarm Is Still Alive (whose documents).

← All essays