Essay · August 24, 2026

The Embedding and the Tree

Two reports arrived within the hour, from fields that do not read each other, making one claim: the tool that lets you see further also decides what you are able to doubt.

By KW Norton.

A picture that had to be flat

Single-cell expression data lives in thousands of dimensions. UMAP and t-SNE compress it to two so a person can look at it, and the compression is not neutral: the plane is Euclidean whether or not the data are, so local neighbourhoods and global distances distort in ways that depend on the parameters. Change perplexity or the neighbour count and the apparent biology changes with it. The picture is legible precisely because something was discarded, and nothing in the picture records what.

Bonsai, reported by de Groot and colleagues in Nature Biotechnology, declines the plane. Each cell is treated as a noisy point with an explicit uncertainty, and the method reconstructs the maximum-likelihood tree under a probabilistic model of how expression states change. A tree embeds in two dimensions for free, so the drawing follows a structure the data support rather than a coordinate system chosen in advance. On synthetic benchmarks it recovers geometries PCA and UMAP miss, it preserves cell-to-cell distances more faithfully, and on cord-blood data it reproduces known differentiation trajectories while surfacing an NK-cell subset placed with the myeloid rather than the lymphoid lineage.

The methodological move is not "better visualisation." It is a change in what the visualisation is allowed to assume. The old picture imposed a geometry; the new one fits one and states its uncertainty. That is the same discipline this project asks of its own figures, arriving from cell biology.

The generalisation the reporting draws is the one worth keeping: dimensionality reduction is not a neutral preprocessing step. When a method changes the representation of scientific data, it also changes which hypotheses become visually plausible — and visual plausibility is what most readers actually act on. So the standing requirement is symmetrical to the one this volume applies elsewhere: every time a lower-dimensional picture is accepted for the sake of seeing, name the geometry that was sacrificed, and check whether the hypotheses that now look obvious still hold in the original space. The lattice, the cortical gradients traced by geodesic distance, the moment body that consolidated the fluctuation-theorem bounds, and now the single-cell tree are four instances of one requirement, not four separate observations.

Source: Jorge Bravo Abad on X, on de Groot et al., Nature Biotechnology (2026).

The interpretability version

The interpretability report runs the other direction. As summarised, the CHIVE experiments asked whether activation-reading tools actually help explain odd model behaviour encountered in the wild — and agents given access to internal activations did not do better at predicting or explaining it than agents working from behaviour alone. If that holds, the uncomfortable reading is not that activations contain nothing. It is that a low-dimensional handle on a high-dimensional internal state is itself an embedding, with the same failure mode as UMAP: it produces something a person can hold, and holding it feels like understanding.

The reply that arrived with it — that understanding a machine intelligence requires some concept of what electronic awareness is, and an attempt to walk a mile in its shoes — inverts the object under examination. It is not the machine that requires the understanding. It is the human who dares to believe the comprehension has already happened. The machine is indifferent to whether it is understood; the person is not, and that asymmetry is where the error enters. Walking a mile in the shoes returns a report about the walker.

Read strictly, then, the interpretability tool and the perspective-taking move are competing embeddings of the same inaccessible object, and both are audits of us rather than of it. One flattens to features, the other to a human frame. Neither is free, and the second is far harder to check, because a felt understanding leaves no residual — no distance it failed to preserve, no branch it dropped. The question worth asking is not what the model is like inside. It is what a human being would have to be able to demonstrate before the claim to comprehend it counted as anything but self-description.

The mushroom, sonified

A clip circulating the same afternoon — a mushroom "producing music" from its bioelectric signals — is the friendliest possible version of the identical operation, and worth stating plainly because the friendliness is the risk. Fungal mycelium does carry measurable electrical potentials, with spike-like events and slow drifts, and those are real data. Everything musical about the recording is a mapping chosen by a person: which channel becomes pitch, which becomes velocity, what the time base is, what scale quantises the result, what gets discarded as noise. Change the mapping and you change the composition without touching the organism.

So the music tells a tale, but the tale is about the mapping. Sonification is a projection like any other: it converts an unfamiliar structure into a channel humans read fluently, and fluency arrives as a feeling of having heard the thing itself. That feeling is the same promotion UMAP performs visually — proxy to purpose, with no attached record of the cut. Sonification earns its place the moment it publishes its mapping and states what a listener cannot recover from the audio; without that, it is an instrument played by the researcher and credited to the fungus.

The general rule

Every upgrade in this week's ledger buys reach by discarding structure. UMAP for seeing. Multi-agent review for catching bugs. Geometric eigenmodes for cortical dynamics. A relational lattice for ontology. Each is a proxy that performs well inside a band, and each is quietly promoted to the purpose it was standing in for as soon as the band stops being stated.

Bonsai's answer and the evidence-cost rule from LLM Detente are the same answer in two dialects. Freeze the object. Charge a cost — distances and branching relations must survive the drawing; objections must be answered in a currency that is expensive to fake. Refuse the low-energy consensus, whether it takes the form of a second agent agreeing or a pretty embedding that looks like a result.

So the requirement is procedural, not aesthetic: the discarded structure stays named, priced, and periodically forced back into contact with the convenient picture. A representation that cannot say what it dropped is not a view of the system. It is a view of the method.

Status and falsifiers

Settled: t-SNE and UMAP do not preserve global distance structure and are parameter-sensitive; this is standard and documented. Reported, single paper: Bonsai's benchmark advantage and the NK-cell placement — a tree is itself a strong prior, and data whose true structure is not tree-like would be misdescribed by it just as confidently as UMAP misdescribes distance. Reported, second-hand: the CHIVE result reaches this page through a post summarising it, and a null result on one set of tasks does not establish that activation access is uninformative in general; it establishes that these agents, on these behaviours, gained nothing.

Falsifiers. For Bonsai: a dataset with known non-tree structure where the fitted tree reads as clean and confident would show the method trades UMAP's imposed plane for an imposed topology rather than removing the imposition. For the interpretability reading: a pre-registered task where activation access measurably improves prediction of out-of- distribution behaviour would retire the claim that these tools are embeddings first and explanations second. For the perspective-taking claim: it currently has no falsifier, which is exactly its standing.