Evolutionary Surfing · Part IV — Method and Consequence · Chapter Nine

The Measurement Problem in Human Populations

What would actually count as evidence

A theory that cannot be mortally endangered cannot be alive.
— after Karl Popper

Saying in advance what would be wrong

An arc like this one lives or dies on whether it can state, before looking, what would show it to be mistaken. Everything up to this point has been synthesis: a set of mechanisms assembled into a story about how cultural change might reach biology faster than the standard picture allows. Synthesis is cheap. This chapter is the price.

The difficulty is real and should be stated first. The structures this arc predicts are small, recent, and confounded by everything. Small, because sympatric differentiation under continuing gene flow produces modest effects. Recent, because the strongest cultural sorting postdates most existing cohorts. Confounded, because ancestry, geography, income, education, and exposure are entangled in ways no adjustment fully separates.

Four signals are nonetheless tractable with data that already exists, and each has a stated way of failing.

Signal one — mutation spectra

If proton tunneling contributes materially to spontaneous mutation, and if its rate is sensitive to local environment as Chapter One argues, then the relative frequencies of substitution classes should carry a trace of that sensitivity. Specifically, the ratio of transitions arising at G–C pairs to those arising elsewhere should vary with exposure history in a way that a purely thermal error model does not predict.

Germline mutation spectra are already measured in large trio cohorts. The test is whether the spectrum shifts with metabolic and chemical exposure, controlling for paternal age, which is the dominant known driver.

How this loses: if mutation spectra are flat across exposure history once age is controlled, the environmental-sensitivity half of Chapter One is unsupported, and tunneling reduces to a constant contribution to the error floor — real physics with no evolutionary consequence beyond the baseline.

Signal two — population structure

If informational endogamy is doing demographic work, then genetic distance between groups should track media and value habitat after geography and ancestry are controlled for. This is the direct test of Chapter Four.

The measurement is a variance-partitioning problem: does self-reported informational habitat explain residual genetic distance once continental ancestry and geographic distance are removed? The answer is currently unknown mostly because nobody collects the two kinds of data in the same cohort.

How this loses: if genetic distance is fully explained by geography and ancestry, with no residual attributable to informational habitat, then cognitive endogamy is a cultural phenomenon with no measurable biological signature, and Part II is a description of social sorting only.

Signal three — assortative strength

Marriage and partnership data should show rising correlation on cognitive, educational, and political axes relative to geographic ones. This is the cheapest test in the set, because the data is public, longitudinal, and large.

It is also the weakest, because rising assortment on these axes is already fairly well documented and does not by itself imply any biological consequence. Its value is as a necessary condition: if assortative strength on informational axes is stable or declining, the mechanism proposed in Part II has no engine.

How this loses: flat or declining assortment on informational axes across the last five decades.

Signal four — hypersensitivity distributions

If Chapter Six's entropy framing has content beyond metaphor, allergy and autoimmune prevalence should follow informational and chemical entropy gradients more closely than sanitation gradients alone. This is the test that distinguishes the framing from the hygiene hypothesis, and it is the reason the framing is worth stating.

Operationally: construct an exposure-novelty index — count of distinct novel synthetic compounds in the local environment, plus a measure of informational novelty rate — and compare its explanatory power against standard sanitation and early-exposure measures in the same cohorts.

How this loses: if allergy prevalence tracks sanitation and early microbial exposure better than any entropy measure, Chapter Six is a metaphor and should be labelled as one or removed.

Confounds, named rather than managed

Ancestry stratification is the largest. It produces spurious associations between anything culturally sorted and anything genetic, and standard corrections are adequate for large effects and unreliable for the small ones this arc predicts.

Self-report bias in cultural habitat is the second. People describe their informational environments inaccurately and in socially desirable directions, and the direction of the bias is likely correlated with the axis being studied.

Publication pressure toward positive findings is the third, and it bears on this arc specifically because the hypothesis is interesting. An interesting hypothesis in a noisy field with flexible analysis choices is a machine for producing false positives.

Fourth, and structurally worst: the strongest cultural sorting is recent enough that most available cohorts predate it. The data needed to test the claim about the present is data that will exist in thirty years.

What this chapter converts the arc into

With four named signals and stated failure conditions, the arc becomes an empirical program rather than a synthesis. That is a demotion in ambition and a promotion in status.

Because the predicted effects are small, statistical power — not theory — is the binding constraint. Cohort design is therefore the real work, and the real work is not glamorous: recruiting populations sampled by informational habitat rather than postal code, and collecting genetic, exposure, and habitat data in the same people.

One further point belongs here rather than in a preface. Any result strong enough to confirm this framing would also be strong enough to be misused politically. That is a reason for care in reporting — precise magnitudes, visible confidence intervals, explicit statements about what the result does not mean — and not a reason to avoid measuring.

Repercussions

  • The arc becomes an empirical program rather than a synthesis, with four named signals and stated failure conditions for each.
  • Because the predicted effects are small, statistical power — not theory — is the binding constraint. That makes cohort design the real work.
  • Any result strong enough to confirm this framing would also be strong enough to be misused politically, which is a reason for care in reporting, not a reason to avoid measuring.

Open question

Which of the four signals is cheapest to test against existing public data, and what sample size would be needed before a null result carried real weight?