In 1906, Francis Galton attended a livestock fair in Plymouth where visitors paid sixpence to guess the weight of an ox. Galton, a committed elitist, expected the exercise to demonstrate the ignorance of the average man. Instead, across the 787 valid entries, the crowd’s mean came within a pound of the animal’s 1,197-pound dressed weight: an accuracy no theory of expert judgment would have predicted from a six-penny raffle. His account, Vox Populi, ran in Nature the following spring.
The result is usually filed under “wisdom of crowds” and left there. But the interesting question was never that the crowd was accurate. The interesting question is why? and under what conditions the accuracy disappears.
The crowd at Plymouth worked because it disagreed:
The farmer estimated from years of hauling feed;
the butcher worked backward from carcass yields;
each entrant drew on a different contact with the animal world, and each brought independent errors.
Aggregate enough independent errors and they cancel; what remains is signal. The magic ingredient was never the size of the crowd, but in fact was the statistical independence of its mistakes.
This condition is fragile: though the failure mode is specific. Talking, by itself, can help a crowd when what circulates is information. What corrupts a crowd is the circulation of anchors: a posted number, a confident voice, a shared almanac that everyone consults before guessing. Then the errors correlate and a correlated crowd is confidently, collectively wrong. Diversity of information, not diversity of appearance, is what makes many minds smarter than one.
Hold that thought, because we are currently running the largest correlated-crowd experiment in history…
The Consensus Machine
A large language model is, at its core, a lossy compression of the statistical distribution of human text. Train on the corpus of everything written, learn the conditional probability of the next word, and what emerges is something remarkable: a fluent approximation of what the crowd of all authors would most plausibly say next. This is a genuine achievement, and I use these systems daily.
But notice what the architecture is. An LLM is Galton’s crowd… pre-computed, compressed, and frozen into weights. When you query it, you are asking the consensus a question. Repeated sampling deserves a fair hearing here: asking the model many times and taking the majority does improve performance on bounded reasoning problems, because different samples wander down different search paths.
What repeated sampling cannot do is manufacture independent evidence. The farmer, asked ten times, may check his arithmetic; he cannot acquire the butcher’s eyes. A blind spot shared by the whole training distribution survives every draw from it.
The industry’s other instinctive fix, ensemble several frontier models, inherits the same flaw for the same reason.
The major models share training corpora, share architectures, share the same human-feedback conventions. And the correlation is no longer conjecture: a study spanning more than 350 models found their errors substantially correlated… on one benchmark, when two models both erred, they gave the same wrong answer 60 percent of the time, and the largest, most accurate models were the most correlated of all, even across distinct providers. Plymouth’s condition fails at the foundation.
The crowd of models is a crowd that has read the same almanac.
And the condition is deteriorating. Model output is now flowing back into the training corpus at scale, and researchers have documented the failure mode this produces, models trained on their predecessors’ output lose the tails of the distribution first, the rare judgments before the common ones.
Collapse is a hazard rather than a destiny: it bites hardest where generated text replaces real data indiscriminately, and careful curation, accumulation of human data, and output filtering can slow it. But the economics lean toward the hazard. Synthetic text is nearly free; provenance is expensive; and every incentive in the pipeline points toward more of the former.
The crowd is collapsing, gradually and by default, into a single voice. The almanac is now being written by the crowd that reads it. I gestured at this loop over a year ago in Influencing the Shoggoth… the observation that everything we write becomes the priors of tomorrow’s machines; what I underestimated was how quickly the loop would begin consuming itself.
Emerging ‘Distillation Hazard’
There is a second mechanism accelerating the collapse, and it deserves a plain-language explanation because it will shape the economics of this industry for the next decade.
The practice is called distillation. A frontier model, the kind that costs billions of dollars and years of compute to train, can be queried millions of times, and its answers harvested. Train a much smaller model on those harvested answers, and the small model absorbs a startling fraction of the large one’s capability at a rounding-error fraction of the cost. The student never studies the world. The student studies the teacher’s answer key.
The dispute this has ignited is usually framed in legal and moral terms, and the irony there is genuine. The frontier labs trained on the accumulated writing of the internet: books, forums, journalism, code, without asking anyone, and called it fair use. Now the same maneuver is being run against them, and they call it theft.
Their defense rests mostly on terms-of-service contracts, which are notoriously difficult to enforce against a determined actor in another jurisdiction. And beneath the legal question sits a colder economic one: if 80% of a billion-dollar model’s value can be extracted for a few million dollars, the moat that justified the billion was never a moat. That thread: who captures the value of intelligence when intelligence can be photocopied, deserves its own essay, and I am deliberately parking it here.
Because from the vantage point of this one, distillation matters for a different reason entirely. The hazardous form, worth naming precisely, is black-box distillation: training on a teacher’s outputs alone, with no independent data and no external verification.
Done that way, a distilled model is a copy of a copy of the average. The frontier model already compressed the crowd into a single voice; the student compresses that voice again, and each pass of the photocopier tends to strip away more of the tails… the rare judgments, the minority readings, the residual variance that survived the first compression. (A student trained with its own data, or checked against the world, can escape this and can even beat its teacher on narrow tasks. The escape routes all run through the same door: fresh contact with reality.) What emerges from a distillation economy is a proliferating ecosystem of models that looks like diversity: hundreds of systems, dozens of vendors, a healthy competitive landscape by any conventional map. But trace the lineage and most of them descend from the same two or three teachers. Many bodies, one mind.
Galton’s crowd degrades in stages. First everyone reads the same almanac. Then, cheaper still, everyone reads somebody’s book report of the almanac. The distillation hazard is that the second stage arrives disguised as competition… the crowd appears to be growing precisely as its independent information approaches zero. Epistemic diversity is a property of lineage: of where the data came from, what objective was optimized, what contact with reality was maintained. A census of interfaces measures none of it.
Monoculture as Systemic Risk
Why should an investor care? Because the failure mode scales with adoption.
An adapted framework I have been developing borrows from condensed-matter physics, a phenomenon called Anderson localization, for which Philip Anderson shared the 1977 Nobel Prize. The physics, compressed: electrons move through a crystal as waves. In a perfectly regular crystal, those waves propagate freely. Introduce enough disorder: random impurities, irregularity in the lattice and something counterintuitive happens. The waves stop traveling. They become trapped in small pockets, and conduction ceases entirely. Disorder, which intuition says should merely slow things down, instead acts as a containment system.
Let’s map this onto economies… as analogy and hypothesis, offered in that spirit; economies are not quantum crystals, and heterogeneity can occasionally transmit risk as well as trap it and the implication still inverts a decade of systemic-risk thinking.
Model firms as lattice sites, supply relationships as the pathways between them, and the heterogeneity of firm behavior, differences in how firms substitute inputs, hold inventory, write contracts, as the disorder.
In a heterogeneous economy, shocks hit a firm and die out locally; every neighbor responds differently, and the shock finds no coherent path to propagate. In a homogeneous economy, where firms optimize identically, a local shock travels freely and aggregates into macro volatility. Homogeneity, on this view, is the systemic risk factor. Diversity of behavior, usually dismissed as inefficiency, is the immune system.
Now watch what happens as underwriting, credit, capital allocation, and policy analysis all converge on the same handful of foundation models and, one tier down, on the swarm of distilled descendants those models have spawned.
Every institution consulting the consensus machine is a lattice site behaving identically, and the apparent variety of vendors offers no protection: a bank running a distilled model and a hedge fund running its teacher are two sites with the same potential, however different the logos.
We are engineering the disorder out of the decision layer of the economy, polishing the crystal at precisely the moment we should be protecting it. This concern has now reached the institutions charged with watching for it: the Bank of England’s financial-stability committee has warned that widespread reliance on a small number of common AI models could push firms into correlated positions and amplify shocks. The errors will not be larger. They will be shared, arriving everywhere at once, correlated across every balance sheet that outsourced its judgment to the average. A monoculture of reasoning fails the way a monoculture crop fails: all at once, to the same pathogen.
The view from the edge
Internally we test opportunities for emrgnce against a framework we call the Proximity Inversion: the observation that in most institutional systems, the people closest to a problem hold the richest signal but the least capital and authority to act on it, while capital and authority sit furthest from the signal. Every layer between the edge and the center: reports, codes, standards, committee summaries, abstracts reality to make it legible, and destroys information in the process. By the time the signal reaches the people empowered to act on it, the part that mattered is gone.
Re-run the LLM through that lens and the diagnosis sharpens. The language model is the ultimate instrument of distance. It ingests the world only after the world has been flattened into text… described, summarized, published, averaged. Everything the vessel operator knows and never wrote down, everything the farmer senses before the almanac prints it, is invisible to the machine by construction. The consensus machine is legibility’s final form: an intelligence built entirely from the abstractions, with the proximate signal boiled away at every step. It does not merely sit far from the edge. It is made of distance.
Where variance returns
So the question worth asking about the next decade of machine intelligence: what would restore the Plymouth condition… independent errors, productive disagreement, a mechanism for settling it?
The answer, I suspect, lies in systems that model mechanisms rather than word frequencies.
A mechanism claim: this cause produces this effect, through this pathway, unless this condition intervenes, has a property that a statistical compression lacks: it is falsifiable at the level of the claim.
An honest caveat first: mechanism-based models can inherit each other’s mistakes too; several teams can encode the same wrong causal structure, and often do. Their advantage is narrower and more valuable than automatic independence. Two mechanism-based models that disagree do so legibly, thus we can locate the dispute at a specific causal claim and adjudicate it against data, or against the world. Two consensus machines that disagree do so opaquely; there is no ledger of why and hence nothing to settle.
Disagreement between mechanisms is information.
Disagreement between compressions is noise.
Which points at the deeper conclusion. Variance returns wherever intelligence is forced back into contact with independently observed reality… proprietary measurement, sensors at the edge, tacit expertise, local feedback loops that never passed through the published corpus. The mechanism is the vehicle; the contact is the cargo. An intelligence that must reconcile its causal claims against observations nobody else holds is, by construction, capable of independent error and independent error, Plymouth taught us, is the raw material of collective truth.
There is an old optical principle buried in this. Depth perception requires two eyes set apart, two structurally different vantage points on the same scene. A single averaged view, however sharp, is flat. Parallax is the reward for maintaining genuinely separate perspectives, and it is the only way any observer, biological or artificial, has ever perceived depth.
The economics follow. Every marginal dollar allocated by the consensus machine makes the consensus more crowded, and makes the independent estimator more valuable.
Correlation-driven intelligence competes away its own edge with each new adopter; mechanism-driven intelligence appreciates as the field homogenizes around the average. In markets, in underwriting, in any domain where being right matters more when others are wrong together, the returns to variance are rising exactly as fast as the supply of variance is falling.
Galton’s crowd was wise because nobody in it had read the same book.
We are now training every mind on the same book, the cheaper minds on a summary of it… and calling the result intelligence.
The ox, for the record, weighed 1,197 pounds; the mean of the 787 tickets came to within a pound of it. Somewhere in that pile, at least one entrant had written down the true weight exactly, one independent judgment, indistinguishable from the noise around it until the aggregation surfaced it.
The lesson was never that crowds are wise. The lesson is that independence is precious, rare, and at this particular moment in the history of machines… for sale at a discount.










Provides insight beyond common knowledge.