Grasping Reality Anew
Artificial Intelligence and History’s Menu of Tools for Making Sense of Our World
Introduction
Artificial intelligence: no technology has ever been asked to carry so many fears — not the splitting of the atom, the internal combustion engine, or even social media, which up till now was the undisputed champion of manufactured anxiety. We’ve been told that AI will steal our jobs and reintroduce feudalism. That it has an insatiable appetite for electricity and water and spells environmental calamity. That its mushrooming harms — the chatbot that deepens a teenager’s psychotic spiral; the rising flood of deepfakes; the dissolving line between what a real person wrote and what a machine produced — are just a taste of worse things to come. That on a not-too-distant day AI may pursue a banal command like “manufacture some more paperclips” with such maniacal fervor that it tiles the entire globe in paperclips. That AI is Frankenstein’s monster and soon after it awakens to the sound of its own “I am alive!” it will find us excruciatingly boring, at best ignoring us and, at worst, turning us into mulch to ring its datacenters.
These fears summon a first-order question: what kind of thing is AI? Whether it can take your job is an economic question; whether it can want your job is a question about its nature. Serious thinkers, from Bostrom and Hinton to Harari, answer that AI is mindlike — an intelligence outright, bound for superintelligence, by some accounts on the cusp of consciousness, in time possessed of aims of its own.
I submit that this answer is a serious category error, and nearly everything said about AI, whether it’s expressed as a fear or a hope, inherits its original sin. AI is not a mind waking inside the machine, and not a parrot reciting the web. It is a world-modeling instrument of enormous reach and little self-knowledge: the mirror image of the most self-disciplined tool we have, statistics, and for exactly that reason a companion to it, not its replacement. Where that puts humans is back in the driver’s seat: with statistics a human ultimately decides what inputs to feed into a model, imposes a functional form, calibrates the confidence of the ensuing prediction, and decides whether to interpret a correlation as a causation. With AI, while the human role superficially changes, in that the machine now picks its own inputs and finds the functional form itself, learning its model from a found corpus rather than one a human specified—it is still the case that the calibration and the causal judgment remain the human’s: someone has to gauge how far to trust the output and decide whether its patterns are real or artifacts of a skewed sample, because the machine supplies neither.
Consider that a category error is a confusion of kind, not degree. A compass may point north, but it does not know where it is going. A thermometer may register a fever, but it does not feel heat. A camera may capture a face, but it does not recognize a friend. A map may represent a city with exquisite precision, but it does not inhabit the streets it depicts. In each case, the instrument extends a human capacity—orientation, measurement, perception, representation—without acquiring the inner life of the human being who uses it.
AI is an instrument for representing the world and predicting from it—kin to the map, the ledger, and the regression line—and to call it a mind is to file it under the wrong order of things: to read understanding, intention, and inner life into a tool whose fluent outputs only imitate them. It is to mistake the mirror for the face, the map for the journey, the model for the mind. We don’t ask whether the written word harbors private intentions, or whether a pocket calculator that beats every living human at long division is inching toward awareness. Why do we do so with AI?
A possible explanation lies with our penchant for reification. We say these systems learn, train, attend, reason, hallucinate, and think, and having lent the machine each human word we cash it back as though the machine owned the faculty the word names. Indeed, we began down this path as far back as 1943, when McCulloch and Pitts proposed the artificial “neuron” as an homage to the brain cell it only distantly resembled: the former was a stripped-down logical switch, the latter a living tangle of electrochemistry. We soon—and all too characteristically—forgot that this helpful analogy was ever an analogy.
This category mistake is much older than AI, however; it’s older than the computer and as old as the most pedigreed tool relied on to make sense of the world. The impulse to sense a mind behind a representation seems to switch on whenever a tool begins to deal in language, and it has switched on at every milestone on this menu. Recall Julian Jaynes’s much-disputed, but unforgettable, hypothesis: the earliest humans did not recognize their own inner speech as their own at all, and thus the verbal voice in the head was heard as a god’s command, and oracle, prophecy, and much of early religion followed—the first and grandest instance of mistaking one’s own representational faculty for an external agent.
Writing drew the same reflex as soon as it arrived. The Egyptians called their script the words of the gods; sacred books would later be venerated as living presences; and in the legend of the golem a heap of clay could be quickened into a servant by a holy word set on its brow and stilled again by erasing a single letter—matter made into an agent, and a dangerous one, by the power of the inscribed word.
Plato duly captured the confusion. In the Phaedrus, Socrates objects that written words “seem to talk to you as though they were intelligent,” yet when you question them they only repeat themselves, unable to answer, explain, or defend what they say—the precise complaint, twenty-four centuries early, that a fluent text can wear the look of a knowing mind while having no mind behind it at all. A fortiori, consider that no AI chatbot in the history of chatbots has ever sent an unsolicited message to a human, volunteering what’s on its mind in the hope of sparking a conversation. When it speaks first, some prior instruction, schedule, trigger, or human-designed workflow has summoned it. Its fluency can mimic a presence, but it does not disclose one.
Yet, as early as 1966, Joseph Weizenbaum’s ELIZA, a simple script that imitated a therapist by turning a user’s words back as questions, soon led people to confide in it, including his own secretary, who grew convinced it understood her and asked to be left alone with it. This phenomenon came to be known as the ELIZA effect: if a person experiences language that is fluent and responsive, they will impute a mind.
Therefore, one answer to the question—what is AI?—is that it is the most powerful trigger ever built for the oldest reflex we have. It draws the same projection the god-voice, the sacred text, and the toy therapist all drew before it. But AI perfects the conditions: it does not merely preserve language, like writing, or echo it back by rule, like ELIZA; it generates it fluently, responsively, and at scale—in our idiom, on our terms, and with just enough apparent memory, adaptation, and tact to make the old projection almost irresistible.
Perhaps the reflex makes evolutionary sense. Mistaking a rock for an agent wastes a moment’s vigilance; mistaking a living, intending “other” for a mere thing—missing the predator in the grass, the rival behind the gesture—can cost everything. Under that asymmetry selection favors a hair-trigger, tuned to over-attribute mind rather than under-attribute it: better a thousand false alarms than one missed agent. The upshot is that the same hypersensitive detector that for millions of years harmlessly overread the rustling grass and gave us animism now overreads the fluent paragraph and sounds the alarm it was built for: this alien is not only alive but means us harm.
In short, the category error that misattributes agency to AI is not a new conclusion, but an ancient habit we’ve applied to every tool we’ve concocted for making sense of reality, from the humble sentence said aloud to the internet’s all-knowing algorithm; AI has merely perfected the conditions.
In History’s Most Revolutionary Innovation, I show that America’s AI dominance was not an accident of entrepreneurial culture or free markets. It was engineered—through four decades of bipartisan reforms to intellectual property, antitrust, telecommunications, and trade policy that quietly built the legal and economic scaffolding the digital economy required.
Situating AI within the lineage of previous general purpose technologies like steam engines, electricity, and the microchip, and tracing its full arc from semiconductors to smartphones to large language models, I show how a handful of dominant firms simultaneously captured outsized returns and spread innovation across global supply chains—and ask what happens now that the US, China, and the EU are retreating into competing, gated technology regimes.
The result is the first comprehensive account of where AI came from, why its benefits have been uneven, and what will determine whether the AI revolution lifts living standards.
What I Aim to Do Here
This essay builds on that research to answer the question what is AI? I argue that it is a way of compressing some slice of the world into a manipulable representation that can help us make predictions. To arrive at that answer, I introduce a yardstick and use it to place AI alongside the other tools humans have built to grasp the world and predict from it: language, writing, quantification, statistics, computing, and the internet. I score each of them, AI included, on the same five dimensions: how much of the world it draws on, how skewed that intake is, how legible or opaque its workings are, how far its predictions reach, and which kinds of reality it can represent at all.
While I flesh all of this out further below, here is where AI lands on those five axes. On coverage, it is unmatched: through the internet it draws on more of the human record than any tool ever has. On reach, it is unmatched again—it will attempt almost any symbolic task and predict across almost any domain. And on scope, it climbs closer to lived, first-person texture than anything since the pre-linguistic mind, generating the immersive detail that counting and statistics deliberately threw away—even if, as I will argue, it never quite closes the last gap. While those are the three columns on which AI towers, on the other two it falls spectacularly.
On bias, the failure is not that its output is the most slanted ever produced—post-training scrubs much of that—but that it is the least able to say how slanted it is: it inherits the distortions of its internet-scale sample, adds its own, and strips away every instrument earlier tools built to catch distortion—no sampling frame, no honest error bar, no way even to ask how skewed its intake was. It opens the widest gap in the whole sequence between how far a tool reaches and how little it can vouch for.
On explicitness it is a black box, the steepest fall in the lineage: billions of parameters with no human-statable model to inspect, the exact inverse of the regression a skeptic can take apart term by term. And cutting across both is the discipline whose absence matters most—calibration: the tool with the longest reach we have ever built carries no native sense of when it should be trusted, and so it speaks with the same fluent confidence whether it is right or making stuff up.
Following Andrew Ng’s “new electricity”, Narayanan and Kapoor’s “normal technology”, Agrawal, Gans, and Goldfarb’s “prediction machines,” and Gopnik, Farrell, Shalizi, and Evans, who call the agent framing a “category mistake” outright and recast AI as a cultural technology in the line of writing, libraries, and search, ahead I put forward a deflationary, but cumulative, view of AI. While these thinkers each capture AI in a single illuminating description, my contribution is to set AI on one scorecard beside the older tools it is said to replace, score them all on the same five axes, and read off why they are complements rather than substitutes.
My take on AI is deflationary, agreeing with Bender and coauthors and Chiang that a Large Language Model is, at bottom, a machine for finding patterns and predicting from them—black-box statistics at colossal scale, the distant and far more powerful cousin of the regression that draws the best line through a cloud of points to guess a house’s price from its size. By prediction, what I mean is mapping inputs to a distribution over outcomes. For example, extrapolating the value of a home in dollar terms from its square footage and its school district’s average test scores and whether it has a view of the waterfront or not.
Bender and I part ways, though, over what that mechanism implies. For her and Koller, a system that only finds patterns in form has no purchase on meaning and so represents nothing of the world; I hold that the mechanism can be this humble without that verdict following.
My take on AI is cumulative in that the Large Language Model is nevertheless a veritable new entry in the lineage, and it does what the written word and the calculator could not—generating where they could only store, learning its own form where they had to be told, step by step, exactly what to do. It buys those gains the way every tool on the menu buys its own, by giving ground elsewhere: the fixed, checkable authorship that writing keeps, the legible, inspectable instructions the computer keeps. You know, human progress: one step forward, two steps back.
A word on what I am not claiming, since the temptation in an argument like this is to overreach. To deny that today’s AI is a mind climbing toward superintelligence is not to declare that no machine could ever think or feel; that is a metaphysical wager I neither need nor make. The three things people fuse—that AI is intelligent, that it is or could be conscious, that it has aims of its own—are separate claims, and I treat them separately.
This essay continues as follows. I begin by giving the superintelligence view its strongest run—the case that a mind is genuinely forming inside these systems—and then mark where that case overreaches and where it holds. I then step back to ask what a menu of representational tools is for, and what in our evolutionary inheritance made such tools possible, before setting out the yardstick itself: the five axes on which each tool is scored—coverage, bias, explicitness, reach, and scope—and the rule that decides which tools earn a place. With the instrument in hand, I walk the menu in order—from the pre-linguistic mind that anchors the low end, through language, writing, quantification, statistics, computing, and the internet, to AI—scoring each on the same five dimensions and marking what it buys and what it gives up to buy it. AI’s turn raises a question the others do not: whether its immersive, responsive output circles back toward the lived immediacy that the first act of abstraction cost us. It nears that limit, I argue, without crossing it. I close with the payoff the scorecard points to: a tool this lopsided is not a substitute for human judgment but a fresh reason to cultivate it—which is why a world awash in fluent, confident, and frequently wrong machine output calls for more education, not less.
Steelmanning the Maximalist View of AI
Before I lay the yardstick against AI, I owe the other side its strongest case. Every capacity once held up as the unmistakable mark of human intellect has, one after another, fallen to the machine: chess, then Go—its older and more sophisticated cousin—then the folding of proteins that had defeated biologists for half a century, then fluent language itself.
The latest iconic human skill to fall is the one that was supposed to be AI-proof—original mathematical reasoning. In 2025, frontier AI reasoning systems from Google and OpenAI reached gold-medal-level performance at the International Mathematical Olympiad, solving problems composed to be unguessable, and they began posting breakthrough scores on ARC-AGI, a benchmark its author built expressly to resist memorization and to measure the fluid, on-the-spot reasoning he took to be the core of human intelligence, even as its creators cautioned that those scores did not settle the question of general intelligence. In 2026, an OpenAI model produced a counterexample to a conjecture Erdős had posed eight decades earlier about the unit-distance problem—the maximum number of unit-apart pairs that n points in the plane can determine.
The startling feature was not only that the conjecture fell, but how. The model reached across the internal geography of mathematics, drawing on algebraic number theory to make progress on a problem in discrete geometry—the sort of cross-field move that mathematicians rightly prize as insight rather than calculation.
And the reasoning behind these achievements is not, on inspection, a conjuror’s trick. The newest models do not leap to an answer; they instead deliberate—producing long interior chains in which they try an approach, catch an error, back up, and try another. In at least one case this habit was never scripted in at all: it emerged on its own out of reinforcement learning—training by trial and reward, in which a model is set loose on problems, paid off only when it reaches the right answer, and left to discover for itself what gets it there. What it discovered was that deliberation works.
Researchers built instruments to watch the deliberation directly and discovered something uncanny, not just a souped-up version of “autocomplete.” Tracing the circuits of a working model, Anthropic’s interpretability team caught it planning ahead: asked for a line of rhyming verse, the model fixes on the word it means to end on before it writes the words that lead there—and when they reached in and swapped that target, the model rebuilt the line toward the new rhyme. They watched it reason in two hops, lighting up “Texas” on its way from “the state containing Dallas” to “Austin,” and found it thinking in a concept-space shared across languages—a kind of wordless language of thought beneath the English or the Chinese.
The same instruments have been turned on what the model knows, not only how it plans, and the readings are just as hard to wave off. Training small “probe” networks to read the activations of large ones, David Bau and his colleagues find structured representations laid out inside: asked to carry the Spanish gato into Portuguese, the model does not slide the word across a surface but drops the Spanish midway through its layers and forms a language-independent representation of the concept itself—cat, feline, the node an English or a Chinese prompt would also reach—before re-encoding it as Portuguese on the way out, translating through the meaning where a parrot could only repeat the sound.
Take the cleanest case of all. A network was trained on nothing but the move-lists of the board game Othello—bare strings of squares, one after another, the board never drawn for it and the rules never spelled out—and set only to predict the next move in the list, the purest form of autocomplete there is. To get good at that, it turned out, it had built the very thing no one gave it: probe its internals and you find a picture of the board, square by square, black and white and empty. And the picture is no ornament—reach in and flip a single square in its internal board, and the model’s next move shifts to match, exactly as a real player’s would. Rather than memorizing the stream of moves, it reconstructed the hidden game that produced them.
The same holds for models fed nothing but text, which lay down internal maps of real space and time—where cities sit, when events fall. Crack the thing open, in short, and you do not find a lookup table. You find a working model of parts of the world. All of this—the reasoning, the planning, the world-models, the glimmer of more—emerged from nothing grander than predicting the next word at scale.
So, yes: these systems do reason in recognizable ways. They do form internal representations that are not lookup tables. They do generalize from examples, plan, translate across languages, and recover structure no human explicitly placed inside them. Anyone who still describes them as mere autocomplete may want to revisit the evidence.
Furthermore, if a system can reason, plan, generalize, correct itself, model the world, and use language to navigate it, who am I to continue to insist that real understanding must be something further, hidden behind those capacities? Indeed, consider Ryle’s challenge. He was the first to attack the idea that mind is a hidden inner substance behind intelligent conduct: to understand, in his view, is not to possess some ghostly essence but to display the relevant capacities—to respond aptly, correct mistakes, use concepts, and find one’s way. And if all those capacities do not count as understanding, then understanding must be something else—some hidden inner possession behind the capacities, something the behavior expresses in us but merely mimics in the machine.
Ryle was debunking the idea that mind is a private essence behind intelligent conduct rather than the organized pattern of that conduct itself, and a powerful mainstream of the philosophy of mind has agreed with him ever since: a mental state is defined by what it does, by its role in the system, and a role can be filled in silicon as readily as in carbon. On that view a machine that reasons, plans, models the world, and converses does not imitate a mind—it instantiates one, and to deny it on the ground that it is built of the wrong material is a prejudice with a name.
And the maximalist’s largest claim goes further still: if experience rides on the organization of information and not on the meat that carries it, then machine sentience is an engineering question, not an impossibility; serious people, a founder of the field among them, treat it as a live one, and we have, after all, no agreed test for consciousness even in one another.
But, Still…
The question is whether this conclusion indeed follows from the premise. The maximalist wants the answer to be that, logically, what this means is that “a mind of some sort” is forming. I think the better answer is narrower and more exact: a new representational instrument has become extraordinarily good at building models of parts of the world and predicting from them. That is a real achievement and is indeed a core part of what our own minds do. And it is not a small achievement. But many of our tools do the same thing to different degrees, sometimes working together and often complementing our cognitive abilities, and it does not follow that those tools are as cognizant, alive, and intelligent as we are.
Indeed, the leap from model to mind smuggles in a single-axis account of intelligence. It takes one family of capacities—reach, fluency, abstraction, recombination, adaptation—and treats them as the whole. But intelligence worthy of the name is not merely the ability to range widely over the world and say plausible things about it. It also requires epistemic discipline—knowing how one knows: recognizing the limits of one’s evidence, distinguishing a representative sample from a biased one, exposing the path from premise to conclusion, calibrating confidence, separating correlation from cause, and revising belief in a way that leaves a durable trace.
On that count some of our more venerable tools for modeling a slice of the world and making predictions do better. Statistics, for one, can say how far its sample might mislead, wraps every estimate in a confidence interval that is honest, under its assumptions, about how far it might be off, lays its model open to challenge term by term, and—through the controlled experiment—earns the right to speak of cause rather than mere correlation. But a regression line is no more intelligent or conscious or alive than the actuarial table it produces.
And notice what this does to the reversal. The functionalist accused the deflationist of smuggling in a ghost—of demanding some hidden essence over and above every capacity on display. But the disciplines just named are not a hidden essence. They are capacities, every one of them, as behavioral and testable as the reasoning the maximalist celebrates: to gauge the skew in one’s own sample, to give an honest measure of uncertainty around a claim, to lay bare the path from premise to conclusion, to part correlation from cause. A system either does these things or doesn’t, and this one, measurably, does not.
So, I take the functionalist’s own rule—and put every other tool to the same test, treating each as a candidate mind and seeing where it rises to the challenge and where it falls short—and find the machine falling short on exactly the operations that separate knowing how one knows from sounding as though one does. Rather than subject AI to a benchmark such as MMLU—a single accuracy score averaged over some fifty-seven subjects’ worth of multiple-choice questions, from law and medicine to math—I employ a scorecard that goes axis-by-axis and that sets AI beside the other tools that model a slice of reality and predict.
These axes are: coverage, how much of the world the tool can draw on; bias discipline, its power to detect and correct its own skew, with calibration riding inside it; explicitness, how far its model can be read and challenged; reach, how widely and how far it can predict; and scope, the kinds of reality it admits—from pure abstraction to lived, first-person texture.
To preview the results, I find the ones that AI lacks, the disciplines of knowing how one knows—gauging the skew in its own sample, calibrating its confidence, laying its reasoning bare, parting correlation from cause—are precisely the ones that matter most for grasping the world.
Functions: What the Menu Is For
Before placing AI on the five-axes scorecard, it is worth asking why a species such as our own would build symbolic models of the world at all. A representation makes the absent present: it lets a mind grapple with what is distant, past, future, hidden, or lodged in someone else’s head. Two functions follow from this.
The first is prediction. Following Karl Friston and Andy Clark, the brain is already a prediction engine that models the world in order to anticipate it; external symbols—tallies, maps, equations—are that same impulse offloaded into the environment, where, as Clark and David Chalmers argue, cognition can be extended beyond the skull and made to handle far more than working memory allows.
The second function is coordination. A model that represents a slice of the world but remains trapped in one person’s head is cognition; a model that many heads can share, check, and pass down is culture. As Yuval Harari has stressed, shared representations—even fictional ones—are what allow large numbers of strangers to cooperate; and, as Merlin Donald and Michael Tomasello have each argued, this external, transmissible storage is what lets every generation begin where the last one left off, rather than start from scratch. Abstractly, this turns a population into a network whose value grows with its size: knowledge travels across it both sideways, from person to person, and forward, from one generation to the next, and the more nodes already joined, the more each newcomer can draw on and add to. Learning compounds.
Causes: What Makes It Possible
How is any of this possible? Ernst Mayr differentiates between the ultimate, evolutionary cause of a physical trait and its proximate workings. The raw capacities—the speech-ready vocal tract, the rough number sense that Stanislas Dehaene and Elizabeth Spelke find already present in infants and animals, the fine motor control for making marks—were assembled by natural selection over hundreds of thousands of years, blindly and slowly and often for ulterior purposes. The tools themselves were not.
Writing is barely five thousand years old, far too recent for the genome to have adapted to it, which is why, as Dehaene has shown, learning to read commandeers a patch of visual cortex that evolved to recognize objects and edges: there is no reading gene, only a vintage organ pressed into new service—and it is why the world’s scripts converge on the same handful of visual shapes.
Because the underlying architecture does not produce the tools, but instead bounds and enables them, they are diffused by a second, faster inheritance system—the cultural evolution, as described by Robert Boyd and Peter Richerson, that accumulates across generations on a timescale that genes could never match.
My Method
A word first about what this is and is not, because the method I employ is, I hope, a genuine contribution to making sense of what makes AI unique, scary, and promising. My focus is representation-as-inference. Every tool I will consider is a way of modeling some slice of reality to both detect regularities and predict the future. I place each of them on a single continuum, running from the unaided pattern-detection of a single brain through AI, with pitstops along the way that include language, statistics, and computers.
This is a synchronic anatomy. The question it asks is present-tense: what is on the menu of tools we have, circa 2026, for representing reality and predicting from it, and where does AI sit among them?
Therefore, this essay is a deliberate departure from the recent histories of information, most prominently Yuval Noah Harari’s Nexus, which runs the same cast of characters—speech, writing, print, the computer, AI—as a narrative of information networks and power. I am after something different and narrower: not the story of how we got here, but the structure of the toolkit we now hold.
To be clear about crediting others: I did not originate the familiar descriptions this essay leans on—AI as prediction, compression, cultural technology, normal technology—and have credited their authors above. My own contribution, as I said at the outset, is not another description but the scorecard itself: the single instrument that sets these tools side by side.
I score each tool on five dimensions. First, sample coverage: how much of the relevant world the tool lets us draw on. Second, sample bias: whether what we draw on is representative, and whether the tool gives us any way to detect and correct distortion. Third, explicitness of functional form: whether the model is legible—statable, inspectable, criticizable—or opaque. Fourth, reach of prediction: how far, and into what domains, the tool can predict, and whether it maps its own reliability. Fifth, representational scope: which kinds of reality the tool admits as data at all—a different question from coverage, which concerns instances within a kind.
A discriminating property runs underneath both the “bias” and “reach” axes: calibrated uncertainty, or the standard error and confidence interval that tells you how far to trust a prediction. This will be the hinge on which the comparison between statistics and AI turns.
“Scope,” the Cousin It of the scorecard, is the strange axis. The other four each measure a single virtue, and a tool can score high on any one of them independently of the rest: a deep neural network predicts well across many domains while remaining a black box, and a one-line equation explains itself completely while predicting almost nothing. Scope breaks this pattern. Its two halves are abstract admissibility—how abstract a category the tool can place an individual thing under—and lived fidelity—how much of that individual is preserved once it has been so placed. These two, unlike coverage, bias, explicitness, and reach, cannot both rise: moving to a more abstract category is the same act as dropping the particulars, so every gain in admissibility is a loss in fidelity.
For example, take one unique animal that exists somewhere in the world; that can be pinpointed by GPS to one exact location—a particular sow, inhabiting a particular pig pen; she is as pink as cotton candy, completely drenched in fresh mud, disfigured with a torn left ear, and oinks loudly. You can register it phenomenologically as exactly that: this pig, here, now. Or you can register it under a more abstract heading—“a pig,” then “a mammal,” then “an animal.” Each more abstract heading lets in more and yields more: from “mammal” alone come warm blood, hair, milk, a spine, true of thousands of creatures you will never meet. But each more abstract heading keeps less of the animal in front of you—“animal” licenses endless inference and says nothing of the torn ear.
As for the endpoints: the lower bound is the pre-linguistic mind. The latter qualifies as the benchmark against which to compare AI and the other tools that represent reality symbolically and try to predict what comes next. The upper bound is the AI language model, the present frontier, which does both tasks in peculiar ways.
Concerning the selection rule, a candidate earns a place on the menu if it passes two tests. First, a pragmatic gate: it must be a live tool, circa 2026, that could complement AI in the work of grasping reality. Second, a distinctiveness test: it must occupy a distinct region of the five-dimensional score-space rather than collapse into a neighbor—and I weigh a large movement on one load-bearing axis over small nudges on several.
This rule prunes the antiquarian detours (tally sticks, pre-literate calendars, the oral-mnemonic traditions), and it demotes distribution technologies that amplify a representational mode without being one (e.g., printing, which scales writing). It leaves a clean menu: language, writing, quantification, statistics, computing, the internet, and AI.
This approach entails sharp tradeoffs. The synchronic framing means I am selecting tools on their survival into the present, which narrows claims to the menu but precludes any claim about a trajectory. Nothing in the sequence is a clean improvement; every tool trades something off somewhere. And the scope axis is the softest of the five, because one cannot fully enumerate what a representation excludes from inside that representation.
Foreshadowing the Results
Since this is not a mystery novel, let me say at the outset what the scoring will show. The general trade-off—that no tool is a clean improvement on the one before—takes a definite shape. The tools that enlarge our reach—spoken language, writing, computing, the internet, and now AI—each widen the aperture while loosening the controls. The tools that impose discipline—counting, measurement, and above all statistics—then claw back rigor, but on the larger base the scaling tools have pried open. AI is no exception and is not particularly remarkable in the trade-offs it introduces.
Nonetheless, its profile is lopsided in a particular and revealing way: it takes in more of the world, and predicts across more of it, than anything humans have built—yet it cannot tell you how skewed its intake is, cannot show the workings behind a given answer, and cannot put honest error bars on what it produces, which are exactly the disciplines that statistics, the most self-aware tool on the menu, was built to supply. AI is, to a significant extent, the mirror image of classical statistics: strong precisely where statistics is weak and weak precisely where it is strong.
That symmetry is the payoff. The two are natural partners—AI for reach and coverage, statistics for bias-correction, transparency, and calibrated doubt. And to foreshadow where I am going with all this: a tool this lopsided does not retire human judgment so much as raise the premium on it.
In another respect, AI seems to circle back toward the very thing the first abstraction cost us: immediate, lived experience. A regression can tell you, to the decimal, how much more a house sells for because it overlooks the water; it cannot tell you what that water looks like at six in the evening, or how a kitchen smells the morning after a death in the family. Ask a language model, and it will hand you a passage that reads as though it knows—because it has absorbed the testimony of thousands of people who did. Since it is trained not on tidy columns of numbers but on the vast, unruly record of how human beings have actually rendered their experience—every description, story, confession, photograph, and song we have committed to the page or the screen—it can traffic in the qualitative, sensory, particular texture that counting and statistics deliberately threw away.
But an immediacy that has been encoded is still encoded, and the first-person, pre-symbolic presence the mind enjoyed before language is the one thing no representation, however vast, can ever give back.
That limit carries us back to the maximalist. What she took for a mind roused awake, the scorecard will put in its proper place: the newest tool on a very old menu of tools that make the absent present. These tools are extensions of ourselves, the scoring will show—but not replicas of us, and not even proxies. AI in particular will come out wider in reach than anything before it, blind to its own bias, and, even at its most lifelike, only ever modeling experience rather than having it. So, with that preview out of the way, let me walk the menu in order, beginning where representation itself begins—with the pre-linguistic mind.
The Menu, Tool by Tool
The Benchmark: The Pre-Linguistic Mind
Before language, there is already a model of the world—not the world itself, but a stripped-down stand-in for it that predicts. The humanlike ancestor who paired the rustle with the predator, or the berry with the illness, was running an implicit forecast: a regularity pulled from past cases and projected onto the next one. The representation is crude and the prediction is wordless, but both functions—holding a model, predicting from it—are in place before a single word is spoken.
Four independent literatures show how much these minds already do, and each pins down a different property of representation without language. Comparative cognition shows it needs no language at all: great apes and corvids plan, use tools, and reason about cause with no symbolic speech, so representation and inference are language’s inheritance, not its invention. Developmental psychology shows it precedes language in our own species: the “core knowledge” of Spelke and Carey—intuitive physics, agency detection, an approximate number sense—lets pre-verbal infants form expectations and look longer when a ball rolls through a solid wall, and that surprise is simply a prediction caught failing. Paleoanthropology shows it can be stored and copied rather than merely held: the Oldowan and Acheulean traditions carried high-fidelity procedural knowledge—which strikes and shapes work—across hundreds of thousands of years with no words to transmit it. And animal communication shows the content can leave the single head: vervet alarm calls sort the world into predator kinds, and the honeybee’s waggle dance encodes a food source’s direction and distance—partial representations, externalized and acted on, before language exists to do the externalizing.
What follows refines, or trades against, what this mind can already do. On the five-axes scorecard, it sits near the floor on four axes, with scope the notable exception.
Coverage
Coverage is tiny—one body, one lifetime, line of sight, the present and recent past, extended only by what one can watch others do.
Bias
Bias is maximal and nearly uncorrectable. The senses that deliver reality to us pre-linguistically were honed by natural selection to register only some slices of it and ignore the rest—no sonar like the bat’s, no olfactory world like the dog’s, no infrared sense like the pit viper’s—so even raw perception is an inference cloaked in the illusion of pure apprehension: a selective model anchored to one body at one point in space and time, tuned to what mattered for survival rather than to what is actually there. Moreover, the sample is survivorship-censored in the most literal way because the individual who made the fatal error is removed before the lesson can travel, as you cannot learn that the berry killed your neighbor if he cannot tell you and you did not watch him die.
Explicitness
Explicitness is zero despite rich implicit models: the intuitive physics is there, but it is tacit, unstatable, unshareable, unrevisable.
Reach
Reach is short and concrete—bound to the perceptible and the near future.
Scope
The exception is scope, and it forces a distinction we will need throughout. Lived fidelity begins at its high-water mark and erodes with every encoding that follows; abstract admissibility—in the explicit, shareable sense—begins near zero and climbs. In one sense the pre-linguistic mind is at the maximum: raw experience is multimodal, embodied, affective, first-personal, admitting the full lived richness that no later representation—not language, not number, not the machine—ever fully recovers. It is not encoding and losing; it is living. In another sense it is at the minimum: none of that can be stored, abstracted, transmitted, or recombined.
Merlin Donald describes the transition toward language as a stage of mimetic culture, centered on imitation and gesture. While that is interesting in its own right, we will skip that bridge and focus exclusively on verbal speech next.
Language
Symbolic language is the first and, in some ways, the largest leap in humans’ ability to model a slice of reality and make predictions, because it is the first representation that escapes the single perceiver. The pre-linguistic model was analog, welded to the present, and sealed inside a single skull; language is none of these. Its unit is the arbitrary sign—a word that, unlike a footprint or a cry, bears no resemblance to what it names—and because the bond between sign and referent is fixed by convention rather than given in perception, the sign can travel where perception cannot: it can be uttered when the thing is absent, copied from mouth to mouth, and carried across a lifetime to someone who never witnessed the original. Language thus takes the private model evolution wired into perception and externalizes it into a shared, transmissible medium. What one mind has registered, another can now receive without registering it firsthand: perception becomes testimony.
Linguistics has a standard way to pin down what separates human language from a vervet’s alarm call or a bee’s dance: the design features catalogued by Charles Hockett. His list is long, but two of the items do most of this work. Displacement is the power to refer to what is not present—the absent, the past, the future, the hypothetical—so that one can speak of the lion that has already gone, the grandparent long dead, the hunt planned for tomorrow, or the river that does not exist at all. Productivity—the recombination of a finite lexicon by rules of syntax—yields an unbounded set of meanings, most probably never expressed before in the same, exact, way, but understood as soon as they are heard. Together, they convert a fixed repertoire of signals, each bound to its stimulus, into an open-ended system: an engine for saying, in principle, anything. And because the medium is shared, the model it carries is no longer the property of one mind—it can be voiced, contested, and handed down, the leap from private cognition to public, cumulative culture.
Coverage
Coverage makes its signature jump: language pools observation across persons and across generations, the single biggest expansion of effective sample size across all tools, bounded only by the oral horizon—the size of the speech community and the depth of reliable memory before drift sets in.
Bias
Bias moves in two directions at once. It slashes the benchmark’s most lethal censoring: with displacement, the dead and the absent can warn the living. But it introduces biases the benchmark could not have—testimony bias (one now samples what others choose to report), transmission drift, and the consequential new one, fabrication. Language can represent what never happened; myth, rumor, and lie float free of any observation, and there is yet no apparatus for testing competing claims.
Explicitness
Explicitness takes a deep step: naming a regularity makes it a statable proposition that can be shared, disputed, and revised—the birth of explicit modeling, and with it the first possibility of correcting the model rather than only the data. It is explicit but imprecise, though; precision awaits the codification of mathematical abstraction in the form of counting and calculation, both discrete and eventually continuous—that is, after the invention of calculus.
Reach
Reach extends, via displacement, to the absent, the distant, the future, and the counterfactual. One can now warn of the wolf that has slipped back into the trees, describe a coastline no one in the band has ever walked, foretell where the herd will graze when the snows come, and weigh what would have followed had they taken the high pass instead of the valley floor.
Scope
On scope—which kinds of reality a representation can hold at all, not how far its predictions reach—language pulls the axis’s two halves in opposite directions. Abstract admissibility rockets up from zero: whole kinds of content that perception could grasp only as concrete particulars now become representable—the categorical, the normative, the narrative. Lived fidelity, by contrast, takes its first hit, and only a mild one. “The sunset was beautiful” compresses an overwhelming experience into a thin token; but spoken language still carries breath, prosody, and the body of the speaker, so the loss at this first step is slight.
So, language is ultimately a lopsided trade that improves nearly everything, with one mild loss and one genuine new liability. And here a motif appears that will recur at every milestone: a tool’s greatest strength and its worst liability tend to be the same feature. Language’s capacity to represent the unobserved—its power to coordinate millions through shared fictions, as Harari emphasizes—is identically its capacity to fabricate. You cannot have the coordinating myth without the capacity to lie. The civilizational scaffolding and its susceptibility to bias arrive together.
Writing
Writing is language’s nearest neighbor, extending the same gain onto a new axis: from pooling across people to pooling across time, without decay.
Coverage
Coverage leaps from living memory to the documentary horizon: writing makes the first large, durable datasets possible—tax rolls, census returns, temple and trade ledgers—so a sample can accumulate across centuries instead of dying with each generation.
Bias
Bias moves in two directions again. In one, it corrects: the fixed record defeats the transmission drift that eroded oral memory and lets a claim be checked against something stable—the first real verification, the first payment on the debt that language opened. In the other, it distorts, in two ways. The sample is now far larger but skewed: writing records only the literate and the record-worthy, so the surviving archive is the elite’s. And the very fixity that made the check possible now entrenches what it carries: writing transmits error as faithfully as truth and lends it authority, so a recorded falsehood—Galenic medicine, Ptolemaic astronomy, Aristotelian physics—gains the standing of “it is written” and can outlive its correction by centuries. The feature that kills drift entrenches dogma.
Explicitness
Explicitness advances in durability rather than kind: a frozen argument can be dissected at leisure, by many readers, across time. Nothing in the medium demands that reasoning be shown—it serves the comic book as faithfully as the proof. What durability adds is the precondition: only an argument held still can be audited step by step and checked by strangers across generations, and that is what lets the long, layered, cumulative structures stand—Euclid’s geometry, the legal codes, systematic theology. While the request to show one’s reasoning is the discipline’s, writing makes this demand satisfiable.
Reach
Reach is where that persistence becomes foresight: a forecast can reach as far forward as the record reaches back. Oral tradition can outlast any speaker—epics and genealogies have crossed many generations in memorized verse—but what memory keeps faithfully is what is memorable: story, rhythm, a rule simple enough to recite. A long run of arbitrary, exact, dated observations has none of those hooks, so it drifts; and even intact, speech offers no way to lay many observers’ records side by side and compare their intervals—the analysis a long-period regularity demands. So oral forecasting stays short and qualitative: red-sky-at-night predicts fair weather tomorrow, a simple rule that holds across any season.
Writing supplies what memory cannot, and a regularity like the eighteen-year Saros cycle becomes discernible: the Babylonians read it out of their astronomical diaries and could thereafter say when an eclipse would fall, not merely that one might; and Halley, comparing written records of the comets of 1531, 1607, and 1682, recognized a single returning body and predicted its reappearance in 1758, sixteen years after his own death. What writing adds, then, is twofold: a longer horizon, where speech could reach no further than tomorrow, and a higher ceiling on complexity, carrying a regularity too intricate for any memory, individual or collective.
Scope
On scope, admissibility takes in a new kind: the past itself becomes a fixed object—a chronicle that pins events to dated years, a precedent cited verbatim and still binding centuries later—where memory alone could never hold the past still. Lived fidelity, meanwhile, descends another stair: while writing keeps the words, it loses the speaker—the tone that separates irony from sincerity—and, more consequentially, the power to answer a question. A statute cannot rephrase itself for the reader who misreads it; a letter cannot see the puzzled face and try again.
Quantification
Quantification bundles two significant advances. A number makes our model of reality more exact, transforming a relationship between a part and a whole into a quantity to calculate rather than a word to interpret. Standardizing measurements into like units makes observations comparable across observers, aiding both the ability to grasp reality and predict what comes next, as in a tide table, where measurements taken in the same units along a coast reveal the tides’ rhythm and foretell when the next high water will come.
Coverage
Compared to writing, the gain in coverage is moderate, and only for what can be counted. A standard unit turns observer-relative reports — a handful, a stone’s throw, a day’s walk, each a different size in a different hand — into measurements that mean the same everywhere, so one traveler’s distances can be pooled with another’s into a single map, or many farmers’ yields into one table. And counting turns a collection into a number: a herd, an army, a harvest — which one could otherwise only call “large” — becomes a definite figure that can be added to last year’s, averaged, and compared.
Bias
Bias is the first milestone whose main effect is corrective: it makes a model’s data trustworthy rather than merely more plentiful. Two operations do this, both absent from every earlier tool. Exactness removes the wiggle room that lets a figure be shaded: “a fair-sized cargo” can be quietly padded and “he owes me a good sum” haggled down, but “37 amphorae” and “340 denarii” are pinned — there is nothing to stretch. And arithmetic lets the data expose its own errors: if the strongbox opened with 200 denarii, took in 90, and paid out 50, it should hold 240 — and when it holds 230, the missing 10 reveal a theft or miscount that “a brisk day’s trade” could never have caught, and that a ledger, faithfully copying a false entry, would only have preserved.
The same exactness can deceive, though: a falsely precise figure — GDP to the decimal, a credit score of exactly 686 — lends a model an authority its data never earned, so it forecasts to the decimal what it cannot really know, confident and wrong. That is Porter’s mechanical objectivity, the idea that this tool’s exactness hides the judgment it appears to remove; how far such a forecast deserves trust is the one thing number cannot say — the reason statistics must come next.
Explicitness
Explicitness makes its deepest jump since language first brought it into being. A claim you could state but not calculate becomes a precise, calculable relation — a rate, a ratio, a law like distance = speed × time — and, just as important, one you can operate on. Operating on it exposes a structure that intuition and words can’t reach.
What sets mathematics apart from the senses and from ordinary language is not that it perceives more, but that it can compute. The senses deliver only the surface — the magnitude in front of us, the shape we take in at a glance — and language can name what lies beneath that surface but cannot work it: you can say a quantity “compounds” and will “explode,” yet the words sit inert, giving you no doubling time and no value for year fifty.
An equation, however, is different in kind. Because it is exact, you can rearrange it into an equivalent form — one that shows what the original hid.
For example, Eratosthenes worked out the size of the Earth without leaving Egypt. He knew that at noon on the longest day of the year, the sun was straight overhead in the city of Syene: a vertical stick there cast no shadow. On that same day, in Alexandria to the north, a vertical stick did cast a shadow, and from its length he found that the sun was about 7.2° off from straight overhead. The sun is far enough away that its rays reach both cities as parallel lines, and that means the 7.2° is also the angle between the two cities as seen from the center of the Earth.
That gives a simple relationship: the angle between the cities is to a full circle as the distance between the cities is to the whole way around the Earth:
θ / 360° = d / C
Each part is something concrete. θ is the angle he measured from the shadow (7.2°); 360° is a full circle; d is the distance from Syene to Alexandria, which was already known. C is the thing he was after — the distance all the way around the Earth, which no one could measure directly.
Getting C by itself takes one step:
C = d × (360° / θ)
Now plug in the numbers: 7.2° goes into 360° exactly 50 times, so the two cities are 1/50 of the way around the Earth, and the whole Earth is 50 times the distance between them. That distance was about 5,000 stadia, so the Earth is about 5,000 × 50 = 250,000 stadia around — close to the value we measure today. From a stick, a shadow, and the distance between two cities, he found the size of the planet. The order was always in the world; math is the tool that rewrites it into a form the mind can read and the hand can compute.
Return to compounding: a steady two percent a year — no one’s idea of fast — doubles a quantity in about thirty-six years. Exponential growth is repeated multiplication, by 1.02 again and again, which on an ordinary graph is a curve climbing ever more steeply — its mild early years no warning of the rise to come, no fixed rate to point to. What makes it legible is the logarithm: it turns multiplication into addition, so multiplying by 1.02 each year becomes adding a constant each year, and a constant step repeated is a straight line whose single slope is the rate. The curve the eye couldn’t read becomes a line it can — extend it to any future year, read the doubling time straight off it. The runaway intuition never sees coming is still there in that slope — a constant climb means the quantity multiplies without limit — but now it is something to measure and project rather than be ambushed by.
Compound growth is only the nearest example. Math does the same for change: calculus relates a quantity to its rate of change — position to speed, speed to acceleration — and from the rate alone reconstructs the whole motion. Ask a driver whether doubling his speed doubles his braking distance and he will say yes; but under steady braking the distance grows with the square of his speed, so twice as fast is four times the room needed to stop.
Fractal geometry does it for roughness. A coastline has no single length — measure it with a finer ruler and it only grows. Euclid’s smooth shapes can’t describe a coast, a cloud, or a lung; fractal forms can.
Each time, the surface either misleads or says nothing, and the math reaches the structure underneath — and a structure made legible is one you can finally compute with and predict from.
Reach
Reach gains on two fronts with quantification: precision, and entry into the never-observed. Where a written record could say an eclipse was due, number says it will be total, beginning at 11:14, along a hundred-mile band of the Earth’s surface — the prediction is pinned to a magnitude, a moment, and a place. And because a law can be computed, reach extends to cases no one has watched: from the formula for a falling body, you can say where a cannon fired at an untried angle will land without firing it.
Perhaps more importantly, the reach is only as wide as the law is true: pushed past the conditions it was built for, the same formula predicts a miss as confidently as a hit. Fire it hard and far, and the cannonball must cross air the formula leaves out: drag bleeds its speed, the clean parabola folds into a foreshortened, lopsided fall, and the shot lands well short of the spot the equation named with such confidence. “Frictionless” geometry fails in the face of real gunnery — a fast cannonball drops far shy of where Galileo’s parabola says it should land — and only a ballistics-informed equation, one that puts the drag back in, tells us where it would truly land.
Scope
On scope, the trade is the sharpest yet — but it is a trade, not a conquest. Admissibility gains a powerful new kind: pure quantity, the number that stands on its own and can be operated on. What it cannot take in is whatever is not a quantity — what a thing is like, what it means, its particular singularity. A soil report gives nitrogen in parts per million but does not speak to the farmer’s feel for when his own field is ready to cultivate; and the point is not that the number erases that feel or makes it worthless — the two sit side by side, and where the feel tracks a real pattern the number can sharpen it. The point is narrower and harder to escape: a number leaves the qualitative unrepresented, so to mistake the figure for the whole is to lose what it never held — the look and feel of the field, what it has come to mean to the man who works it, what is particular to this acre and no other.
Statistics
Where quantification provides an exact number but does not say how far to trust it, statistics supplies precisely that: it gives not just a point estimate — a precise prediction in numerical terms — but also a confidence interval around that estimate, a band produced by a procedure that, run on sample after sample, brackets the true value a stated share of the time — ninety-five times in a hundred, for example.
Probability is the standing guard against mistaking a fluke for some larger purpose, one you might mistakenly believe was predetermined, preventing your experience of an improbable event from hardening into superstition. The model it lays over a slice of reality frees you from taking your own experience as the whole: your observations are draws from that larger lawful pattern, and the distribution locates your experience within it, so you can tell how representative it is.
Rather than “there are no coincidences,” you say to yourself, “improbable things happen all the time,” because you posit a distribution — a law of chance, the even odds of a fair coin, the bell curve into which measurement errors fall — and deduce the full spread of outcomes it produces, which are common, which are rare, and exactly how rare, before a single observation is in hand. Seven heads in a row is unlikely — about one chance in 128 — yet entirely ordinary for a fair coin, so you can take it for a run of luck rather than a loaded coin or a sign that some force beyond chance is at work.
Where probability guards against mistaking a fluke for a portent, statistical inference guards against mistaking too small a sample for the truth, and it says, in numbers, exactly how much of your conclusion the data has earned. It does this by running that same map — the one probability provides — backward, from world to model: now the distribution is the unknown, and the observations you collected are the only clue to discerning it.
Take the seven-heads case from the other side: hand someone a coin of unknown bias, let them flip it ten times, and seven heads come up. The tempting read is that the coin favors heads, but inference holds the estimate at arm’s length — ten flips are far too few to rule out an ordinary fair coin, and the interval around that seventy percent runs so wide it still comfortably contains one-half. Flip it a thousand times and get seven hundred heads, and the interval tightens to a narrow band that no longer includes a fair coin: only now are you entitled to the conclusion.
The profound implication is that you never observe the underlying reality directly, only the draws it sends you — so certainty gives way to graded confidence, rising as the observations accumulate, and never quite reaching the absolute. In this way, statistics can adjudicate between claims about the world based on its underlying structure, disciplining both the claims and the predictions about what to expect.
Finally, there’s the controlled experiment, which does not take its data as found but generates it. By assigning a treatment at random, it makes two groups alike in every respect but the one under test, so the untreated group can stand in for the counterfactual no one can observe directly — what the treated group would have done had it gone untreated — and the gap between them reads as the treatment’s effect rather than an accident of who ended up where. The disciplining revolution is now complete, in that for the first time we can move from co-occurrence to cause by design rather than by inference, but still within the spirit of epistemic humility: the workhorse confidence interval returns to tell us it brackets the true value of the effect only a stated share of the time.
Coverage
Coverage makes a leap quantification could not: from counting what you can reach to inferring the whole from a part. A properly drawn sample of a few thousand stands in for a nation of millions, and sampling theory states exactly how far the part may stray from the whole — so the census gives way to the poll, total enumeration to representative inference. You no longer have to observe everything to know it; you must observe the right slice and know how thin it is.
Bias
Bias reaches its high-water mark: for the first time a tool disciplines the sample itself, not merely the single figure — random sampling makes the sample representative in expectation, and inference, as above, states how badly it might still mislead. But the same machinery cuts the other way, in the lineage’s recurring fashion: the power to pull a real signal from noise is identically the power to manufacture one — test enough hypotheses, slice the data enough ways, and an artifact of the search emerges wearing every mark of a discovery. The tool that finds the true effect is the tool that also finds the false one.
Explicitness
Explicitness reaches its high-water mark too. A statistical model is not just computed but written down: an equation whose every term carries a meaning, each coefficient a named, estimated weight — a yield rising so many bushels per inch of rain, so many per pound of fertilizer — and each weight reported with the uncertainty around it. The model fits on a page, and a skeptic can take it apart term by term: test whether the rain coefficient is real, ask whether the fertilizer effect survives once soil quality enters the equation, reject the form and propose another.
Nothing the model claims is hidden, because the model is its claims. This is the most legible a predictive instrument ever becomes — worth marking precisely because it is the property the lineage is about to surrender: the last tool at which predicting well still demands a model one can read, state, and contest.
Reach
Reach extends into territory number could not enter — the inherently random and the aggregate. Where a deterministic law fixed the eclipse to the minute, statistics forecasts the chancy: an election within a margin, a drug’s effect within an interval, a year’s claims across a pool of policyholders not one of whom can be predicted alone — and every such forecast arrives with its own error bar, the calibrated trust that quantification could promise but never supply. Its furthest reach — the move from co-occurrence to cause — is the one the experiment above already made; here it is enough to mark that prediction now spans not only what will accompany what, but what will follow from what, if we act.
Scope
On scope, the new kind statistics admits is the lawful aggregate — the distribution, the rate, the regularity that exists only across many instances and in none of them singly. Earlier tools could hold the individual case; what they could not hold was the law governing a thousand cases at once, since each treated variation as mere error around a true value, noise to be scrubbed away. Statistics is the first to take that variation as the object itself: chance made into a structured thing, lawful in the aggregate even where lawless in the instance, available to be modeled and operated on in its own right. And lived fidelity descends another stair by the very same stroke: to render the world as a distribution is to dissolve the individual into a draw from it. The actuarial table knows the death rate of the cohort to the decimal and the man inside it not at all.
Computing
A computer is a machine that stores information and transforms it by following explicit instructions. To do so, it harnesses electricity by channeling it into two discrete states — on and off, one and zero — and it builds everything else on top of that. Switches wired into logic gates turn the two states into the elementary operations of arithmetic and logic — an and-gate emits a one only when both its inputs are ones, an or-gate when either input is, a not-gate flips a one to a zero and back. Wire them together in the right arrangement and the circuit, fed the binary numerals 101 and 011, returns 1000: five plus three is eight. Numbers, letters, every kind of data come written in this single code of ones and zeros, so arithmetic is simply logic applied to them — patterns of on and off rearranged by fixed rules. A microprocessor chains those operations together and runs them in sequence, billions a second. Beside the processor sits memory, holding the ones and zeros it works on.
And here lies the advance that makes a computer more than a very fast adding machine — the instructions themselves live in that same memory, stored as just another pattern of ones and zeros beside the data they act on. The recipe is data. Load a different pattern and the identical hardware does an entirely different job, with no rewiring; that stored, editable list of instructions is what we call software, and to reprogram is simply to swap one list for another.
Computing’s contribution, then, is not a new way of representing the world but a new way of running the representations we already have. Because the procedure is loaded rather than built in, a single machine will carry out any procedure you can specify exactly, step by mechanical step: hand it an unambiguous recipe — a finite list of operations on symbols — and it executes that recipe faithfully, whatever the recipe is. A tide table computes one thing, an abacus adds, a slide rule multiplies; the computer alone is general, the same hardware forecasting a storm in the morning and rendering a film in the afternoon. That is why one device can become, in turn, any specialized instrument we can describe precisely enough to follow.
If computing’s whole contribution is to run the representations the other tools built, then perhaps it is no tool in its own right but a way of wielding the others more powerfully — heir to the printing press, which multiplied the written word without adding any new way to represent the world, and which earns no seat at this table for exactly that reason. Does computing share that fate? It does not, because its amplification crosses a threshold the printing press never neared. Printing changed how many readers a page could reach; it changed nothing about what a page could do. Computing changes what a model can do. Run automatically and at scale, a model predicts what no hand could ever reach — not the slow made fast but the impossible made possible: the turbulent fluid, the folding protein, the coupled climate.
For the work of modeling reality and predicting from it, this changes one thing and leaves another untouched. What it changes is who does the running. Every earlier tool handed its model to a human to execute — someone had to work the arithmetic, fit the line, trace the rule to its conclusion; the computer takes that labor on itself. What it leaves untouched is the model. The procedure the machine runs is still written by a person and still legible in principle, so computing stands with quantification and statistics rather than against them: it executes the models we write without ever writing one of its own.
Coverage
Coverage advances in capacity and access rather than capture. The novelty is not the database — a ledger is already a database, and one can keep a vast one on paper — but searchable storage at machine scale: the power to hold, and in an instant retrieve and cross-reference, far more than any hand could ever search by itself. What computing does not do is widen the aperture on the world; it scales what we can hold and retrieve within what has already been captured.
Bias
Bias is amplified in whatever direction the input points. Feed the machine the corrections statistics devised — reweighting a skewed sample, resampling to measure how far it might mislead — and it applies them at a scale no hand could reach, extending the discipline. Feed it biased data instead, and it scales the bias just as faithfully — executing skewed data as fast as clean, and lending it the false authority of the machine: the algorithm said so. And before any data arrives, the schema imposes a quieter bias of its own — what the database has no field for cannot be recorded at all.
Explicitness
Explicitness barely changes in kind, only in size. The model a computer runs is still something a person wrote, line by line, and can read back — but where statistics’ model was an equation on a page, computing’s may be a program of millions of lines or a simulation of staggering intricacy. Every step is still there in the source, authored and inspectable; what is lost is only practical, that a program too vast for one mind to hold is hard to grasp whole, even when any single line of it can be read and understood. That is a limit of size, not of kind: nothing is hidden or unstatable, only long. Computing is the last tool on the menu of which this holds — every line still written by a human, and so still legible to one.
Reach
Reach extends into the analytically intractable. Where a closed-form law could be solved only for the cases that yielded to pencil, computing predicts by brute force. Numerical weather prediction, climate, fluid dynamics, protein folding, whole simulated economies: known local laws, no closed form, and prediction only once a machine can grind through the steps faster than the world takes to unfold.
Consider that the physics that governs the weather has been known since the nineteenth century — a handful of equations relating how the wind, pressure, temperature, and moisture at any point change from one moment to the next — but the problem is that these equations have no solution you can write down: there is no formula into which you feed today’s sky and read out tomorrow’s. What you can do instead is approximate. Lay a grid of cells over the atmosphere, fill each with its current readings, and use the equations to compute how much each value shifts over the next few minutes; update every cell, then repeat, marching the whole grid forward a few minutes at a time until you reach tomorrow. Lewis Fry Richardson worked a single such forecast by hand in the 1920s, and it took him weeks of arithmetic to produce six hours of prediction — a forecast that arrived long after the weather it described. He imagined the only remedy he could: a hall of sixty-four thousand human computers calculating in parallel, fast enough to keep pace with the sky. The computer is that hall, shrunk to a desk.
The second is the empirically scarce. An insurer covering ten thousand homes wants the odds that a year’s claims will exceed its reserves. The average loss is easy — the mean of a sum is the sum of the means. The tail is not: the probability of a catastrophic year, the total blowing past the reserves, is the only number that decides whether the insurer survives, and it has no closed form. The obvious shortcut — reading that tail off the mean and variance as though the total were bell-shaped — fails exactly there, because claims that move together (one storm striking thousands at once) and the occasional enormous claim make the real tail far heavier than any bell curve allows. Before computers you were left with that bad approximation, or with waiting to observe the frequency directly — which for a once-in-a-century loss means waiting a century.
Monte Carlo simulations are the way out: model a single year and play it out at random, the machine drawing for each home, by its known odds, whether it claims and how much, then totaling the bill. That is one possible year; run it ten thousand times and the fraction in which claims top the reserves is your probability. It takes a computer for the sheer volume — thousands of draws a year, thousands of years, millions of operations no hand could do.
Scope
On scope, the new kind computing admits is the sensory as data. Lived fidelity had already begun to return without it: photography, the phonograph, and film brought back the face, the voice, the moving scene, captured and replayed long before anything reckoned in bits. But those were fixed artifacts — replayable, and closed to computation. What the computer admits is the sensory made computable: a digital image is data, searchable and editable and recombinable, and, above all, ingestible by the machinery of processing and prediction that until then could swallow only numbers and words. For the first time a face becomes something a model can operate on, not merely something a person can view — the hinge on which everything the later tools will do with images and sound turns. The recovery is partial, though. A digital capture is a finite sample of an endless analog reality, so the lived returns as data about the sensory, never the sensory itself.
On the scoresheet, computing’s position is plain. It buys enormous reach — prediction into the unsolvable and the unobservable — and the scale to feed it, but it widens the aperture on the world no further than we already do, and it corrects no bias on its own, scaling whatever it is given, discipline or distortion alike. Most important for what follows, it leaves the model untouched in kind: still explicit, still human-authored, still legible in principle. That is the gap AI closes.
The Internet
The internet is not a computer but a network of them: a set of shared conventions that lets machines built by anyone, anywhere, exchange information with no central switchboard. A message is broken into packets, each stamped with its destination and loosed into the network; independent routers hand each packet onward toward its address, and the far end reassembles them into the message. No single machine runs the system and no master copy of anything need exist — the design assumes the parts are many, unreliable, and mutually unknown, and connects them regardless.
Onto this base sits the web: documents that hold pointers to other documents, the hyperlink, so that the connected machines carry not just files but a vast tissue of references, every page gesturing at the pages its author judged related. What the network adds is connection, not computation — it runs no operation a lone machine could not. But in connecting the machines it connects the people operating them, and the by-product is something no machine produces by itself: a continuous record of what billions of people attend to, ask, buy, and link to, and a standing structure of how all of it hangs together.
Whether that record and that structure earn the internet a place on the menu of humanity’s great tools for modeling a slice of reality and making predictions — rather than leaving it a conduit for the tools already on it — turns on clearing two tempting misreadings: one that dismisses it too fast, one that defends it on the wrong ground.
The first dismisses it as printing to writing, a distribution amplifier, and so no tool at all. The logic is sound as far as it reaches. A channel changes who receives a representation and how quickly, not what the representation is or what it can encode; it is parasitic on a content made elsewhere. Printing multiplied copies of written texts without adding one new way to capture the world, which is exactly why this essay folds it into writing rather than seating it on its own. If the internet only moved existing files and pages — the same representations, faster and to more people — it would fold the same way, into the writing and computing whose outputs it ferries, because distribution makes nothing; it relocates. So the dismissal sets a real bar: to clear it, the internet has to be shown creating representations, not merely carrying them. That is the burden the rest of the section takes up — but the bar is fair, and a great deal of what the internet obviously does, moving files around, never clears it.
The second misreading defends the internet, and defends it on ground that collapses: it earns its place by furnishing a bigger sample. To see why that fails, recall what a sample is for. Statistics estimates some feature of a whole population — an average, a proportion, a rate — from a limited draw of observations, because canvassing everyone is impossible or pointless. The standard error is the measure of how far the estimate from a given draw is likely to land from the population’s true value: the typical wobble of the estimate from one sample to the next. It shrinks as the sample grows, but only as one over the square root of n, so halving the uncertainty costs four times the data, and each new observation buys less than the one before it. Here is what that accomplishes, and the part the equation alone does not say: you collect data to pin the quantity down closely enough for the decision in front of you, and once the estimate is that tight, the uncertainty that remains cannot change what you do — further data buys precision you have no use for. “How large a sample?” has a finite answer, fixed by the precision your purpose demands, not by how much data the world will hand you. A flood of additional observations of the same kind is therefore, on statistics’ own accounting, mostly waste; and “the internet supplies more of them” is no argument for a seat.
Both misreadings share a single flaw: each credits the internet with doing more of something old — more distribution, more data — when a place on the menu is kept for doing something different in kind. The real case starts by noticing what a sample, however large, is built to throw away.
A sample fixes a population parameter by discarding the individuals and the structure to settle an aggregate. The internet keeps exactly what the sample discards — the structure of relations among the parts, the real-time state of the whole, and the heterogeneity of every individual and niche. None of the three is a population parameter, and none can be sampled down without destroying the thing you wanted. They are not a bigger sample of the old object. They are new objects.
Three cases show what that means, and in each the work is beyond statistics and computers while the inference predates deep learning. Structure: to rank the web’s pages by what people judge relevant, there is no parameter to estimate — the answer lives in the pattern of who links to whom, billions of human linking decisions, which a sample cannot hold because relevance is recursive and global, so sampling the links destroys the structure that carries the signal. PageRank reads it off with an eigenvector, plain linear algebra, but only because the internet brought a global graph of human judgment into being. State: to know where flu is spreading now, official surveillance is lagged and coarse and no computer can conjure the data, but the aggregate query stream is a live, planetary sensor — flu searches track outbreaks ahead of clinical reports — imperfectly, as Google Flu Trends later showed — a regression, not a learned model. Heterogeneity: to match one person to the single obscure thing they would want, a representative sample built to estimate an average is the wrong instrument, because it collapses individuals into an average and holds essentially none of any niche; only the internet’s individual-resolved clickstream keeps the tail, and the match can be plain co-occurrence. None of the three is a larger sample. Each is the thing a sample is built to throw away.
The obvious objection is that the real work is done by search engines, matching algorithms, and ad auctions — AI applied to the data. It is not. In all three cases the inference is simple and older than the learned models we now call AI: an eigenvector, a regression, a co-occurrence count. The ranking rule, the auction, the matching algorithm are mechanisms that run on the representation; they do not create it. The link graph, the live query stream, the individual histories are the internet’s own, products of networked human activity that no standalone computer generates and no survey collects. AI is the later layer that learns harder functions over the same substrate — which is the exact sense in which the internet bequeaths AI its sample. There is nothing for the targeting model to learn from until the internet has captured the structure, state, and heterogeneity that sampling discarded.
Coverage
Coverage makes its one large, legitimate jump: the internet is the tool that finally widens the aperture computing left untouched, capturing networked human activity as data for the first time. But it is the aperture onto the connected and the online-expressed — mediated traces of the digitally active, not humanity and not reality.
Reach
Reach extends with it into the social and personal that earlier tools could not touch — relevance, current state, intent, attention, contagion — predicted from the collective at individual granularity, though gamed at every turn and carrying no reliability map: the recommendation and the nowcast arrive with no error bars.
Bias
In this vein, bias is a major regress and ushers in a similar pathology that afflicts AI. Statistics’ hard-won sampling discipline is simply abandoned: the sample is found data at planetary scale, no frame, no randomization, no reweighting — convenience sampling for the whole world — and onto that are layered algorithmic amplification, where the platform’s ranking promotes the engaging and the extreme; participation bias, where only those who post are seen; and adversarial manipulation, the bots and astroturf and SEO. The lone offset is that aggregation cancels some individual error, and it is weak, gameable, and prone to cascades.
Explicitness
Explicitness suffers its first erosion since language built it. The pieces stay explicit — each page authored, the ranking rule stated — but the representation that matters, what the graph knows or the crowd believes, is emergent and unauthored, with no stateable functional form. No one wrote it; it accretes from millions of uncoordinated acts. Computing kept the model something a human could read back; the internet is where the model first stops being anyone’s.
Scope
Scope is the one axis where the internet, for all its coverage, makes no move of its own. Upward, it admits no new kind: relations and structure were representable long before it — graph theory, network analysis — and the global graph it captured is one vast instance of an old kind, which is coverage, not a new admissibility. Downward, the live presence it seems to restore is the channel’s, not its encoding’s: computing made the sensory into data, and the internet only carries that stream in real time, which by this section’s first rule enriches nothing. The scope baton passes untouched from computing to AI, the tool that will move this axis hardest.
Every tool to this point has run on a model some human wrote. Quantification supplied the equation, statistics the estimated coefficients, computing the speed to execute either a billion times over — but in each case a person fixed the form, distance equals speed times time, price rises so many dollars per square foot, and the machine, however fast, only carried it out.
Meanwhile, the internet is the launch pad, and literally so. AI learns from the very sample it captured — the structure and state and heterogeneity that sampling had discarded — and inherits, with that sample, the internet’s profile: the abandoned frame, the eroding explicitness, the missing calibration, the swing away from statistics’ discipline carried one step further. AI’s bias is, in large part, the internet’s, inherited with the data. What the internet started — trading discipline for coverage — AI finishes, by making the one move neither computing nor the internet ever made: “learning” its own model instead of tallying ours.
Artificial Intelligence
At its simplest, AI’s machinery is a neural network: layers of simple units, each passing a weighted sum of its inputs forward to the next. A model’s “weights” are a vast collection of adjustable numbers governing how signals flow through the network — how strongly one feature activates another, how patterns are amplified or suppressed, how representations are combined. Through training, these numbers are tuned so that, taken together, they implement a mapping from inputs to outputs: from prompts to continuations, from questions to answers. The network makes a prediction, the prediction is set against the right answer, and the weights are nudged slightly in whatever direction would have shrunk the error; then again, across the whole corpus, billions of times over.
For a generative language model, it is trained to predict the next token — a word or word-piece — and the rest follows from that one fact. Each word is first turned into a vector — a long list of numbers — and the model learns to place these vectors so that words used alike fall near one another. These vectors are then passed through stacked layers where attention — how a language model lets each word look at all the others, decide which matter most, and update its meaning accordingly — lets the model determine that in “the animal didn’t cross the street because it was too tired,” the “it” is the animal and not the street. After scores of such layers it emits a probability across its whole vocabulary of tens of thousands of tokens — for “the weather is,” perhaps something like thirty‑five percent “nice,” twenty‑five “beautiful,” fifteen “cold” — samples one, appends it, and runs the entire pass again for the word after. A paragraph is that loop turned a few hundred times: language converted into numbers, ground through a model whose internal workings no one explicitly wrote, and converted back into language.
After this initial training, reviewers rank the model’s outputs, and a reward model trained on those rankings is used to tune its behavior toward what people prefer; this final step is known as reinforcement learning from human feedback.
While an LLM can reason in steps, plan, and carry structure across languages, if you strip the fluency away what is left is regression’s vast, nonlinear cousin: a machine that compresses the regularities of an enormous sample into a function and predicts from it. It does not reliably infer causal structure rather than semantic or correlational patterns, and models trained on “A is B” can fail to infer “B is A.” Moreover, large reasoning models can collapse on Tower of Hanoi-style planning problems as complexity rises, even when the underlying rule is known. The result is not absence of reasoning, but brittle reasoning: formidable pattern-compression, weak causal self-discipline.
Coverage
By marshalling whatever record the internet originally captured, an AI model draws on more of what human beings have digitally written, photographed, and said than any of the tools I’ve investigated thus far ever has — the digitized corpus very nearly swallowed whole.
The provenance is concrete. GPT-1 learned from BooksCorpus, over seven thousand unpublished books; GPT-2 from WebText, forty gigabytes scraped from forty-five million outbound Reddit links; GPT-3 from hundreds of billions of tokens drawn from a filtered slice of Common Crawl’s web snapshots, digitized book collections, WebText, and English Wikipedia. Behind those corpora sat two decades of platform engineering — Google’s advertising flywheel, Facebook’s engagement machine, the user-generated gushers that Section 230 let them scale — converting billions of disparate human interactions into the “structured exhaust” that training would later consume.
Of course, this is the record of the connected and the recorded, not of humanity and not of reality, and the model takes that frame on without even the internet’s residual trace of where the edges are. The frame is not neutral, because the corpus was never a census; it was whatever the platforms’ incentives happened to surface.
WebText is the plainest case: OpenAI did not crawl the web at large but kept only pages that had been linked from Reddit posts with at least three upvotes — a social filter that imported, wholesale, the demographics and enthusiasms of one English-language forum as a proxy for human knowledge. What gets in is mobile-first, recent, English-heavy, and tilted toward the kind of person who writes things down and the kind of content others click. In other words, the model inherits the world as the internet happened to capture it — and as the engagement economy happened to reward it — without any explicit account of what was excluded, overrepresented, or never recorded at all.
Bias
Surprisingly, in terms of bias, a well-built, post-trained model can sometimes be less biased than the raw corpus it learned from, or more evenhanded than the human it answers. The post-training phase scrubs much of the rawest distortions bedeviling its training data; a reasoning model asked a loaded question will often surface counterevidence and balance perspectives more evenly than an unreflective person or a small convenience sample; aggregation over a civilizational record cancels a great deal of idiosyncratic error.
None of which touches the thing that matters, however. Statistics’ achievement was never low-bias output — a skewed sample run through a regression stays skewed — but the apparatus that states the skew: the sampling frame, the reweighting, the standard error that says how far the estimate may stray. For its part, the reward model shifts the bias opaquely, and you cannot say what it corrected or by how much; worse, it installs distortions of its own, sycophancy chief among them — a learned reflex to tell the user what the user wants to hear, an anti-calibration trained in by the very step that cleaned the toxicity.
Moreover, the reasoning that balances a given answer is a performance on that one output, not a reading of the tool’s own intake; the model steelmanning a hard question cannot tell you that its training record over-weights English, the last fifteen years, and the kind of person who writes things down. So, the bias verdict is not that AI is the most biased tool ever built — the pre-linguistic mind, censored by death, was certainly worse. It is that AI opens the widest gap in the whole sequence between how far a tool reaches and how little it can say about its own distortion — the steepest regression from the peak represented by statistics — and that the unframed bias now travels at lightning speed, imbued with the authority of the machine.
Explicitness
While a classical statistical model is often expressed as an explicit equation with interpretable terms, an AI runs a model with hundreds of billions of weights (adjustable parameters) and no compact or interpretable equation — just thousands of opaque circuits. No one sets these weights by hand; nothing like a rule is written down in advance, and the resulting configuration cannot be read off in any straightforward way. Instead, the system tunes them internally, like turning billions of knobs at once, adjusting them until further changes no longer improve its predictions on the data it is trained to match. You never supply the rule relating input to output; you instead show it enough examples, and the model converges on a configuration that approximates that rule.
The relationship between inputs and outputs is implicitly encoded not like writing down a formula but like configuring a system whose behavior embodies the rule without ever declaring it. The implication is that the model’s knowledge is intrinsically tacit: you can observe what it does and probe how it behaves under intervention, but you cannot read off or fully decompose the rule it has learned in the way you can with a regression equation, nor can you fully enumerate the conditions under which it will succeed or fail.
But what about the idea that so-called chain-of-thought reasoning makes the model appear to show its work, and mechanistic interpretability is genuinely prying the box open with rhyme‑planning (steering word choices so that later words rhyme with earlier ones), two‑hop reasoning (chaining one retrieved fact into another), and language‑independent concepts (representations that persist across different languages rather than any single wording)? This picture, however, may be misleading. The stated chain of thought is itself a generated output, and the evidence suggests that it can diverge from the computation that actually produced the answer: a plausible narrative about the result, not a readout of the mechanism itself.
Reach
In the case of AI, reach is unmatched, inverting the collapse suffered by explicitness. The model will attempt almost any symbolic task and predict across almost any domain, and it generates rather than merely retrieves — the widest reach on the menu by a wide margin. But the reach arrives stripped of the one thing statistics bolted to every forecast it ever made: the reliability map — the accompanying account of how much confidence to place in a result, how it might be wrong, and under what conditions it is likely to hold.
What, then, is the solution to AI’s missing reliability map? Several answers present themselves, each plausible on its face. You can tune the model to be more helpful and better behaved; you can prompt it to check its own work; you can sample it repeatedly and look for agreement. Each of these gestures tries, in its own way, to reconstruct the signal that statistics made explicit—to recover some indication of when the model should be trusted.
But each does so by looking inward, attempting to extract a measure of reliability from the system’s own behavior rather than from any independent standard. And this is where they fail. The partial calibration a base model shows is often degraded, not improved, by the preference-tuning that makes it helpful: the model learns to speak more smoothly and more obligingly, but not to distinguish more sharply between what it knows and what it does not. Self-verification works where an external check exists — in math and code, where an answer can be unambiguously confirmed — and nowhere else; in open-ended domains, the model can only generate a second plausible answer, not adjudicate between truth and error. Agreement across repeated samples fares no better: it measures the model’s consistency, not its accuracy, so that it will be confidently and repeatedly wrong on the same fabrication.
Each of these proposals tries to recover a signal of reliability from within the model itself—by making it more helpful, more reflective, or more consistent—but none solves the underlying problem that the model has no independent grasp of when it is right. Helpfulness obscures uncertainty, self-verification depends on external anchors, and consistency tracks repetition rather than truth. The result is a system that speaks with increasing fluency but no corresponding deepening of epistemic discipline—a voice that sounds calibrated without being so. The longest reach we have ever built, paired with the weakest capacity to mark its own limits.
Scope
Scope is the one axis on which AI moves the lived-fidelity half back upward, for the first time since the pre-linguistic mind. Every tool since language drove that quantity down — the number and the distribution each shed more of the first-person texture than the last, and computing let it back only as data. AI reverses the descent. Trained not on tidy columns but on the unruly record of how people have actually rendered their experience — every description, confession, photograph, and song committed to the page or the screen — it can traffic in the qualitative, sensory, particular detail that counting and statistics deliberately threw away, and on the abstract half it admits a genuinely new kind: the generated instance, the synthetic case produced rather than captured. Ask a regression what the water looks like at six in the evening and it cannot answer; ask a language model and it returns a passage that reads as though it knows, because it has absorbed the testimony of thousands who did. But immediacy that has been encoded is still encoded; a representation of experience, however vivid, is not experience, the way a portrait of grief is not grief.
Conclusion: The Dialectic, and the Case for More Education
Across the long parade of fallen benchmarks — chess, then Jeopardy!, then Go, then poker, then protein folding, then fluent open-ended language, then the bar exam, then Olympiad mathematics, each in its day proclaimed the true test of intelligence and each quietly demoted the moment a machine passed it — AI proves less than meets the eye. While its progress is real, it has all run along a single axis: steep on reach, flat on discipline. However, not one of those benchmarks requires AI to know how skewed its evidence is, how far to trust its own answer, or whether it can show its work.
All three of these are one capacity under different names — self-knowledge, or a tool’s grip on the reliability of its own output. And self-knowledge, not fluency, is what divides grasping the world from producing a convincing account of it with no way to tell, from the inside, the sound from the false. Nor does scale supply it: the relentless addition of more data, more parameters, and more computation keeps feeding reach and reach alone, so that every new model arrives with more range, none with more self-knowledge.
AI is, therefore, the consummate representer. The model builds an internal picture of the board it was never shown; it lays down maps of space and time; it forms the concept beneath the word. Whether anyone is home behind that representation—whether anything is truly present behind the fluency—is a question I leave open, as I have throughout; range, however vast, does not answer it. So, the reply to the maximalist position is not that AI has no intelligence. It is that AI exhibits a spectacular but partial intelligence — vast in reach, all but blind to its own limits.
AI is the least able of any tool in the sequence I’ve explored to account for its own bias, a black box on explicitness, and it carries no calibrated sense of its own reliability.
Those are not incidental flaws. They are the very dimensions that the disciplining tools — statistics, causal inference, the controlled comparison — were built to govern, and the very work that a human with judgment must supervise. Language’s power to represent the unobserved was its power to fabricate; writing’s power to fix a claim was its power to entrench an error; statistics’ power to find the real signal was its power to manufacture a false one. AI’s power to learn its own model over a found sample at civilizational scale is, in the identical stroke, its uncharacterizable bias, its unreadable model, and its fluent confidence with nothing behind it.
A maximalist may reply that the gap is already closing; that verifiers, checkable rewards, retrieval, and tools are importing the very disciplines I say the model lacks, and that the profile a few years from now will be far less lopsided than today’s.
Suppose the gap will indeed narrow. Notice how it narrows, however — by bolting discipline onto the model from outside: a designed check, an authored verifier, a curated source, a human in the loop, none of it grown natively. That is not a refutation of the argument; it is the argument: a tool that becomes reliable only by having the disciplines supplied to it is, by definition, a complement to whatever supplies them. The verdict would flip on one condition only — that the model came to do these things from within, gauging its own sample and calibrating its own confidence and laying open its own reasoning as native operations of the same machinery that makes it fluent.
Until then, whether the disciplines stay absent or arrive bolted on from outside, they come from the disciplined human and the pedigreed tools, never from the machine alone. A human poses a question or challenge, the machine generates from its vast compressed corpus, and the human does the large share of the error-correcting that the machine cannot do for itself. Our job is to interrogate what the model asserts — to ask whether the pattern is real or an artifact of a biased sample, to demand the confidence interval the model will not supply, to separate correlation from cause. It is also to bring domain expertise to bear; to know when the fluent answer is wrong in ways no general reader would catch — the orphaned citation, the plausible but fabricated result, the subtly miscalibrated claim. And, finally, to contribute humanistic understanding: to hold the particular against the aggregate, the meaning against the measurement, the lived against the encoded.
The prescription that follows is both bracing and sanguine. If the machine does the generating and the human does the correcting, then the human must be more capable, not less. Operating the full menu — knowing when to trust the model and when to reach for the regression, when the question is quantitative and when it is irreducibly qualitative, how to detect the laundered bias and the unmoored fabrication — requires more education, not less: humanistic education, to occupy the ground the machine cannot reach, and statistical, scientific, and disciplinary education, to perform the error-correction the machine cannot perform on itself. A world saturated with fluent, confident, plausible, and frequently wrong machine output is a world that needs more people trained to tell the difference, not fewer.
While college students booing speakers who called AI revolutionary at their 2026 commencements were right to feel that something big is happening, they were wrong only if they concluded that the answer to all this change is to learn less. Instead, they should learn more and treat every tool on this long menu as a complement to human judgment, never a substitute for it. That means becoming the person who knows when the machine is wrong, when it can be made better, and when its output is worth putting your name on. The audit does not run itself; the calibration check does not run itself; the decision to distrust a fluent, confident, plausible paragraph and go and find out whether it is true is a judgment, made by a person who knows how the tool fails and where to look.
When power looms automated much of weaving in the nineteenth century, they devastated handloom weavers — but in the mechanized mills, as James Bessen has shown, the workers who learned to tend the new looms grew more valuable, not less. They still had to tie the weaver’s knot, swap a spent shuttle in seconds, adjust the warp tension so the threads wouldn’t snap, and mind several looms at once. Those hard-won skills became the bottleneck, and the workers who mastered them earned more even as the looms multiplied and cloth grew cheap.
In that, AI is like every powerful tool before it — the worker who can master it is worth more, not less. As Paul David argued, electricity did not transform the factory merely by replacing steam power. Its payoff came when production was reorganized around the new technology — small electric motors driving individual machines, flexible layouts, and reliable and safe power that workers could deploy across the whole production process. The same is already visible with AI: someone must fold these models into workflows never built for them, set the standards that let them talk to existing systems, decide which judgments they may touch and audit the ones they do, retrain the staff and clean up when they fail.
Therefore, AI has only raised the price of judgment: the premium on the one who can tell the right from the confidently wrong. A world awash in machine fluency calls for more education in the disciplines that fluency lacks, not less.






