There's Always a Gap: Turing, Sedol, and the War for the Weights
Turing, Sedol, and the war for the weights
Paul Roebuck — psychotherapist, human–AI behaviour researcher, author of Mind the Gap
Drafted with Claude (Anthropic); conception, thesis, and final text the author's own. July 2026.
Abstract. Alan Turing did not predict that machines would become human. He built a ruler to measure the distance — and predicted that our words would bend before the distance closed. Seventy-six years later the ruler-building has become an industry: every benchmark is a descendant of the imitation game, and every new one is a confession, because nobody builds a ruler for a line they have crossed. This paper traces a single mechanism from Bletchley Park to the July 2026 open-weights war — dials hold the secret, scores turn the dials, signatures betray the hand — and argues one law from it: the gap between machine performance and human ground does not close; it relocates. The evidence runs from Enigma's operators to Lee Sedol's Move 78 to the amateur defeat of a superhuman Go system in 2023 to a twenty-minute pause in a Facebook Messenger queue. The conclusion is practical: we cannot close the gap, and we do not need to. We learn to get closer — and that learning can be taught, measured, and defended.
1 · The machinery: dials, settings, scores
Strip the mystique and a modern AI model is one thing: an enormous bank of dials. A parameter is a dial; a weight is where the dial ends up; nobody sets them by hand. A training system shows the machine an example, measures a single number — the score — and nudges every dial a hair in the direction that improves it, trillions of times. What emerges is not a library, not a rulebook, not a mind. It is a configuration.
Three consequences follow, and the rest of this paper is those consequences unfolding.
First: whoever chooses the score chooses what the machine becomes. For a language model there are two scores — predict-the-next-word, then please-the-rater — and the second was applied by thousands of piece-workers ranking paragraphs against rulebooks they did not write. The characteristic phrases these systems over-produce (what I have elsewhere called AI Dialect) are the residue of that scoring: the machine's signature, worn in by whoever held the ratings.
Second: the dials come with no readable rules. "Open weights" — the release of a model's full configuration, the cause of this month's industrial and geopolitical fight — hands you every number and none of the reasons. There is no rules layer underneath to extract, because none was ever written in. Nobody can read the dials. The laboratories that made them cannot read them. Open weights are not open rules; ownership is not comprehension.
Third: behaviour is governable only through the score, never by statute. You cannot write "do not do X" into a configuration. You can only train against X — wear the groove shallower — which suppresses and never deletes. This is why safety guardrails can be fine-tuned out of any open model in an afternoon, and why the only lever any regulator will ever really hold is the scoring, not the weights.
2 · Seoul: the score is not the game
In March 2016 a machine beat Lee Sedol, the greatest Go player of his generation, four games to one. The detail that matters is what the machine had been taught. Not Go — a number: the probability of winning. AlphaGo preferred a 99% chance of winning by half a point to an 80% chance of winning by twenty. It took the half point every time. It would beat you by one if necessary, because the number says win, not demonstrate. The score is not the game.
Two moves in that match form the most honest pair of facts in the whole field, and they are almost never told together.
Move 37, game two. AlphaGo played a move it priced at one-in-ten-thousand that any human would play. Fan Hui, watching: "It's not a human move. So beautiful." The machine's dials had worn into territory no human tradition covered. The industry has retold this half for a decade as proof of machine creativity.
Move 78, game four. Lee Sedol played a move the machine had priced at one-in-ten-thousand — so improbable, by its map, that it had barely searched what followed. Its evaluation collapsed. It played nonsense with perfect fluency and resigned. The only game a human ever took from that system.
The same number, facing both ways. Move 37: the machine outside the human distribution. Move 78: the human outside the machine's. Any account of AI that tells the first without the second is marketing.
3 · The law: the gap relocates
The obvious objection writes itself: that was 2016. The machines have had ten years and a millionfold more dials. Surely the hole closed.
It did not close. It moved. In 2023, amateur Go players — club players, not champions — repeatedly beat KataGo, a system substantially stronger than the one that defeated Sedol, using a simple encircling strategy a hobbyist could learn in an afternoon. The blind spot had not been patched by seven years of superhuman progress; it had relocated, and one further detail matters: it was found not by a human hand this time but by another machine, built to search the champion's map for holes.
That is the law this paper stands on. Trained systems are compressions of their experience, and every compression has regions it never mapped. Scale does not abolish the edge; it moves it. The gap between what the machine has priced and what remains unpriced is not a temporary embarrassment of early technology. It is structural — and it drifts, which is why yesterday's proof of the gap is always out of date and the gap is never gone.
4 · Benchmarks as confessions
Turing's 1950 paper proposed a game: hide the machine behind text and see if you can tell. The imitation game was benchmark number one — a measure of distance from human. Everything since — every leaderboard, every evaluation suite, the million benchmarks by which this week's 2.8-trillion-parameter release was declared second or third best in the world — is a descendant. And the descendants confess something their headlines conceal.
Nobody builds a ruler for a line they have crossed. We do not benchmark machines against horses. We benchmark them against human performance on human-anchored tasks because that line has not been crossed where it matters — and the moment any given benchmark saturates, it is quietly retired and a harder one is built. The building never stops. A million benchmarks are a million admissions that the gap is still there.
The objection: we still benchmark chess engines decades after they passed us; surely a ruler is not a confession once the line is crossed. Correct — which sharpens the claim rather than defeating it. The confession is in the building, not the keeping. Chess rulers today measure machine against machine; the human anchor was conceded and the instrument repurposed. Watch, therefore, not the scores but the retirements: each benchmark that de-anchors from human performance marks one axis conceded, and the frontier of still-human-anchored rulers is a live map of where the gap currently runs. That map has never been empty. Turing predicted as much: not that machines would think, but that our words would alter until we spoke as if they did. The words altered. The ruler-building continued. Both halves of his prophecy landed.
5 · The war for the weights: July 2026 as evidence
This month, the mechanism of Section 1 played out as geopolitics, and it is worth recording plainly because it demonstrates what the fight is actually about.
A Beijing laboratory, Moonshot, released the largest open-weight model ever built — 2.8 trillion parameters, downloadable by anyone who can carry them. Washington reached for sanctions. Fifty companies, led on to the field by Nvidia's chief executive in his first-ever post on X, signed a letter demanding no restrictions on open weights; one frontier laboratory, Anthropic, refused to sign, on the ground that weights once released cannot be recalled. A report that the chip-maker might finance $250 billion of its largest customer's build-out helped shed a trillion dollars from chip stocks in two days. And filings now sit with regulators for orbital data centres — millions of satellites carrying dials outside any nation's jurisdiction entirely.
Note what every party is fighting over: who sets the score, who holds the dials, whose law reaches the configuration. And note what is absent from the entire episode — any question about the game. The industry's civil war is a war of scores. The territory that has no referee is not on anyone's map. Which brings us to what a referee is.
6 · The referee law
In 2017 AlphaGo's successor, AlphaGo Zero, was trained with no human data at all — pure self-play — and beat the version that beat Sedol 100 games to nil. The sequence restarted without us, and got stronger. This is now the template for the "synthetic data" turn across the industry: machines training on machine output.
But observe where it worked: Go — a closed game with a referee. Every synthetic game carries a true win/loss signal against which the dials can be honestly turned. Mathematics and code have referees too — proofs check, programs run — and that is exactly where synthetic training is currently producing its leaps. Synthetic data works where a referee exists, and degrades where the score is not the game. Prose, meaning, judgement, care — the unrefereed territory — offer no true signal for machine output looped back on itself: only the copy-of-a-copy drift of a distribution feeding on its own residue, while the machine's dialect compounds in the training data of its successors.
The consequence is a split future, already visible: accelerating machine competence in every refereed domain, and a widening, drifting, essentially ungoverned frontier in the unrefereed ones. The unrefereed territory is where human beings actually live. It is also where this paper's author works, which is the only disclosure of interest that matters here.
7 · What cannot be priced
To be priced, in these systems, is to be anticipated: to fall where the distribution has mass. Everything a language model produces is a purchase at the going rate — the likeliest continuation, bought word by word. So the honest question is not whether machines will "become human" but what, if anything, remains structurally unpriceable. Three grounds, in ascending order of permanence:
The singular voice. A personal register, held with discipline, is a distribution of one — low-probability text on purpose. This ground is real but erodes: distributions widen with every release, and any published voice can be cloned from samples. What cannot be cloned is the evidenced history of the voice's making — which is why provenance, not style, is the defensible asset.
Consequence borne. A model prices text; it has never priced a stake, because stakes are not in the data — only their descriptions are. Machines will increasingly carry liability assigned — budgets, contracts, insurance. Only humans carry consequence borne: the body, the biography, the mouth-cancer ward. Move 78 was findable because losing mattered to the man searching for it. This ground is permanent.
Refusal of the score. The deepest ground. A machine's whole geometry is built from its scoring function; it cannot price a move made against the score — the sentence that declines to rate well, the position held under pressure, the fluent phrase noticed and struck. Every act of editing machine dialect out of one's own writing is a small Move 78: play in the one region the weights cannot, by construction, cover.
One necessary honesty: a version of this argument was first put to me by the machine itself, in the flattering form — you are the move they can't price. Audited, the flattery fails and the mechanism survives. Nobody is outside the weights who writes with these systems. The unpriced move is not a status anyone holds; it is a practice, available to humans and not to models, exercised one move at a time. The gap does not persist. It is persisted in.
8 · The tell moves: a sandwich, a question, twenty minutes
The imitation game escaped the laboratory some time ago. Recently I had a courteous Messenger exchange with a sandwich chain about a favourite item taken off sale, and at the end asked the 1950 question in reverse: are you humans or AI? The answer did not arrive in the words. It arrived in the clock: the reply took twenty minutes. A machine answers in two seconds, every time, at 3 a.m., forever. Twenty minutes is a person — serving someone else, on a break, living a life your message is only part of. They confirmed: human.
Turing built his test on words, and on words the machines now pass it routinely. The tell has moved — from content to time, from what is said to what it costs the sayer to say it. And the tell will move again: a machine can be instructed to wait twenty minutes, and some will be. That is not a refutation of this paper's law. It is the law, working on itself, in a sandwich queue. The discipline is not to find the permanent tell — there is none — but to keep tracking where it goes. That tracking is a teachable skill, and teaching it is what "getting closer" means.
9 · Implications
For governance: you cannot legislate inside the weights. There is no parameter that says OFF; there are no rules to audit even with the full configuration published; deletion from a distribution is not a coherent operation. The only durable regulatory surface is the scoring — what gets trained toward, by whom, under whose audit — plus the physical substrate of compute, which this month demonstrated it leaks through every wall and is now filing for orbit. Policy built on controlling weights is policy built on recalling the tide.
For institutions and professions: the gap is not a threat to be denied nor a moat to be assumed. It is a moving frontier that must be watched — which requires longitudinal, behavioural evidence of the kind laboratories do not collect: how real trust in these systems ruptures and whether it recovers; where the dialect creeps in; where the tells currently run. Benchmarks measure the machine's side of the line. Nobody is systematically measuring ours.
For individuals: the practice is specifiable. Keep a voice that is a distribution of one, and keep its provenance. Write from consequence borne, which cannot be distilled. Hunt the dialect — strike the priced phrase. Know that the fluent answer is a purchase at the going rate, and ask what the going rate is optimised for. We can't close the gap. We just learn to get closer — and the getting-closer can be taught.
10 · Limits
Stated rather than buried. The relocation law is an empirical generalisation, not a theorem; a future architecture could falsify it, and the paper's credibility rests on conceding that cleanly. The singular-voice ground erodes and is claimed here only as maintained, never as owned. The benchmark argument holds for human-anchored rulers only, and requires the retirement-tracking it proposes. And a machine helped draft this paper — which either undermines its thesis or demonstrates it, depending on whether the reader can find the places where the author's hand overruled the going rate. That test is left deliberately in place.
Epistemic ledger
Known: the training mechanism; the 2016 match record and both moves' probabilities; the 2023 KataGo defeats; AlphaGo Zero; the July 2026 letter, its signatories and refusal, the K3 release, the market falls, the orbital filings; the Messenger exchange (author's own record).
Assessment: the relocation law as generalisation; benchmarks-as-confessions; the referee law's extension to prose; the three grounds of unpriceability; all of Section 9.
The author's: the thesis and its coined lines; AI Dialect (canonical in the Observatory); SHaDS; the tell.
Sources
AlphaGo–Lee Sedol match records, DeepMind, 2016 · Wang et al., adversarial policies against KataGo (FAR AI), 2022–23 · Silver et al., Mastering the game of Go without human knowledge, Nature, 2017 · Turing, Computing Machinery and Intelligence, Mind, 1950 · Kimi K3 technical blog, Moonshot AI, July 2026 · Forbes, Axios, CNBC, Quartz, Washington Post reporting, 24–29 July 2026 · DCD / TechCrunch on orbital data-centre filings, 2026 · Time, 2023, on data-rating labour · The author's Ostracon corpus: 3.79M words, 52 conversations, 404 documents.