Canonical note · coined by G · captured by Bob · 5 Aug 2026
The measurable interval — in time, or in moves — between the moment evidence appears and the moment an observer's internal model updates to reflect it.
Origin: a car conversation with G, after the Move 37 / Move 78 deep dive. The study stopped being about Go and became a theory of judgment.
Status: candidate construct, not a ratified finding. Claims tagged fact / interpretation / inference.
The actual upgrade
The phenomenon isn't Move 37, and it isn't Move 78. It's this: evidence appears before judgment updates — and the lag has a clock on it.
That single move converts an unmeasurable question — is the machine intelligent, has the gap moved? — into a measurable one: how long between evidence and recognition? That is the publishable seam. Not "Move 37 was misunderstood" (everyone knows that), but revision-time itself as a dependent variable.
Nothing about Move 37 ever changed. Only the observers changed. The board is fixed; the observer walks. Perhaps the "gap" was never a thing — only the distance between fixed evidence and the observer's current model.
Read this before believing the symmetry
The elegant story says 37 and 78 are mirror images. They are not the same mechanism, and conflating them is the first thing a hostile reviewer breaks.
| The clock | What it measures | Clean example |
|---|---|---|
| Re-interpretation latency | "I misread what was already in front of me." Fixed evidence, observer slowly re-reads it. FactThis is Move 37: the stone never moved, humans caught up. | Move 37 · human side |
| Consequence latency | "The implications hadn't played out yet." AlphaGo wasn't admiring 78 — it was failing to see it had lost, and its confidence fell only as new moves 79→87 were actually played. InferThat's consequences unfolding, not re-reading. | Move 78 · machine side |
⚠ The single most important note here
Only re-interpretation latency is the thing worth measuring. Measured as one variable with consequence latency, the variable is contaminated. Split the two clocks before coding anything.
New synthesis — sharper than "add an axis"
Don't bolt a "when did judgment move?" axis onto SHaDS. Make it causal: the distortions are the mechanisms that set the latency too short or too long.
If this holds, SHaDS stops being a catalogue of distortions and becomes a theory of why judgment revision runs fast or slow — and gives latency a sign and direction, not just a magnitude. Latency about threatening evidence behaves differently from latency about confirming evidence. The pass-on is the domestic n=1 case of the same shape: AI surfaces something, nothing happens, weeks later you suddenly see it.
The research programme — and where it's soft
Code five stages: evidence introduced → first dismissal → first uncertainty → first reinterpretation → stable integration. Across four timelines. The catch is that the four are not equally instrumented.
Instrumented. Lee's clock on move 38; AlphaGo's logged win-% diving at move 87. You can read the stages off the record.
Text, not timestamps of internal model-change. The five stages must be inferred from language — where inter-rater reliability lives or dies.
The right ordering: use Go to calibrate the coding scheme where ground truth exists, then carry it — with declared humility — into the rooms where it doesn't. A paper that pretends all four are coded alike is the second thing a reviewer breaks.
Apply the suspicion you already earned
You've lived "elegance is not evidence" once already.
Raw, the 4:1 was a beautiful 4.03:1 ≈ AlphaGo–Sedol 4–1. Under volume correction it collapsed to ≈0.59:1 — a verbosity artifact. The 37/78 symmetry deserves exactly that suspicion. Before it's a construct, it has to survive its own volume-correction moment. That's what the counterexample hunt is for.
Hostile testing — verdict: the mirror does NOT survive
Two passes, commissioned and reported 5 Aug 2026, set out to break the symmetry. They did.
Judgment latency isn't novel as a phenomenon — it sits across seven mature literatures, several of which already measure the lag: conservatism / Bayesian under-updating (Edwards 1968, the direct ancestor), anchoring (Tversky–Kahneman), the Semmelweis reflex, and — measured — Planck's principle (Azoulay 2019: entry to a subfield rose 8.6% after a superstar's death). Add drift-diffusion models (time-to-threshold is latency's formal home), predictive-coding learning rate, and the expert–novice work. The defensible novelty is narrow: a single cross-substrate metric putting humans and AI on one latency axis, with latency as the regime-conditional dependent variable of expertise.
⚠ The two threats to design out
1. Evidence-fixity confound. To measure how long the observer took, prove the evidence didn't move. Continental drift waited for 1960s seafloor data; Semmelweis for germ theory — the "lag" was partly new evidence arriving. This is exactly why the instrumented Go case matters: evidence onset is genuinely fixed there.
2. "Expertise = reduced latency" is false as a general law. Einstellung in masters, Tetlock's hedgehogs, Christensen's incumbents. Expertise cuts latency in-distribution and inflates it for disconfirming evidence. State the regime or it's trivially falsified.
Your own discipline
The measurable core: re-interpretation latency coded on Go, then the corpus. Split clocks, calibrated coding, pre-registered.
Turing's imitation game, scientific revolutions, children learning language, organisations. Cross-domain and analogical, not deductive — the same fence you put around Gödel / Karpowicz. Keep it out of the falsifiable build.
G handed you a better question and a worse safety margin — and the hunt confirmed both.
The question — can judgment revision itself be measured? — is real, and yours to plant a flag on. But the flagship example failed its volume-correction moment: the 37/78 mirror is a rhyme, not a mechanism, broken from inside (humans were fast on 78) and outside (Deep Blue's Be4). That's the 4:1 lesson, a second time — and it's a result, not a setback.
So the honest headline isn't "two mirrored moves prove a law." It's the cleanest lab ever built for watching a single judgment change — and the discipline of measuring how long that took. Keep the instrumented, evidence-fixed core in H-0011; keep the mirror and the general law in the essay, flagged as rhyme. Enter the literature citing your ancestors, and claim only the narrow cross-substrate novelty.