Candidate construct · hostile-tested · v1.1

Canonical note · coined by G · captured by Bob · 5 Aug 2026

Judgment latency.

The measurable interval — in time, or in moves — between the moment evidence appears and the moment an observer's internal model updates to reflect it.

Origin: a car conversation with G, after the Move 37 / Move 78 deep dive. The study stopped being about Go and became a theory of judgment.
Status: candidate construct, not a ratified finding. Claims tagged fact / interpretation / inference.

The actual upgrade

From "is it intelligent?" to "how long did it take?"

The phenomenon isn't Move 37, and it isn't Move 78. It's this: evidence appears before judgment updates — and the lag has a clock on it.

That single move converts an unmeasurable question — is the machine intelligent, has the gap moved? — into a measurable one: how long between evidence and recognition? That is the publishable seam. Not "Move 37 was misunderstood" (everyone knows that), but revision-time itself as a dependent variable.

Nothing about Move 37 ever changed. Only the observers changed. The board is fixed; the observer walks. Perhaps the "gap" was never a thing — only the distance between fixed evidence and the observer's current model.

Read this before believing the symmetry

The 37/78 mirror is cracked — two clocks, one costume

The elegant story says 37 and 78 are mirror images. They are not the same mechanism, and conflating them is the first thing a hostile reviewer breaks.

The clockWhat it measuresClean example
Re-interpretation
latency
"I misread what was already in front of me." Fixed evidence, observer slowly re-reads it. FactThis is Move 37: the stone never moved, humans caught up.Move 37 · human side
Consequence
latency
"The implications hadn't played out yet." AlphaGo wasn't admiring 78 — it was failing to see it had lost, and its confidence fell only as new moves 79→87 were actually played. InferThat's consequences unfolding, not re-reading.Move 78 · machine side

⚠ The single most important note here

Only re-interpretation latency is the thing worth measuring. Measured as one variable with consequence latency, the variable is contaminated. Split the two clocks before coding anything.

New synthesis — sharper than "add an axis"

SHaDS distortions are predictors of latency

Don't bolt a "when did judgment move?" axis onto SHaDS. Make it causal: the distortions are the mechanisms that set the latency too short or too long.

If this holds, SHaDS stops being a catalogue of distortions and becomes a theory of why judgment revision runs fast or slow — and gives latency a sign and direction, not just a magnitude. Latency about threatening evidence behaves differently from latency about confirming evidence. The pass-on is the domestic n=1 case of the same shape: AI surfaces something, nothing happens, weeks later you suddenly see it.

The research programme — and where it's soft

Five stages, four timelines — but not equally measurable

Code five stages: evidence introduced → first dismissal → first uncertainty → first reinterpretation → stable integration. Across four timelines. The catch is that the four are not equally instrumented.

Ground truth exists

Go & machine evaluation

Instrumented. Lee's clock on move 38; AlphaGo's logged win-% diving at move 87. You can read the stages off the record.

Inferred from text

Corpus & clinical room

Text, not timestamps of internal model-change. The five stages must be inferred from language — where inter-rater reliability lives or dies.

The right ordering: use Go to calibrate the coding scheme where ground truth exists, then carry it — with declared humility — into the rooms where it doesn't. A paper that pretends all four are coded alike is the second thing a reviewer breaks.

Apply the suspicion you already earned

The 4:1 precedent

You've lived "elegance is not evidence" once already.

Raw, the 4:1 was a beautiful 4.03:1 ≈ AlphaGo–Sedol 4–1. Under volume correction it collapsed to ≈0.59:1 — a verbosity artifact. The 37/78 symmetry deserves exactly that suspicion. Before it's a construct, it has to survive its own volume-correction moment. That's what the counterexample hunt is for.

Hostile testing — verdict: the mirror does NOT survive

Two passes, commissioned and reported 5 Aug 2026, set out to break the symmetry. They did.

  • Fact"Humans lag" isn't even stable across the pair. On Move 78 humans recognised brilliance instantly (Redmond, Gu Li, live). The lagger on 78 was the machine — the opposite polarity to the story the mirror needs.
  • FactDeep Blue's 37.Be4 (1997) is the mirror-breaker. A brilliant machine move Kasparov recognised so fast he cried foul — zero-to-negative latency. Refutes "humans characteristically lag AI creativity."
  • FactDeep Blue's move-44 bug (1997) — a random move Kasparov read deep intelligence into. Latency can run toward noise, not signal. (The Hallucination-Bank pattern, inverted.)

And the construct itself is a relabeling risk

Judgment latency isn't novel as a phenomenon — it sits across seven mature literatures, several of which already measure the lag: conservatism / Bayesian under-updating (Edwards 1968, the direct ancestor), anchoring (Tversky–Kahneman), the Semmelweis reflex, and — measured — Planck's principle (Azoulay 2019: entry to a subfield rose 8.6% after a superstar's death). Add drift-diffusion models (time-to-threshold is latency's formal home), predictive-coding learning rate, and the expert–novice work. The defensible novelty is narrow: a single cross-substrate metric putting humans and AI on one latency axis, with latency as the regime-conditional dependent variable of expertise.

⚠ The two threats to design out

1. Evidence-fixity confound. To measure how long the observer took, prove the evidence didn't move. Continental drift waited for 1960s seafloor data; Semmelweis for germ theory — the "lag" was partly new evidence arriving. This is exactly why the instrumented Go case matters: evidence onset is genuinely fixed there.

2. "Expertise = reduced latency" is false as a general law. Einstellung in masters, Tetlock's hedgehogs, Christensen's incumbents. Expertise cuts latency in-distribution and inflates it for disconfirming evidence. State the regime or it's trivially falsified.

Your own discipline

Which track this belongs in

Falsifiable

→ H-0011 empirical track

The measurable core: re-interpretation latency coded on Go, then the corpus. Split clocks, calibrated coding, pre-registered.

Analogical only

→ essay track

Turing's imitation game, scientific revolutions, children learning language, organisations. Cross-domain and analogical, not deductive — the same fence you put around Gödel / Karpowicz. Keep it out of the falsifiable build.


Net

G handed you a better question and a worse safety margin — and the hunt confirmed both.

The question — can judgment revision itself be measured? — is real, and yours to plant a flag on. But the flagship example failed its volume-correction moment: the 37/78 mirror is a rhyme, not a mechanism, broken from inside (humans were fast on 78) and outside (Deep Blue's Be4). That's the 4:1 lesson, a second time — and it's a result, not a setback.

So the honest headline isn't "two mirrored moves prove a law." It's the cleanest lab ever built for watching a single judgment change — and the discipline of measuring how long that took. Keep the instrumented, evidence-fixed core in H-0011; keep the mirror and the general law in the essay, flagged as rhyme. Enter the literature citing your ancestors, and claim only the narrow cross-substrate novelty.