← Back to the gauntlet

How the gauntlet judges a theory

A verdict is only worth something if the theory got a fair fight. The gauntlet is built so that every theory — whether it ends confirmed or refuted — was argued at its strongest, attacked at its weakest, and scored by the same published rubric. This page is that rubric.

Input
A community theory, sourced and attributed — the claim as its believers actually make it
Court
advocate → two independent skeptics → rebuttal → three-judge panel
Output
a locked verdict with per-sub-claim labels, four axis scores, and a "what would change our mind" list
Provenance
the same adversarial machinery that produced the locked endgame thesis

Step 1 — Steelman first

Before anything gets attacked, the theory is written up in its strongest defensible form, with attribution: where it originated, who popularized it, and the best version of the argument as the community actually makes it. A gauntlet that quietly weakens a theory before judging it isn't a court, it's a strawman factory — so the steelman is the document of record, and every later attack must engage it, not a caricature.

Step 2 — Decompose into falsifiable sub-claims

Most theories are bundles. "X is secretly Y" usually smuggles in two or three separate assertions — an identity claim, a backstory claim, a foreshadowing claim — that can succeed or fail independently. The steelman is broken into two to five falsifiable sub-claims, each one testable against the page, with one markedcentral: the claim the theory cannot survive losing. Sub-claims are what make amixed verdict honest — the ledger shows exactly which parts landed and which died.

Step 3 — Refute, rebut, judge

  1. Refute. Two independent skeptics attack the steelman, refute-by-default: their job is to kill the theory using on-page citations only. Hearsay inside the story counts as hearsay, not fact.
  2. Rebut. An advocate answers the attacks honestly — permitted only to narrow the claim, never to widen it to dodge a hit. A concession is recorded as a concession.
  3. Judge. A three-judge panel scores what survives on four axes, and the scores aremedian-aggregated so no single judge can drag a verdict:
    • Coverage — the fraction of the relevant evidence the theory explains without strain.
    • Contradictions — the count of direct on-page beats the theory must explain away.
    • Parsimony — the inverse of how much machinery the theory invents beyond what is on-page.
    • Forward-robustness — the odds the theory survives the plausible space of future chapters.

Sub-claim labels and the overall verdict are then derived mechanically from the score medians — the rubric below is applied as arithmetic, not as a vibe check. Judges score; they don't get to pick the word.

The verdict rubric

VerdictThreshold
confirmedcomposite ≥ 0.85 and zero surviving fatal attacks
plausiblecomposite ≥ 0.60
strainedcomposite 0.35–0.60, or the case depends heavily on chapters not yet written
refutedcomposite < 0.35, or a standing kill-shot contradiction on the page
mixedthe sub-claims split — the central claim and its satellites earn different labels

Locked verdicts, living scores

Every verdict is locked at publication with the chapter frontier it was judged against. It is never silently edited. When new chapters land, the theory's falsifiers — the published "what would change our mind" list — are re-checked, and any movement is appended as a dated addendum with a current-verdict marker. The original judgment stays visible either way: a court that quietly rewrites its old rulings has no track record.

Theories that survive as live bets also spawn entries on the prediction scoreboard, where they are scored against new chapters alongside everything else the archive has ever called — hits and misses both stay up.

Where the method comes from

This is the production version of the tournament machinery built for the archive's biggest bet — thelocked answer to "what is the One Piece?". The four axes, the refute-first posture, the median-aggregated judge panels, and the narrow-only rebuttal rule are all inherited directly from that process;its methodology page shows them running at full scale, including the complete scorecard of how seven competing frames were narrowed to one thesis.

← Back to the gauntlet · The prediction scoreboard →

Message in a bottle

Found a bug, want a feature, or just want to say something? Your current page is attached automatically.

How bad is it? (optional)