# Registered Predictions — v5

**Issued 2026-08-24, before the October data exists. v1 through v4 are preserved verbatim
and are not edited. This document supersedes v4 for scoring. Scored set: 25 predictions,
membership unchanged since v3. Scoring date 2026-10-31.**

Every figure in this document is computed by `scripts/figures.py` from the frozen data at
build time. None is typed.

---

## Why there is a v5

Ten independent reviewers examined the v4 build. Their findings, plus one external fact
(a candidate's party switch), require corrections to **facts, instruments and rules — not
to any bet**. As always: a false claim about data in hand is an error and gets fixed; a
registered forecast stands as written whatever it costs. No prediction's claim, threshold
or confidence changes in v5.

## Ruling 1 — the universe, and dual-universe scoring registered before October

`ANALYSIS_SPEC_v1.md` §0 prints the selection rule: 2024 House margin ≤ 10 **or**
presidential margin ≤ 10. **Three of the 59 districts violate it** — OH-09 (pres margin
10.5), TX-28 (10.3), TX-34 (10.1). `build.py`'s own comment says marquee additions are
"flagged as such"; the flag was never published. That is a defect in the published record
and it is hereby fixed: the three are declared **named exceptions**, marked in
`districts_59.csv`.

The consequence is not small, and it is the reason this ruling must be made now, blind:

| | As published (59 districts) | Under the printed rule (56) |
|---|---|---|
| pooled Pearson r | 0.5756 (n = 112) | **0.7116** (n = 106) |
| incumbents-only r | 0.3108 (n = 48) | **0.4993** (n = 45) |

Under the strict universe, A1b — published as a near-certain miss — would score a **HIT**,
and A1a′ would score a **MISS** (|Δz| = 0.2269 against its 0.20 band).
Choosing a universe in October, after seeing that table, would be choosing the verdict.

**The ruling, registered now:** every A-block prediction is scored on the **as-published
59-district universe** — the universe the bets were priced against. The rule-compliant
figures are computed by the same script and **published beside every A-block verdict,
unscored and labelled**. Neither I nor anyone else gets to swap universes in October.

## Ruling 2 — B4's instrument has a defect, declared per §5.1

`style_analysis.py` pools each cell's samples with a bare newline join.
**72 of 468 samples end without
terminal punctuation — 66 of them social** — so
sentences merge across sample boundaries, the social cell's reading grade inflates, and
the quote-to-post register gap shrinks. In exactly the direction that makes B4's
candidate-median conjunct hold.

Computed both ways, across the 19 candidates with both cells:

| Candidate register-gap median | |
|---|---|
| shipped pipeline (the bug) | **2.38** — B4's ≤ 2.5 conjunct *holds* |
| corrected (terminal punctuation restored per sample) | **3.54** — the conjunct *fails* |

The spec froze `style_analysis.py` by name, so under §5.1 this is **a defect in the frozen
instrument, logged as such** — not a judgment call. **B4 is scored on the corrected
pipeline**, both values are published with the verdict, and on today's data that means B4
is now an expected **miss**. Its stated 55% confidence stands unchanged; the instrument
was wrong, the bet was the bet. B1 and B3 use the same pipeline and inherit the same
correction; their thresholds are unaffected by direction.

## Ruling 3 — C2's baseline, corrected for a fact in the world

Kevin Kiley (CA-06) left the Republican Party on 2026-03-09. The roster baseline is
**57 D / 56 R / 2 IND**, not the 57/57 the C2 card called "exactly 1.00 — which is what
makes the claim meaningful." The card's arithmetic is re-based to 57:56 (baseline odds
1.018, not 1.000); C2's claim, threshold (OR ≥ 1.30) and confidence are unchanged — the
threshold clears the corrected baseline by the same margin that made the claim meaningful.
The roster's party column is corrected with the change logged; G1/G2, which compare D to
R, drop Kiley from both sides by construction.

## Correction 4 — the D2 promise failed, and the claim comes down, not the record

The page promised the two withheld predictions were published "as hashes, so the count
verifies without disclosing the content." **That was false for both**: the full texts of
D2 and E4 are printed in `PREDICTIONS_registered_2026-08-22.md` — a file the page links,
hashes and tells readers to download — and D2 concerns whether a named candidate actively
contests his race. The fact-check gate hash-verified both printed texts against the
published withheld-hashes, confirming the disclosure. v1 is preserved verbatim; scrubbing
it would be worse than the failure. **What changes is the claim:** the page now states
that both texts were published in v1 in error, that the withholding failed, and that the
failure is logged. D2 remains scored privately — privately meaning the verdict, not the
text, which is already public by my own error. E4 remains retired.

## Correction 5 — the story's central uncertainty was false

The correction story claimed the cause of the 0.82 collapse could not be determined
because the raw FEC files were gone. **The question was fully answerable from the
published record.** Restricting today's frozen CSV to the batches 1–8 coding universe
(n = 92, 38 incumbents under the FEC incumbency
code August used) reproduces every August correlation:

| | Published in August | Batches 1–8, today |
|---|---|---|
| pooled Pearson | 0.82 | **0.8231** |
| pooled Spearman | 0.86 | **0.8587** |
| incumbents Pearson | 0.70 | **0.6963** |
| incumbents Spearman | 0.60 | **0.6004** |
| three longest removed | 0.54 | **0.5412** |
| within-party incumbents | 0.81 D / 0.63 R | **0.8095 D / 0.6263 R** |

Seven for seven. The mechanism: **the district list never expanded — the coding roster
completed.** The Pearson/Spearman mix-up hypothesis the story offered is **retracted by
name**: both statistics were recorded separately at the time and both reproduce. The
rewritten story carries this as its ending, together with the fact that I asserted
uncertainty I could have resolved from my own records.

*(One vintage caveat, found by review and kept visible: the means August published beside
those correlations — 13.2% D-incumbent, 12.6% R — reproduce on the FULL current data
under the FEC incumbency code (13.21% /
12.58%), not on the 92-row subset. The August record
mixed two vintages inside one sentence; the correlations were batch-1–8, the means were
not.)*

## Also in this issue

- **Tijerina erratum:** `EVIDENTIARY_STANDARD_v1.1.md` still says Tijerina's age cell is
  blank and "those blanks are correct." The roster says 52, exact DOB, properly
  superseded in the Source Log. The published rulebook went stale; an erratum notice
  accompanies it. The rulebook text itself is not edited.
- **Exculpatory qualifiers restored:** `roster_115.csv` reduced every legal matter to one
  or two words, stripping the qualifiers the evidentiary standard requires to travel with
  a charge. The CSV now carries the full legal detail and service-record columns the
  master coding always had.
- **Licence:** code MIT, data and text CC BY 4.0, FEC source data public domain. Both
  pages and the bundle now say so; "all rights reserved" is gone.

## Scoring rules, v5

Rules 1–8 of v4 carry forward unchanged, plus:

9. **Universe ruling (Ruling 1):** A-block verdicts are scored on the as-published
   59-district universe; rule-compliant figures are published beside them, unscored.
10. **Instrument ruling (Ruling 2):** the B-block is scored on the corrected pooling
    pipeline; shipped-pipeline values are published beside every B-block verdict.
11. **Baseline ruling (Ruling 3):** C2 is scored against the corrected 57:56 baseline.
12. Version chain, all hashed and published: v1 (18 registered / 16 published) → v2 (19
    scored) → v3 (25 scored) → v4 (defect corrections) → **v5 (this document, binding)**.

*v5 issued 2026-08-24. Corrections go to v6 with a changelog; nothing here is edited in place.*
