# Registered Predictions — v4

**Issued 2026-08-23, after a second round of four independent blind reviews of v3. v1, v2 and v3
are preserved verbatim and are not edited. This document supersedes v3 for scoring. Scored set:
25 predictions. Scoring date 2026-10-31.**

---

## Why there is a v4

v3 was sent to four independent reviewers — a hostile skeptic, a commissioning editor, an applied
statistician, and a plain-reader/accessibility auditor. None was shown an earlier draft or told
what the others said. They graded it **B−, B−, B− and B/C+**. The target was S.

They found six fatal defects. I recomputed every one before accepting it. **Three of the six are
the same failure this project exists to name** — a number that lives in prose, carried across a
change in its inputs. The 0.82 was the first. The two below are the second and third.

### The defects, named

| # | Found by | Defect | Severity |
|---|---|---|---|
| 1 | Statistician | **The confidence interval on the page was a stale v2 number.** It said "roughly 47% to 87%", computed for a nineteen-prediction set and carried unchanged onto a twenty-five-prediction page. | **Fatal — and the exact failure the page condemns** |
| 2 | Skeptic + statistician, independently | **The F2 self-reference paradox was back.** F2′ and F2″ are the same proposition, both sat inside the set â was fitted over, so â was a fixed-point equation. 6.29% of outcome vectors had no self-consistent assignment. | **Fatal** |
| 3 | Statistician | **A5 — the replacement headline — is a magnitude claim tested with a correlation, and the magnitudes point the other way.** | **Fatal** |
| 4 | Skeptic | **`make_pages.py` hardcodes every published number as a string literal.** A5's own basis figure is a prose number, written into the file on the day the rule forbidding prose numbers was written. | **Fatal** |
| 5 | Skeptic | **"Check it yourself" was false.** The bundle omitted the finance CSV and the tenure table, so all three published scripts failed from it. Editing one unpublished input moved the primary statistic from 0.311 to 0.525 while `sha256sum -c` still returned 15/15 OK. | **Fatal** |
| 6 | Statistician | **The frozen analysis specification binds v2**, so six of the twenty-five scored predictions — including the replacement headline — had no population, estimator, exclusion rule or scoring map anywhere. 24% of the register unspecified. | **Fatal** |
| 7 | Skeptic | **G1's published basis violates the specification's own population rule.** The spec says "Unclear" is excluded; the basis counted the three Unclear candidates as non-veterans. | Serious |
| 8 | Statistician | **A1a′'s printed band contradicts its own binding criterion at both edges.** | Serious |
| 9 | Skeptic | **Scoring rule 3's arithmetic is backwards.** | Serious |
| 10 | Skeptic | **B3's scoring map omits a conjunct its published claim carries.** | Serious |
| 11 | Reader/a11y | **`a11y_audit.py` — the check on the check — measures nothing that is text**, and printed ALL CHECKS PASS on a tree with three AA failures, one of which Phase 4 introduced. | Serious |
| 12 | Editor | **The two pages do not reconcile in front of the reader**: 19 vs 25 predictions, 56 vs 64 challengers, a roster of 115 against charts plotting 112. | Serious |

*(Defects 4, 5, 6, 11 and 12 are fixed in the build and the bundle, not in this document. This
document covers 1, 2, 3, 7, 8, 9 and 10.)*

---

## Defect 3 — A5, and a rule that stopped me deleting it

A5 said the tenure–money association is "mostly a step, not a slope." Recomputed:

| | |
|---|---|
| Step (incumbent mean − non-incumbent mean) | **11.89 pp** |
| Slope | 0.348 pp/yr, se 0.157, 95% CI [0.040, 0.656], R² = 0.097 |
| Climb across the observed range, 1.8 to 43.8 years | **14.62 pp** |
| Residual scatter about the line | **8.64 pp** |

**The slope moves more than the step.** r = 0.311 is low because the scatter is 8.64 points, not
because the gradient is small. I replaced a wrong headline with a differently wrong one, and the
error is from the same family: reaching for a correlation to make a claim about size.

It is also not an independent test. In 20,000 bootstrap resamples, **A1a′ holding implies A5
holding in 20,000 of 20,000 cases**; P(A5) = **0.9996** against a stated 88%.

**My instinct was to withdraw it. My own frozen rule stopped me.** `ANALYSIS_SPEC_v1.md` §5.5:
*"No prediction is withdrawn after 2026-08-22. The only permitted post-registration status is HIT,
MISS, or the single specified UNSCOREABLE branch in E2."* A5 was registered on 23 August. Deleting
it because I have since decided it was a bad bet is precisely what that rule exists to forbid, and
a rule that binds only until it is inconvenient is not a rule.

So **A5 stays in the register and is scored at 88%.** What is withdrawn is its *status*:

- It is **no longer the headline** and is no longer described as a finding.
- **No replacement headline is registered.** I do not currently know whether the right
  characterisation of the A-block is a step or a slope, and the honest position is that the
  question is open. It will not be answered by picking a third framing three days after the first
  two failed.
- **A hit on A5 is not evidence for the claim.** At P = 0.9996 it is near-certain regardless, and it
  should be read as evidence about the weakness of the test I built, which is what it measures.

## Defect 1 — the confidence interval

The page said the 95% interval on my true hit rate spans "roughly 47% to 87%." Provenance is
provable: v2's mean was 70.95% and v2 stated "−24 to +16 points." 70.95 − 24 = 46.95;
70.95 + 16 = 86.95. **A statistic computed for nineteen predictions, printed on a page of
twenty-five.**

Exact Poisson-binomial over the actual 25 confidences:

| | |
|---|---|
| 95% CI on the true hit rate | **56.0% – 88.0%** (14–22 hits) |
| Given the four predictions already known to be misses | **44.0% – 76.0%** (11–19 hits) |
| **Maximum attainable hit rate** | **84.0%** |
| P(hit rate ≤ 47%) under the page's own numbers | **0.0009** |

The old figure was not merely stale, it was unreachable in both directions.

## Defect 2 — the paradox, closed at the root this time

F2′ and F2″ assert the same proposition — that the log-odds calibration shift **â** comes out
negative — and both sat inside the set â was fitted over. Each one's outcome therefore determined
the estimator that scored it. In 20,000 simulated outcome vectors with the four known misses
forced, **6.29% had no self-consistent assignment**: assume hit and â came out positive, assume
miss and â came out negative.

**Fix, fixed now: â is estimated on the 23 predictions excluding F2′ and F2″, and scored against
those two.** The estimator no longer depends on the outcomes it scores, so a fixed point cannot
exist by construction. Re-simulated on the same 20,000 draws, varying F2′ and F2″ both ways each time and checking
whether the estimator moves: **0 inconsistent vectors, 0.00%.**

Under the register's own confidences this gives **P(â < 0) = 0.9356**, which is what F2″'s stated
93% should have been priced at and, by luck rather than judgment, almost exactly is. F2′ remains
frozen at 60% and remains a substantial freebie; that was declared when it was frozen and it
stands.

---

## The scored set — 25 predictions, mean stated confidence 73.80%

**Carried from v3 unchanged except where a correction is named below.** Full text of the nineteen
v2 predictions is in `PREDICTIONS_v2_2026-08-22.md`; full text of the six added on 23 August is in
`PREDICTIONS_v3_2026-08-23.md`. Nothing in either is edited.

### Corrections issued in v4

**A1a′ — printed band corrected.** The binding criterion is and always was
**|z − 0.3214| ≤ 0.20**. The card printed "r in [0.12, 0.48]" for readability; both rounded edges
fall *outside* the criterion (r = 0.12 → z = 0.12058 → |Δz| = 0.20082; r = 0.48 → z = 0.52298 →
|Δz| = 0.20158), and under §0's two-decimal rounding rule those are exactly the values most likely
to be hit. **The z form governs. The r edges are 0.1208 and 0.4788** and are printed to four
decimals or not at all. Confidence unchanged at 85%.

**B3 — scoring map completed.** The published claim carries a TOST conjunct ("*and* the difference
is smaller than 1.5 reading grades") that the scoring map omitted, so the most likely outcome for
an underpowered test — inconclusive — made the claim false while scoring HIT. **Scoring map: HIT
iff the split fails at p ≥ 0.05 two-sided AND equivalence is established at margin 1.5. Any other
outcome, inconclusive included, is a MISS.** Confidence unchanged at 65%.

**G1 — basis figure corrected.** `ANALYSIS_SPEC_v1.md` §G1 specifies that "Unclear" veteran coding
is excluded from the population. The published basis counted the three Unclear candidates —
Gregory Cunningham (R), Eric Flores (R), Kristina Knickerbocker (D) — in the denominator as
non-veterans, which is the opposite of the stated rule.

| | R | D | OR | p |
|---|---|---|---|---|
| As published (Unclear counted as non-veteran) | 17 / 57 | 10 / 57 | 1.998 | 0.1857 |
| **Under the spec's own rule (Unclear excluded)** | **17 / 55** | **10 / 56** | **2.058** | **0.1258** |

The claim and the threshold are unchanged and the hit/miss verdict is unchanged — both p values
clear 0.05. **The basis figure was wrong, and a basis figure being wrong is what started all of
this.** Confidence unchanged at 62%. *(G2 reproduces exactly: 5/57 R, 7/57 D, OR = 0.687,
p = 0.7616.)*

**The challenger figures — three groupings, one rule.** Three different numbers were in
circulation because three different populations were being described:

| Grouping | n | Mean corporate/trade share |
|---|---|---|
| FEC incumbency code **C** (challenger) | 50 | **0.2992%** |
| FEC incumbency code **O** (open seat) | 16 | **1.2756%** |
| **Tenure = 0** — the definition the spec actually binds | **64** | **0.5123%** |

The August figures of "0.3%" and "1.3%" were **correct** for the FEC-code groupings. The figure of
**0.26% printed in `DEFECT_A_BLOCK_2026-08-23.md` is wrong** — it came from intersecting tenure = 0
with FEC code C, a fourth population nothing specifies. **Correction: my own defect report
contained an unreproducible number, inside the table asserting that everything reproduced.** The
spec-binding figure is **0.5123% across 64 candidates at tenure 0**, and that is the only one that
appears in publication from here.

---

## Scoring rules, v4

1. **Scored set: 25 predictions**, unchanged in membership from v3 — A1a, A1b, A2, A3, A4, A1a′,
   A1b′, A5, A3′, A4′, B1, B2, B3, B4, C1, C2, D1, E1, E2, E3, F1, F2′, F2″, G1, G2. Mean stated
   confidence **73.80%**, computed. *(F2 remains a commitment and is not scored.)*
2. **A1a, A1b, A3 and A4 are known misses and are scored as such** in numerator and denominator
   both. A5 is scored despite its characterisation being withdrawn, per §5.5.
3. **Correction to v3's scoring rule 3, which was arithmetically backwards.** v3 claimed that
   registering the corrected bets beside the frozen ones "penalises the original error twice." It
   does not. Adding the six *lowers* the expected Brier score, and the mean stated confidence
   *rose* from 70.95% to 73.80% rather than being diluted. **The real penalty is smaller and
   specific: freezing the four dead bets rather than deleting them costs **0.0871 Brier** (0.1759 frozen against 0.0887 deleted).
   That penalty is genuine. The claim that it was doubled was not, and overstating a self-imposed
   cost is a flattering error, which makes it the kind worth naming.
4. **â is estimated on the 23 predictions excluding F2′ and F2″**, and those two are scored against
   it. This is what closes the paradox and it is not revisable.
5. **Scoreability floor unchanged:** the B-block is scoreable only if ≥ 90 of 115 candidates yield
   ≥ 5 collectable posts each; failing the floor scores every B prediction MISS, not unscoreable.
6. **Unscoreable** applies only to E2's non-resolution.
7. **Scoring date 2026-10-31.** All four versions hashed and published, with the two withheld
   prediction hashes, so the count verifies without disclosure.
8. **Every published figure must come from a script in the repository.** Extended in v4 from the
   A-block to *every* block, because the second and third recurrences of the prose-number failure
   were both outside the A-block.

*v4 issued 2026-08-23. Corrections go to v5 with a changelog; nothing here is edited in place.*

---

## Provenance of every figure in this document

All of the above is produced by `scripts/figures.py`, which computes each value from
`finance_nominees.csv`, `AGE_TENURE` and `PHASE2` and writes `web/figures.json`. The published
pages interpolate from the same dict; nothing is typed twice.

This is the correction to defect 4. v3's rule 5 required A-block figures to come from a script,
and `make_pages.py` then held every published number as a string literal — including A5's own
basis figure, typed in on the day the rule was written. Rule 8 in v4 extends the requirement to
every block, because the second and third recurrences of the prose-number failure were both
outside the A-block.

**Two figures in the first draft of this document were themselves wrong and are corrected above
by the script:** the Brier cost of freezing the four dead bets, which I had taken from a
reviewer at “approximately 0.07” and which computes to **0.0871**, and the share of paradoxical outcome vectors under
v3, which I had hand-run at 6.16% and which computes to **6.29%**.
Neither changes a conclusion. Both are the same mistake at smaller scale, made while writing the
document that names it, which is roughly how often it happens.
