# Registered Predictions — v3

**Issued 2026-08-23. v1 (`PREDICTIONS_registered_2026-08-22.md`) and v2
(`PREDICTIONS_v2_2026-08-22.md`) are preserved verbatim and are not edited. This document
supersedes v2 for scoring. All twenty-five predictions below are scored on 2026-10-31.**

---

## Why there is a v3, one day after v2

v2 existed because four independent reviewers found defects in v1. v3 exists because of
something worse, and I found it myself, by doing what the reviewers told me to do.

Phase 3 of the correction plan was editorial: put the money finding at the top of the page as a
plain fact instead of opening with epistemology. To write that sentence I needed the number. So I
computed it from the frozen artifacts, under the frozen specification, for the first time.

**It did not reproduce.**

| Quantity | Published in v1 and v2 | Reproduced from the frozen data |
|---|---|---|
| Pooled Pearson r | **0.82** | **0.576** |
| Incumbents-only Pearson r | **0.70** | **0.311** |
| n pooled / n incumbents | 92 / 38 | 112 / 48 |

Full verification, including everything that *did* reproduce and two candidate explanations for
what went wrong, is in `DEFECT_A_BLOCK_2026-08-23.md`. The short version: **the figure was never
in a saved script.** It was computed once in an ad-hoc session, written into prose, and cited from
prose thereafter — so when the district universe grew from 92 candidates to 112 it could not be
re-run, and nobody noticed it had gone stale. This is the identical failure mode as the ALL-CAPS
register feature, caught the same way ten days ago. It is now fixed at the root:
`scripts/money_analysis.py` implements the A-block to specification, and no A-block number gets
published again from prose.

### What this costs

Four of the nineteen v2 predictions were priced against a number that was wrong:

- **A1a** — *incumbents-only r ≥ 0.60*, registered at **92%**. June gives **0.311**.
- **A1b** — *pooled r in [0.70, 0.90]*, registered at **88%**. June gives **0.576**.
- **A3** — *partial r ≥ 0.50 among incumbents*, registered at **65%**. A partial correlation can
  exceed a zero-order correlation of 0.311 only under suppression, which is possible but not
  likely.
- **A4** — *2024-cycle incumbents-only r in [0.50, 0.85]*, registered at **70%**. Largely the same
  members, so the 2024 figure is unlikely to sit two-and-a-half times above the 2026 one.

All four are near-certain misses. **They are not being changed.** Under the rule fixed for this
project, a registered prediction is not edited after its basis is found wanting — that is the
precise move this whole exercise exists to make impossible. A1a, A1b, A3 and A4 stand exactly as
written, at 92%, 88%, 65% and 70%, and they will take their misses in October beside the reason.

What gets corrected is the **fact**, not the bet. The 0.82 was never a forecast; it was a claim
about data already collected, and a false claim about data is an error to be fixed, not a
position to be defended.

Then, separately, the corrected bets are registered today, against the number that actually
reproduces, and scored as their own entries. **Both the wrong bet and the corrected bet are on
the record, side by side, for the same quantity.**

### What this does to the calibration claim

Knowing that four predictions are baked-in misses makes **F2′** — *"â comes out negative, I am
overconfident on average"* — close to a certainty. It was registered at **60%** before I knew
this. Leaving it there would hand me a cheap hit on a question I have already resolved, which is
the quiet kind of advantage this project exists to remove. So F2′ is frozen at 60% like the rest,
and **F2″** below states the honest price.

### What the reviewer got more right than I credited

The statistician's zero-inflation charge — that a pooled r near 0.75 can arise from the
challenger/incumbent split with a true within-incumbent gradient of zero — is no longer a
simulation. It is what the data does:

- 64 non-incumbents at tenure 0, mean corporate/trade share **0.51%**
- 48 incumbents, mean corporate/trade share **12.40%**
- pooled r **0.576**, within-incumbent r **0.311**

**The tenure–money relationship is mostly a step, not a slope.** Getting into Congress is what
changes where the money comes from. Staying longer changes it much less than I published. That is
a smaller finding than the one I had, and it is the one that survives.

---

## FROZEN — carried from v2 unchanged, defects and all

These nineteen are exactly as registered on 2026-08-22 (v2) and are reproduced here without a
character changed. **A1a, A1b, A3 and A4 are known to be mispriced and are scored anyway.**

| ID | Claim, in brief | Confidence | Status |
|---|---|---|---|
| A1a | Incumbents-only r ≥ 0.60 | 92% | **known bad — will miss** |
| A1b | Pooled r in [0.70, 0.90], both party signs positive | 88% | **known bad — will miss** |
| A2 | All-435 incumbents within 0.20 Fisher-z of battleground incumbents | 65% | unaffected |
| A3 | Partial r ≥ 0.50 controlling for chair/ranking member | 65% | **known bad — will miss** |
| A4 | 2024-cycle incumbents-only r in [0.50, 0.85] | 70% | **known bad — will miss** |
| B1 | Media background does not predict social reading grade (TOST, margin 1.0) | 80% | unaffected |
| B2 | Zero third-person self-reference in the first 100 written quotations | 70% | unaffected |
| B3 | The under-55/55+ register-gap split fails to replicate (TOST, margin 1.5) | 65% | unaffected |
| B4 | Office-holder quote-to-social gap ≥ 3.5 grades vs candidate median ≤ 2.5 | 55% | unaffected |
| C1 | ≥ 15% of 115 candidates have no collectable social register | 78% | unaffected |
| C2 | Collectability odds ratio D:R ≥ 1.30 against a 57:57 baseline | 60% | unaffected |
| D1 | ≥ 2 more of the 59 races materially change | 88% | unaffected |
| E1 | Between 2 and 4 of the 7 unestablished ages become establishable | 55% | unaffected |
| E2 | The flagged Ohio case resolves to the older value (conditional) | 57% | unaffected |
| E3 | The other four uncertainty markers do not resolve | 70% | unaffected |
| F1 | ≥ 1 item labelled RETRACTED or DOWNGRADED in the October publication | 88% | unaffected |
| F2′ | Log-odds calibration shift â comes out negative | 60% | **underpriced — see F2″** |
| G1 | Veteran share by party stays non-significant beyond the battleground set | 62% | unaffected |
| G2 | Legal-matter share by party stays non-significant | 80% | unaffected |

*(F2 remains a commitment, not a prediction, and is not scored. Full text of all nineteen is in
`PREDICTIONS_v2_2026-08-22.md`, SHA-256 `d0cebea416e4a0677468a5113a10c378777a8f8ebeaa15266c384da50e81c065`.)*

---

## NEW — registered 2026-08-23 against the reproduced baseline

Every figure below comes from `scripts/money_analysis.py`, run on
`scripts/finance_nominees.csv` and `AGE_TENURE`, vintage: receipts through 2026-06-30. All bands
are stated in **Fisher z**, per the v2 rule, and converted to r for readability.

**A1a′ — PRIMARY, corrected.** Among incumbents (tenure > 0) passing the viability filter, the
Pearson correlation between tenure and corporate/trade-PAC share of receipts stays **within 0.20
Fisher-z units of z = 0.3214**, i.e. **r in [0.12, 0.48]**, at the October refresh.
*Fails if:* |z_Oct − 0.3214| > 0.20.
*Basis:* r = 0.3108, z = 0.3214, n = 48, p = 0.031, bootstrap 95% CI [0.058, 0.683].
**Confidence: 85%**
*Declared, so it cannot be claimed as a triumph later:* this is an easy prediction and I am
pricing it as one. The band is 1.34 sampling SDs wide, but October is not an independent draw —
it is the same 48 people with one more quarter of receipts, so a cumulative correlation barely
moves. **This tests continuity, not the finding.** The finding is A5.

**A1b′ — SECONDARY, corrected.** Pooled across all candidates passing the filter, r stays
**within 0.20 Fisher-z of z = 0.6559**, i.e. **r in [0.43, 0.69]**, **and** the sign stays
positive within both parties.
*Fails if:* |z_Oct − 0.6559| > 0.20, **or** either party's r ≤ 0. *(All conjuncts scored.)*
*Basis:* pooled r = 0.5756, n = 112. Within party: DEM r = 0.5183 (n = 56), REP r = 0.6830
(n = 55).
**Confidence: 88%**
*Labelled, permanently:* this number contains the challenger/incumbent contrast and is not a
gradient. It is published as context for A5, not as a finding in its own right.

**A5 — the corrected substantive claim, and the one that matters.** The tenure–money association
is **mostly a step, not a slope**: at the October refresh, the pooled Fisher z exceeds the
incumbents-only Fisher z by **at least 0.20**.
*Fails if:* z_pooled − z_incumbents < 0.20.
*Basis:* 0.6559 − 0.3214 = **0.3344**. The mechanism is visible in the means — non-incumbents
0.51% corporate/trade share, incumbents 12.40%.
**Confidence: 88%**
*This is the replacement headline.* It is a weaker claim than "corporate money rises smoothly
with seniority," and it is the one the data supports. It also directly concedes the reviewer's
zero-inflation charge rather than arguing with it.

**A3′ — corrected.** Among incumbents, controlling for committee chair / ranking-member status,
tenure's partial correlation with corporate/trade share falls in **[0.10, 0.45]**.
*Fails if:* outside [0.10, 0.45].
*Basis:* the zero-order incumbents-only r is 0.311. A partial correlation exceeds its zero-order
correlation only under suppression; absent that, the partial should land at or slightly below
0.311.
**Confidence: 78%**
*Declared limitation, carried forward from A3 unchanged:* one binary control is not an
identification strategy. Majority-party status, committee assignment, seat safety, district
industry mix and total receipts are uncontrolled. No causal claim is made or will be made.

**A4′ — corrected.** Rebuilt on the 2024 cycle, the incumbents-only correlation falls in
**[0.10, 0.55]**.
*Fails if:* outside [0.10, 0.55].
**Confidence: 65%**
*Declared:* largely the same members appear in both cycles, so this is not an independent
replication and will not be described as one. The band is wide because the 2024 district set is
different and I genuinely do not know; pretending to a tighter band would be the same sin twice.

**F2″ — the honest price on calibration.** The log-odds calibration shift **â comes out negative**
— I am overconfident on average.
*Fails if:* â ≥ 0.
**Confidence: 93%**
*Stated plainly:* this is the same claim as F2′, repriced from 60% to 93% for one reason — four
predictions in the frozen set are known misses before the data arrives, so overconfidence is
close to arithmetically guaranteed. **F2″ is not a forecast, it is a consequence, and it should
be read as one.** It is registered rather than withheld because the alternative was to let F2′
collect a cheap hit at 60% while I knew better. Both are scored. The gap between them, 33 points,
is the size of the advantage I am declining to take.

---

## Scoring rules

1. **Scored set: 25 predictions** — the nineteen frozen (A1a, A1b, A2, A3, A4, B1, B2, B3, B4,
   C1, C2, D1, E1, E2, E3, F1, F2′, G1, G2) plus the six registered today (A1a′, A1b′, A5, A3′,
   A4′, F2″). Mean stated confidence: **73.80%**, computed, not estimated.
2. **The four known-bad predictions are scored, not excused.** A1a, A1b, A3 and A4 count as
   misses in the numerator and the denominator both. No asterisk, no separate table, no "adjusted"
   score published alongside the real one.
3. **Both a wrong bet and its corrected replacement are scored** wherever they exist. This
   penalises the original error twice — once in the miss, once in the diluted mean — which is the
   correct direction for the incentive to point.
4. Rules 2 through 6 of the v2 scoring rules carry forward unchanged: exact-negation falsifiers,
   the B-block scoreability floor (≥ 90 of 115 candidates at ≥ 5 posts, **failure counts as a
   MISS**), unscoreable applying only to E2's non-resolution, scoring date **2026-10-31**, and
   publication of all version hashes.
5. **All A-block figures come from `scripts/money_analysis.py`.** Any A-block number appearing in
   publication without a corresponding run of that script is a defect and will be logged as one.

*v3 issued 2026-08-23. Corrections go to v4 with a changelog; nothing here is edited in place.*
