# Registered Predictions — v2
**Issued 2026-08-22, after four independent reviews of v1. v1 is preserved verbatim in `PREDICTIONS_registered_2026-08-22.md` and is not edited. This document supersedes it for scoring.**

---

## Why there is a v2

v1 was sent to four independent reviewers — a hostile skeptic, a commissioning editor, an applied statistician, and a plain-reader/accessibility auditor. They graded it **B, B−, C+ and B−/C+**. They found defects. Some were serious. This document fixes them and **states each one**, because a register that quietly improved itself would be worth nothing.

**Both versions are hashed and published.** Nothing is deleted, softened, or reinterpreted. The v1 predictions stand on the record as originally written, with their defects named beside them.

### The defects, named

| # | Found by | Defect | Severity |
|---|---|---|---|
| 1 | Skeptic + statistician, independently | **F2 was a logical paradox.** It was scored against a mean that included itself. At 10 hits among the other fifteen — the modal outcome, P = 0.215 — hitting implied missing and missing implied hitting. | **Fatal** |
| 2 | Both | **The scoring rule rewarded sandbagging.** Shading every stated confidence down ten points raised the "well calibrated" pass rate from 0.484 to 0.857 with no change in belief. | **Fatal** |
| 3 | Both | **Four falsifiers were looser than the claims they scored**, always in my favour: A1 dropped its incumbents-only clause; A4's band and floor didn't match; B1 was silently directional; B2 claimed "exactly zero" while scoring a rate. | **Serious** |
| 4 | Statistician | **Zero-inflation.** With ~56 challengers at tenure 0, an overall r of ~0.75 appears with a true within-incumbent gradient of exactly zero. The headline figure was substantially a two-group contrast. | **Serious — substantive** |
| 5 | Statistician | **±0.15 in r-units is asymmetric** — 2.7× more permissive upward than downward — and the band was 4 to 11 sampling SDs wide, so it could not fail on noise. | **Serious** |
| 6 | Statistician | **C2 had no baseline**, making the ratio uninterpretable. | Moderate |
| 7 | Skeptic | **Military service and legal record — the two most politically explosive variables in the index — carried zero predictions.** | **Serious** |
| 8 | Skeptic | **No external timestamp.** The page argued pre-hoc and post-hoc are indistinguishable from outside, then offered no way to distinguish itself. | **Serious** (fixed in Phase 2) |
| 9 | Skeptic + statistician | **Confidences were systematically below true odds** on the easy predictions, in a pattern that would manufacture a flattering scorecard. | Moderate |

### What changed structurally

- **F2 is no longer a prediction.** It is a commitment (below). A commitment cannot be self-referential.
- **Calibration is now scored by Brier score**, which is proper — shading down costs you.
- **Every falsifier is the exact negation of its claim.** No limbo zones.
- **Incumbents-only is the primary specification** for the money predictions.
- **All correlation bands are stated in Fisher z**, not r.
- **Nine predictions are re-priced**, most of them upward, in direct response to review 9. Re-pricing after a critique of my pricing is itself a test: if the new numbers are still too low, F2′ should catch it.

---

## A · The tenure–money relationship

**A1a — PRIMARY.** Among incumbents only, the Pearson correlation between years of congressional service and corporate/trade-PAC share of receipts is **≥ 0.60** at the October FEC refresh.
*Fails if:* incumbents-only r < 0.60.
*Basis:* 0.70 on receipts through 2026-06-30.
**Confidence: 92%** *(was folded into an 80% conjunction; re-priced up — a quarterly increment barely moves a cumulative correlation)*
*Why primary:* the pooled figure is inflated by the challenger/incumbent split. This is the number that speaks to the claim.

**A1b — SECONDARY, and labelled as containing a two-group contrast.** Pooled across all candidates, r falls in **[0.70, 0.90]** and the sign is positive within both parties.
*Fails if:* pooled r outside [0.70, 0.90], **or** the sign flips in either party. *(All conjuncts now scored.)*
**Confidence: 88%**

**A2.** Extended to all 435 districts, the incumbents-only correlation lies **within 0.20 Fisher-z units** of the battleground incumbents-only figure.
*Fails if:* |Δz| > 0.20.
**Confidence: 65%** *(z-units, not r-units — the old ±0.15 band was asymmetric and 4–11 SDs wide)*

**A3.** Among incumbents, controlling for committee chair / ranking-member status, tenure's partial correlation with corporate-PAC share stays **≥ 0.50**.
*Fails if:* partial r < 0.50.
**Confidence: 65%**
*Declared limitation:* one binary control is not an identification strategy. Majority-party status, committee assignment, seat safety and total receipts are uncontrolled. This tests whether gavels alone explain it, nothing more.

**A4.** Rebuilt on the 2024 cycle, the incumbents-only correlation falls in **[0.50, 0.85]**.
*Fails if:* outside [0.50, 0.85]. *(Band and falsifier now match.)*
**Confidence: 70%**

## B · How politicians write

**B1.** Across the 59-race sweep, candidates with a professional media background do not differ from those without on social-register reading grade — **and the difference is smaller than 1.0 reading grades** (two one-sided tests, equivalence margin 1.0, α = 0.05).
*Fails if:* a two-sided difference at p < 0.05 **in either direction**, or equivalence is not established at the stated margin.
**Confidence: 80%**
*Fixed:* v1 accepted the null from a failure to reject, and its falsifier fired in only one direction. Both corrected.

**B2.** Among the **first 100 written attributed quotations** collected in the sweep, **not one** contains a third-person self-reference.
*Fails if:* **any single instance**, in any quotation, of any length. *(v1 scored a rate; one violation only tripped it if the quote was under 334 words.)*
**Confidence: 70%**
*Stated openly:* a black-box Beta(1,1) posterior on 23-for-23 gives only **19%** for 100 more. I am departing from it because press releases quote their principal in the first person by construction — a mechanism, not a streak. **If this misses, the mechanism argument was wrong, and that is the more interesting failure.**

**B3.** The under-55 / 55-and-over register-gap split fails to replicate — **and the difference is smaller than 1.5 reading grades** (TOST, margin 1.5).
*Fails if:* the split holds at p < 0.05 two-sided **and** survives removal of the oldest subject.
**Confidence: 65%**

**B4.** National office-holders' written-quote-to-social gap has a median **≥ 3.5 reading grades**, against a candidate median at or below 2.5.
*Fails if:* office-holder median < 3.5, or candidate median > 2.5.
**Confidence: 55%** *(n = 4 at registration; the weakest here and flagged as such)*

## C · Collection

**C1.** At least **15%** of the 115 candidates have no collectable social register under the frozen protocol.
*Fails if:* under 15%.
**Confidence: 78%**

**C2.** Collectability is partisan-skewed. **The roster baseline is 57 Democrats to 57 Republicans — exactly 1.00.** Among candidates whose social cells meet the minimum, the **odds ratio** of collectability, Democrat versus Republican, is **≥ 1.30**.
*Fails if:* OR < 1.30, or it runs the other way.
**Confidence: 60%**
*Fixed:* v1 stated a bare ratio against no baseline. The baseline is now published, and it happens to be exactly balanced — which is what makes the claim meaningful.

## D · The index

**D1.** At least two more of the 59 races materially change before Election Day.
*Fails if:* fewer than two.
**Confidence: 88%** *(re-priced up from 75%; three changed in the previous ten weeks and the threshold is two of 59)*

## E · Ages

**E1.** Between two and four of the seven unestablished ages become establishable by Election Day, under the published evidentiary standard.
*Fails if:* fewer than two, or more than four.
**Confidence: 55%**

**E2.** *Conditional.* **If** the flagged Ohio case resolves, it resolves to the older value.
*Fails if:* it resolves to the younger value. *(If it does not resolve, this is unscoreable — v1 left that outcome unhandled.)*
**Confidence: 57%**, conditional on resolution.
*Stated:* the 182-vs-139-day window assumes birth dates are uniform across the year. US births are seasonal by roughly ±5–8%, which moves this to somewhere in **53–61%**. Publishing 57% to the point without that caveat was overprecision.

**E3.** The other four uncertainty markers do not resolve.
*Fails if:* two or more resolve.
**Confidence: 70%**

## F · Scoring

**F1.** The October publication contains **at least one item explicitly labelled RETRACTED or DOWNGRADED**, referencing a claim published before 2026-09-01.
*Fails if:* no such labelled item.
**Confidence: 88%** *(re-priced from 80%; Laplace on 6-of-6 prior rounds gives 87.5%)*
*Declared conflict:* this outcome is partly under my control. It is scored on the published artifact, not on my judgment, which is the most externally checkable form available — but it is not a clean forecast and should not be read as one.

**F2 — NOT A PREDICTION. A COMMITMENT.**
In October I will publish, for the scored set: the **Brier score**, the **log-odds calibration shift â**, and **â's 95% confidence interval**. No hit-rate threshold is attached, because attaching one is what created the paradox.

**F2′ — the scored calibration claim.** The log-odds calibration shift **â comes out negative** — I am overconfident on average.
*Fails if:* â ≥ 0.
**Confidence: 60%**
*Stated in advance, because it cannot be said honestly afterwards:* with this many predictions the 95% CI on my true hit rate spans roughly **47% to 87%** — about −24 to +16 points. **This round cannot distinguish a well-calibrated forecaster from a badly overconfident one.** Brier and â are published because they are proper and decomposable, not because they are conclusive here. The instrument becomes real past about fifty predictions.

## G · The two variables v1 left untouched

*A reviewer noticed that every v1 prediction sat on campaign finance, writing style, collection, roster and ages — and that **military service and legal record, the two most politically explosive variables the index codes, carried none**. He could not prove that was deliberate. Neither can I. Closing it is better than arguing about it.*

**G1.** The difference in veteran share between the two parties does **not** reach significance when extended beyond the battleground set (Fisher exact, two-sided, p > 0.05).
*Fails if:* p ≤ 0.05.
*Basis:* 17 of 57 Republicans, 10 of 57 Democrats — **Fisher exact p = 0.186, OR = 2.00**, not significant. An earlier version of this claim was already refuted once under a frozen rubric.
**Confidence: 62%**

**G2.** Through the October refresh, the share of candidates carrying any recorded legal matter does **not** differ significantly by party (Fisher exact, two-sided, p > 0.05).
*Fails if:* p ≤ 0.05.
*Basis:* 5 of 57 Republicans, 7 of 57 Democrats — **Fisher exact p = 0.762, OR = 0.69.**
**Confidence: 80%**

---

## Scoring rules, fixed now

1. **Scored set: 19 predictions** — A1a, A1b, A2, A3, A4, B1, B2, B3, B4, C1, C2, D1, E1, E2, E3, F1, F2′, G1, G2. *(F2 is a commitment and is not scored.) Mean stated confidence: **70.95%**, computed, not estimated.*
2. **Every falsifier is the exact negation of its claim.** Any outcome not covered is a defect in this document and will be logged as one.
3. **Scoreability floor, fixed today:** the B-block is scoreable only if ≥ 90 of 115 candidates yield ≥ 5 collectable posts each. **Failing that floor counts as a MISS, not a deletion** — otherwise attrition selects the denominator, and attrition is exactly what C1 and C2 predict.
4. **Unscoreable** applies only to E2's non-resolution, which is specified above.
5. **Scoring date: 2026-10-31**, not "mid-October".
6. **Both v1 and v2 are hashed and published**, along with hashes of the two predictions withheld from public view, so the count verifies without disclosure.

*v2 issued 2026-08-22. Corrections go to v3 with a changelog; nothing here is edited in place.*
