# Frozen Analysis Specification v2

**Issued 2026-08-30. Binds `PREDICTIONS_v5_2026-08-24.md` and every prediction in it.
v1 is preserved verbatim and is not edited; this document supersedes it. Thresholds
below are frozen literals — they are the claim. Basis figures are measurements,
computed by `scripts/figures.py` at issue and stated "as at issue".**

---

## Why there is a v2

v1 bound the nineteen predictions of register v2. Six more were registered afterwards
— **A1a′, A1b′, A5, A3′, A4′, F2″** — and v1 covered none of them. A reviewer put the
consequence plainly: **24% of the scored register had no population, estimator,
exclusion rule, missing-data rule or scoring map anywhere**, which made the register's
own central structural claim false for itself. This document closes that, folds in the
three rulings of register v5, and fixes one ambiguity in v1 that had already produced a
wrong published figure.

**The v1 defect this fixes by name.** v1 §G1 wrote *"documented service; 'Unclear'
excluded and counted"* — a sentence that can be read two ways, and the published basis
read it the wrong way, counting the three Unclear candidates as non-veterans. §0 below
now states the rule in a form that cannot be read twice.

---

## §0 Definitions binding all predictions

Everything in v1 §0 carries forward unchanged except where restated here. Restated in
full so this document stands alone.

**Universe.** The 59 US House districts fixed as a literal table in `scripts/build.py`.
Selection rule as printed there: 2024 House margin ≤ 10 **OR** 2024 presidential margin
≤ 10, from the Cook Political Report competitive list (ratings 2026-08-13, retrieved via
270toWin 2026-08-20). **Three districts do not satisfy that rule and are named
exceptions: OH-09 (presidential margin 10.5), TX-28 (10.3), TX-34 (10.1)** — added as
marquee races, House margins not comparable after redistricting. They are flagged in
`districts_59.csv`. The list is frozen and is not revisable.

**Dual-universe rule (register v5, Ruling 1), binding on every A-block prediction.**
Every A-block verdict is scored on the **as-published 59-district universe** — the
universe the bets were priced against. The **rule-compliant universe** (the 56 districts
that satisfy the printed rule) is computed by the same script and **published beside
every A-block verdict, unscored and labelled**. As at issue the two differ materially —
pooled r 0.5756 against 0.7116, incumbents-only
0.3108 against 0.4993 — which is precisely why the
choice is fixed here, in advance, rather than in October.

**Roster.** 115 candidates in `scripts/nominees.py`. Party composition
**57 D / 56 R / 2 IND**
(register v5, Ruling 3: Kevin Kiley left the Republican Party on 2026-03-09 and filed
CA-06 as an independent). This is the baseline C2 is scored against.

**Incumbent.** Tenure > 0, i.e. any prior elected House or Senate service. As at issue
n = 48 of the 112 candidates with a finance row.
**This group is not "sitting members"** — it contains one former member seeking return
and one sitting member whose FEC record still codes her as a challenger. Publication
says "candidates who have served".

**Tenure.** Cumulative elected House + Senate service in years through 2026-11-03, in
`AGE_TENURE`. Challengers and open-seat non-members are 0.0.

**Corporate/trade PAC receipts.** FEC bulk `itpas2.txt`, summing `transaction_amt` where
transaction type ∈ {24K, 24Z, 24R, 24C, 24F} and the filing committee's bucket is
Corporate PAC or Trade assoc PAC. Memo entries (`memo_cd == "X"`) excluded. Independent
expenditures are not in this sum.

**Committee bucket**, from `cm.txt`, applied in this order, frozen:
`dsgn == "D"` → Leadership PAC (never corporate) · `cmte_tp ∈ {Y,X,Z}` → Party ·
`org_tp ∈ {C,W}` → **Corporate** · `org_tp == "T"` → **Trade** · `L` → Labour ·
`M` → Membership · `V` → Cooperative · otherwise Other.

**Corporate/trade share of receipts** = `100 × (Corporate + Trade) / total_receipts`,
`total_receipts` being `weball26.txt` column 5, cycle-to-date. Not itemised-only.

**Viability filter.** A candidate committee is excluded if receipts < $25,000 **and**
disbursements < $25,000. **Declared gap:** the frozen CSV does not carry disbursements,
so only the receipts leg is enforced. This is stated rather than hidden, and is applied
identically in October.

**Missing finance rows.** As at issue 3 of the 115 have no row in the
frozen finance file: Jennifer Balkcom, Kevin Kiley, Matt Little. One has no federal committee; two have
committees the August bulk files misfiled, documented in the master coding. **They are
excluded from every A-block computation and named in the published excluded-list.** If
exclusions ever exceed 10% of incumbents, the affected prediction scores **MISS**, not
unscoreable.

**Data vintage.** FEC bulk data frozen 2026-08-21. **77 of
112** committees report coverage through 30 June 2026;
the full range runs 16 June 2026 to 18 August 2026. Publication says
this, not "through 30 June". October uses the Q3 bulk files downloaded once, on a date
recorded in the Source Log.

**Reading grade.** Flesch-Kincaid as implemented in `scripts/style_analysis.py`, **with
the terminal-punctuation correction** (register v5, Ruling 2): each sample is terminated
with a full stop before cell pooling. The shipped pipeline joined samples with a bare
newline, and 72 of 468 samples
(66 of them social) end without terminal punctuation,
so sentences merged across sample boundaries and social-cell grade inflated. Both values
are published with every B-block verdict; **the corrected pipeline is what scores.**

**Register cells.** `release` · `quote` (written attributed) · `spoken` (excluded from
written comparisons) · `post`.

**Veteran coding, restated unambiguously.** Each candidate is coded `Yes`, `No`, or
`Unclear` per §F of the evidentiary standard. **In every party comparison, candidates
coded `Unclear` are removed from BOTH the numerator and the denominator — they are not
counted as non-veterans — and their number and names are published with the result.**
As at issue: 27 Yes, 85 No,
3 Unclear.

**Calibration fitting set.** The log-odds calibration shift **â** is estimated on the
**23 scored predictions excluding F2′ and F2″**. Those two are then scored against it.
This is what makes them predictions rather than a fixed-point equation, and it is not
revisable.

**Scoring date: 2026-10-31.**

**Rounding.** All test statistics rounded to 2 decimals **before** comparison to a
threshold, **except** where a scoring map states a Fisher-z criterion, which is applied
at full precision because rounding an r-band edge can place it outside its own rule.
A value landing exactly on a boundary counts as **satisfying** the claim.

---

## §1 The money predictions

Common to all: population as defined in §0, unit = candidate, Y = corporate/trade share
of receipts, X = tenure in years, no covariates unless stated, **no outlier removal ever,
under any justification**. The highest-leverage point (Marcy Kaptur,
44 years, 8.41%) is named in publication
and retained.

### A1a — frozen, known bad, scored anyway
- **Scoring map:** incumbents-only Pearson `r ≥ 0.60` → HIT, else MISS. No other outcome.
- **Basis at registration:** 0.70 (wrong — see the defect report). **As at issue:
  0.3108.** Registered at 92%; expected MISS.

### A1b — frozen, known bad, scored anyway
- **Scoring map:** HIT iff `0.70 ≤ r_pooled ≤ 0.90` **AND** r_DEM > 0 **AND** r_REP > 0.
- **As at issue:** pooled 0.5756; DEM 0.5183, REP
  0.6830. Registered at 88%; expected MISS.

### A1a′ — PRIMARY
- **Population:** incumbents (tenure > 0) passing the viability filter.
- **Estimator:** Pearson r. Also reported, not scored: Spearman ρ (the outcome is
  bounded, skewed, zero-inflated) and a 10,000-resample bootstrap 95% CI, seed 20260823.
- **Scoring map:** `|z_Oct − 0.3214| ≤ 0.20` → HIT, else MISS, where z is the
  Fisher transform. **The z form governs**; its r image is
  [0.1208, 0.4788], printed to four decimals because both
  two-decimal roundings fall outside the rule they state.
- **Basis as at issue:** r = 0.3108, n = 48,
  p = 0.032, bootstrap CI
  [0.058, 0.683].
- **Declared:** October is the same 48 people with one further
  quarter of receipts, not an independent draw. This tests continuity, not the finding.

### A1b′ — SECONDARY
- **Population:** all candidates passing the filter. **Labelled in publication as
  containing the challenger/incumbent contrast; not a gradient.**
- **Scoring map:** HIT iff `|z_Oct − 0.6559| ≤ 0.20` **AND** r_DEM > 0
  **AND** r_REP > 0. All three conjuncts scored.

### A5 — RETRACTED IN SUBSTANCE, STILL SCORED
- **Scoring map:** `z_pooled − z_incumbents ≥ 0.20` → HIT, else MISS.
- **Basis as at issue:** 0.6559 − 0.3214 = 0.3344.
- **Declared, and this is the point:** the claim it was written to express — that the
  association is "mostly a step, not a slope" — **is retracted**. It is a claim about
  magnitude tested with a correlation, and the magnitudes disagree: the step is
  11.89 points, the fitted line climbs 14.62 points
  across the observed range, residual scatter 8.64 points.
  A1a′ holding implies A5 holding in 20,000 of 20,000 bootstrap resamples. **A hit here
  is evidence about the weakness of the test, not about the claim**, and publication says
  so. It is scored because §5.5 forbids withdrawing a registered prediction.

### A3 — frozen, known bad · A3′ — corrected
- **Covariate, both:** binary — committee chair **or** ranking member of any standing
  committee as of the October roster, from official committee membership pages (Tier 1).
  Subcommittee chairs do not count. **The coded list is published before scoring.**
- **Estimator:** partial Pearson correlation of tenure with share, controlling for it.
- **A3 scoring map:** `partial r ≥ 0.50` → HIT, else MISS. Registered 65%; expected MISS.
- **A3′ scoring map:** `0.10 ≤ partial r ≤ 0.45` → HIT, else MISS.
- **Declared limitation, both:** one binary control is not an identification strategy.
  Majority status, committee assignment, seat safety, district industry mix and total
  receipts are uncontrolled. **No causal claim is made or will be made.**

### A2
- **Population:** incumbents in all 435 districts, same filter, same definitions.
- **Estimator:** Fisher z of the all-435 incumbents-only r, minus Fisher z of **A1a′'s
  October value on the same vintage** — referent fixed here, not the June figure and not
  the pooled figure.
- **Scoring map:** `|Δz| ≤ 0.20` → HIT, else MISS.
- **Declared:** the samples are nested; this is not a two-independent-sample comparison
  and no p-value is claimed.

### A4 — frozen, known bad · A4′ — corrected
- **Population:** incumbents in the 2024 competitive set, rebuilt by re-running
  `build.py` against `weball24 / cm24 / pas224 / cn24 / ccl24`. **The 2024 district list
  is fixed from Cook's final 2024 ratings and published before the rebuild is run.**
- **A4 scoring map:** `0.50 ≤ r ≤ 0.85` → HIT, else MISS. Registered 70%; expected MISS.
- **A4′ scoring map:** `0.10 ≤ r ≤ 0.55` → HIT, else MISS.
- **Declared:** largely the same members appear in both cycles. **This is not an
  independent replication and will not be described as one.**

---

## §2 The writing predictions

### Scoreability floor — binds B1, B2, B3, B4
Scoreable only if **≥ 90 of 115 candidates yield ≥ 5 collectable social posts each**
under the frozen protocol. **Failing the floor scores every B prediction MISS, not
unscoreable.** Attrition is what C1 and C2 predict; letting it delete predictions would
let the denominator be selected on a variable correlated with the outcome.

**All four are scored on the terminal-punctuation-corrected pipeline (§0). Both values
are published with every verdict.**

### B1
- **Population:** candidates whose written-official and social cells both meet the frozen
  minimums.
- **Groups:** "professional media background" = pre-candidacy paid employment in
  broadcast, journalism or on-air media, per §H of the evidentiary standard. **The group
  assignment is published before October.**
- **Estimator:** difference in mean social-register FK grade. Two-sided Welch t; **and**
  TOST equivalence, margin 1.0 grades, α = 0.05.
- **Scoring map:** HIT iff `p > 0.05` two-sided **AND** equivalence established at 1.0.
  MISS if `p ≤ 0.05` in either direction, **or** equivalence not established.

### B2
- **Population:** the **first 100** written attributed quotations collected, ordered by
  district code then candidate surname — **the ordering is fixed here so the sample
  cannot be chosen later.**
- **Scoring map:** **any single instance, in any quotation, of any length** → MISS. Zero
  across all 100 → HIT.
- **Declared:** a Beta(1,1) posterior on 23-for-23 gives 19% for 100 more. The stated 70%
  departs from that on the mechanistic ground that releases quote their principal in the
  first person by construction. **A miss falsifies the mechanism, not merely the number.**

### B3
- **Population:** candidates with a computable written-quote → social gap and an
  established age.
- **Estimator:** Mann-Whitney U on the under-55 vs 55-and-over split; **plus** leave-one-
  out removing the oldest subject; **plus** TOST, margin 1.5 grades.
- **Scoring map (v4 correction):** HIT iff the split fails at `p ≥ 0.05` two-sided **AND**
  equivalence is established at margin 1.5. **Every other outcome, inconclusive included,
  is a MISS.** v1's map omitted the equivalence conjunct its own claim carried, so an
  underpowered test could make the claim false while scoring a hit.

### B4
- **Estimator:** median written-quote → social FK gap, national office-holders (benchmark
  file) vs candidates.
- **Scoring map:** HIT iff office-holder median ≥ 3.5 **AND** candidate median ≤ 2.5.
  Else MISS.
- **Basis as at issue, corrected pipeline:** candidate median
  **3.54** — above the 2.5 conjunct, so **expected MISS**.
  On the shipped pipeline it read 2.38 and the conjunct
  held. **The instrument was wrong; the bet was the bet.** Confidence stays 55%.
- **Declared:** n = 4 office-holders at registration. Underpowered and labelled as such.

---

## §3 Collection

### C1
- **Estimator:** count of the 115 with no social cell meeting the frozen minimum ÷ 115.
- **Scoring map:** `≥ 15%` → HIT, `< 15%` → MISS.
- **Published with the result:** the per-candidate reason for every failure.

### C2
- **Baseline (register v5, Ruling 3):** **57 D /
  56 R**, baseline odds 1.018. The two independents are excluded
  from this test. *(v1 fixed the baseline at 57:57, "exactly 1.00"; Kiley's March party
  change makes that false, and the roster column had fossilised the Phase 1 source.)*
- **Estimator:** odds ratio of meeting the social minimum, Democrat vs Republican, on the
  2×2. Fisher exact p reported but **not** part of the scoring rule.
- **Scoring map:** `OR ≥ 1.30` → HIT. `OR < 1.30`, or OR < 1 → MISS.

---

## §4 Roster, ages, calibration, service and legal record

### D1
- **Estimator:** count of the 59 races with a material change — withdrawal, replacement
  nominee, party switch, or suspension — after 2026-08-21, from Tier 1/2 sources.
- **Scoring map:** `≥ 2` → HIT, `< 2` → MISS.

### E1
- **Scoring map:** `2 ≤ newly establishable ≤ 4` → HIT, else MISS. Establishable means
  meeting the frozen evidentiary standard, not appearing on a data broker.

### E2 — the only prediction with an unscoreable branch
- **Scoring map:** resolves to the older value → HIT. Resolves to the younger → MISS.
  **Does not resolve → UNSCOREABLE**, counting as neither.
- **Declared:** the 182:139 window assumes uniform birth dates; US births are seasonal by
  roughly ±5–8%, which places the true probability in the 53–61% band. Publishing 57% to
  the point without that caveat was overprecision.

### E3
- **Scoring map:** `≤ 1` of the four markers resolves → HIT. `≥ 2` → MISS.

### F1
- **Scoring map:** the October publication contains ≥ 1 item explicitly labelled
  RETRACTED or DOWNGRADED, referencing a claim published before 2026-09-01 → HIT. None
  → MISS.
- **Declared conflict:** this outcome is partly under the author's control. It is scored
  on the published artifact rather than on judgment, which is the most externally
  checkable form available — **but it is not a clean forecast and is not read as one.**

### F2 — COMMITMENT, NOT SCORED
Publish the Brier score, the log-odds calibration shift â, and its CI. A commitment
cannot be self-referential; this is why it is not a prediction.

### F2′ and F2″ — same proposition, two prices, both scored
- **Estimator:** â, the MLE of the shift in `y ~ logistic(logit(p) + a)`, fitted on the
  **23 predictions excluding these two** (§0).
- **Scoring map, both:** `â < 0` → HIT. `â ≥ 0` → MISS.
- **As at issue** the register's own confidences give
  P(â < 0) = **0.9356**.
- **Declared:** F2′ is priced at 60%, F2″ at 93% — the same claim, repriced once four
  predictions were known dead. Both are scored. **The 33-point gap is the size of the
  advantage declined**, and F2′ will hit cheaply.

### G1
- **Population:** the roster, **Unclear removed from both sides** (§0).
- **Estimator:** Fisher exact, two-sided, on the 2×2 of veteran status by party.
- **Scoring map:** `p > 0.05` → HIT, `p ≤ 0.05` → MISS.
- **Basis as at issue:** 17/54 R vs
  10/56 D, **p = 0.1223,
  OR = 2.114**. *(The figure first published, p = 0.186 on
  denominators of 57 and 57, counted the Unclear candidates as non-veterans — the
  opposite of the rule. The verdict is unchanged; the basis figure was wrong.)*

### G2
- **Any recorded legal matter** = the legal column is not "None found" (Conviction,
  Charge filed, Settlement, or Civil judgment). **The exculpatory qualifier travels with
  the charge in every published form** — `roster_115.csv` carries the full detail column,
  not the one-word summary.
- **Estimator:** Fisher exact, two-sided, by party.
- **Scoring map:** `p > 0.05` → HIT, `p ≤ 0.05` → MISS.
- **Basis as at issue:** 5/56 R vs
  7/57 D, **p = 0.7616,
  OR = 0.700**.

---

## §5 Adjudication

1. **Every prediction above has a total scoring map. No outcome requires judgment.** If
   one is found that does, that is **a defect in this document**, logged as such in the
   Source Log — not resolved by the author's discretion.
2. **Scoring runs 2026-10-31**, on data downloaded once, on a date recorded in the log.
3. **Every input is published with the result** — the excluded-candidate list, the
   media-background assignment, the committee-chair coding, the per-candidate collection
   failures, the Unclear list — so a third party can recompute every number.
4. **Disputes:** any reader may submit a challenge with the recomputation attached.
   Sustained challenges are published as errata beside the original, never as edits.
5. **No prediction is withdrawn after registration.** The only permitted post-registration
   status is HIT, MISS, or E2's single UNSCOREABLE branch.
   **Case law, added 2026-08-30:** this rule was tested on 2026-08-24 when A5's
   characterisation was found wrong within a day of registration. The prediction was
   **not** deleted. It stands, is scored, and is marked retracted-in-substance with the
   reason published on its card. **A rule that binds only until it is inconvenient is not
   a rule** — and the cost of honouring it is one near-certain undeserved hit, disclosed
   in advance.
6. **Every published figure comes from a script in the repository.** Any figure appearing
   in publication without a corresponding run is a defect and is logged as one.

*v2 issued 2026-08-30. Corrections go to v3 with a changelog; nothing here is edited in
place. Basis figures are as at issue and are not live.*
