Registered in advance · 2026 House Battleground Index

Twenty-Five Bets, Four Lost and One Retracted

Candidates in these 59 races who have never held the seat take 0.51% of their money from corporate and trade PACs?. Candidates who have served in Congress take 12.40%. That gap — whatever small movement the measurement window adds to the exact two figures — is not in dispute. What I published about the rest of it — how much further the figure climbs the longer someone stays — was wrong, and I found out by trying to write it down as a fact. What is true instead: the money arrives with the seat, not with the seniority. Getting to Congress at all is worth about 9 points of corporate-PAC share; each further year served may add only about a third of a point, and even that is uncertain. My August headline was about that small second number, and I had it far too high. The fitted figures, and how I got there, are below.

Getting there cost me five registered predictions. Four are already lost — two measured against the frozen data, two conceded on reasoning I could not compute — and a fifth is retracted; because the test I wrote for it was weak, that fifth will probably still be scored a hit. All five are below, unedited, at the confidence I gave them.

v1 registered 2026-08-22 v5 binding, 2026-08-24 Page last revised 2026-09-02 25 predictions · 4 already lost Scored 2026-10-31

The test of a method isn’t whether it produces impressive claims. It’s whether it can take one away from you.

The finding, after it was checked

This project keeps a verified record of 115 candidates across the 59 most competitive U.S. House districts — occupation, military service, legal record, campaign finance and communication style — built under a written rulebook frozen before the last twenty candidates were coded?, and tested by a separate blind coding session — an AI session, not a second person — given only the rulebook. Blank means “no credible source located.” It never means “assumed none.”

One naming rule, stated once: the specification’s word for anyone with prior House or Senate service is “incumbent” (tenure > 0) — and that group includes one former member running again and one sitting incumbent whose FEC? record still codes her as a challenger, so this page says “candidates who have served” where it means that population, and never calls them sitting members.

Its strongest result was about money and time in office. The version I published in August said corporate money climbs steadily with seniority, at a correlation? of 0.82 across all candidates and 0.70 among incumbents by FEC code — not the same population as the 48 with prior service, a difference the companion page walks through. Recomputed from the frozen data under the frozen specification, those are 0.576 and 0.311. What is left is real but much weaker than I said:

Corporate and trade PAC share against years in Congress Scatter plot of 112 battleground House candidates. The horizontal axis is years of congressional service, 0 to 45. The vertical axis is the share of campaign receipts coming from corporate and trade PACs, 0 to 40 percent. 64 candidates who have never served sit in a flat band at zero years, averaging 0.51 percent. 48 candidates with congressional service spread across the rest of the chart, averaging 12.40 percent, with an upward trend of correlation 0.31. The gap between the two group averages is 11.89 points; the fitted line climbs 14.62 points across the observed range; the points scatter widely around it. A two-term regression in the text finds both a step at arrival and a small, uncertain per-year slope, with the step much the larger. 0% 10% 20% 30% 40% 0 10 20 30 40 the step: 0.51% → 12.40%, never served vs served the slope: r = 0.31 among those who have served Matt Schultz (AK-00) — never served, 0.00% Bill Hill (AK-00) — never served, 0.00% Rhett Marques (AL-02) — never served, 6.45% Amish Shah (AZ-01) — never served, 1.45% Jay Feely (AZ-01) — never served, 0.83% Jonathan Nez (AZ-02) — never served, 0.00% JoAnna Mendoza (AZ-06) — never served, 0.07% Richard Pan (CA-06) — never served, 6.10% Kevin Lincoln (CA-13) — never served, 0.31% Kyle Kirkland (CA-21) — never served, 0.00% Randy Villegas (CA-22) — never served, 0.00% Chuong Vo (CA-45) — never served, 0.00% Marni Von Wilpert (CA-48) — never served, 0.32% Jim Desmond (CA-48) — never served, 0.33% Dwayne Romero (CO-03) — never served, 0.00% Jessica Killin (CO-05) — never served, 0.00% Manny Rutinel (CO-08) — never served, 0.00% Scott Singer (FL-25) — never served, 0.00% Christina Bohannan (IA-01) — never served, 0.07% Lindsay James (IA-02) — never served, 0.00% Joe Mitchell (IA-02) — never served, 2.48% Sarah Trone Garriott (IA-03) — never served, 0.00% Barb Regnitz (IN-01) — never served, 0.00% Matthew Dunlap (ME-02) — never served, 0.00% Paul LePage (ME-02) — never served, 0.10% Sean McCann (MI-04) — never served, 0.19% William Lawrence (MI-07) — never served, 0.00% Thomas Smith (MI-08) — never served, 0.00% Christina Hines (MI-10) — never served, 0.00% Michael Bouchard (MI-10) — never served, 0.00% Eric Pratt (MN-02) — never served, 2.45% Sam Forstag (MT-01) — never served, 0.00% Aaron Flint (MT-01) — never served, 4.57% Laurie Buckhout (NC-01) — never served, 0.71% Jamie Ager (NC-11) — never served, 0.02% Denise Powell (NE-02) — never served, 0.00% Brinker Harding (NE-02) — never served, 1.43% Rebecca Bennett (NJ-07) — never served, 0.00% Rosie Pino (NJ-09) — never served, 0.00% Gregory Cunningham (NM-02) — never served, 0.00% Carrie Buck (NV-01) — never served, 0.00% Martin O'Donnell (NV-03) — never served, 0.00% Cody Whipple (NV-04) — never served, 0.79% Christopher Gallant (NY-01) — never served, 0.00% Michael LiPetri (NY-03) — never served, 0.00% Jeanine Driscoll (NY-04) — never served, 0.00% Cait Conley (NY-17) — never served, 0.02% Peter Oberacker (NY-19) — never served, 0.00% Eric Conroy (OH-01) — never served, 0.94% Derek Merrin (OH-09) — never served, 0.33% Kristina Knickerbocker (OH-10) — never served, 0.00% Carey Coleman (OH-13) — never served, 0.00% Don Leonard (OH-15) — never served, 0.00% Bob Harvie (PA-01) — never served, 0.00% Bob Brooks (PA-07) — never served, 0.84% Paige Cognetti (PA-08) — never served, 0.02% Janelle Stelson (PA-10) — never served, 0.58% Tano Tijerina (TX-28) — never served, 0.00% Eric Flores (TX-34) — never served, 0.23% Shannon Taylor (VA-01) — never served, 0.00% Doug Ollivant (VA-07) — never served, 0.00% John Braun (WA-03) — never served, 1.11% Mitchell Berman (WI-01) — never served, 0.00% Rebecca Cooke (WI-03) — never served, 0.05% Nick Begich (AK-00) — 1.8 years, 5.81% Shomari Figures (AL-02) — 1.8 years, 32.53% Eli Crane (AZ-02) — 3.8 years, 0.13% Juan Ciscomani (AZ-06) — 3.8 years, 10.28% Adam Gray (CA-13) — 1.8 years, 12.97% Jim Costa (CA-21) — 21.8 years, 36.79% David Valadao (CA-22) — 11.8 years, 19.07% Derek Tran (CA-45) — 1.8 years, 5.43% Jeff Hurd (CO-03) — 1.8 years, 9.11% Jeff Crank (CO-05) — 1.8 years, 11.83% Gabe Evans (CO-08) — 1.8 years, 10.63% Jared Moskowitz (FL-25) — 3.8 years, 8.02% Mariannette Miller-Meeks (IA-01) — 5.8 years, 14.68% Zach Nunn (IA-03) — 3.8 years, 17.99% Frank Mrvan (IN-01) — 5.8 years, 15.86% Bill Huizenga (MI-04) — 15.8 years, 24.99% Tom Barrett (MI-07) — 1.8 years, 5.90% Kristen McDonald Rivet (MI-08) — 1.8 years, 10.06% Don Davis (NC-01) — 3.8 years, 20.02% Tom Kean Jr. (NJ-07) — 3.8 years, 9.70% Nellie Pou (NJ-09) — 1.8 years, 11.54% Gabriel Vasquez (NM-02) — 3.8 years, 8.97% Dina Titus (NV-01) — 15.8 years, 14.39% Susie Lee (NV-03) — 7.8 years, 11.52% Steven Horsford (NV-04) — 9.8 years, 27.06% Nicholas LaLota (NY-01) — 3.8 years, 7.93% Thomas Suozzi (NY-03) — 8.7 years, 13.79% Laura Gillen (NY-04) — 1.8 years, 2.47% Mike Lawler (NY-17) — 3.8 years, 8.70% Josh Riley (NY-19) — 1.8 years, 0.60% Greg Landsman (OH-01) — 3.8 years, 4.57% Marcy Kaptur (OH-09) — 43.8 years, 8.41% Mike Turner (OH-10) — 23.8 years, 21.24% Emilia Sykes (OH-13) — 3.8 years, 9.00% Mike Carey (OH-15) — 5.0 years, 39.09% Brian Fitzpatrick (PA-01) — 9.8 years, 16.25% Ryan Mackenzie (PA-07) — 1.8 years, 7.33% Rob Bresnahan (PA-08) — 1.8 years, 9.59% Scott Perry (PA-10) — 13.8 years, 1.88% Henry Cuellar (TX-28) — 21.8 years, 14.81% Vicente Gonzalez (TX-34) — 9.8 years, 19.56% Rob Wittman (VA-01) — 18.9 years, 19.86% Elaine Luria (VA-02) — 4.0 years, 0.11% Jen Kiggans (VA-02) — 3.8 years, 8.51% Eugene Vindman (VA-07) — 1.8 years, 2.02% Marie Gluesenkamp Perez (WA-03) — 3.8 years, 2.69% Bryan Steil (WI-01) — 7.8 years, 18.68% Derrick Van Orden (WI-03) — 3.8 years, 2.81% Marcy Kaptur, 44 yrs — kept in NEVER SERVED YEARS IN CONGRESS CORPORATE / TRADE PAC SHARE
112 of the 115 candidates — 3 have no row in the frozen finance file, for reasons the roster note below spells out — showing corporate and trade PAC share of receipts against years of congressional service, FEC bulk data frozen 2026-08-21 (77 of 112 committees reporting through 30 June 2026; coverage runs 16 June 2026 to 18 August 2026). The 64 who have never served average 0.51%; the 48 candidates who have served in Congress average 12.40%. Among those 48 the upward trend is r = 0.31, with the points scattering 8.64 percentage points either side of the line — which is why the trend looks weak even though the line itself rises 14.6 points across the range. Marcy Kaptur, 44 years and 8.4%, is the single highest-leverage point?; she is circled and kept in, because removing an inconvenient point is how you get a number that doesn’t reproduce. Deleting her takes the incumbents-only correlation from 0.311 to 0.481 — printed because naming an outlier and keeping it proves nothing unless you also print what keeping it costs. Nothing here is scored with her removed. Generated by make_figures.py, published below.

Step or slope: fitted, at last

My first replacement headline said the pattern was mostly a step?, not a slope? — that arriving in Congress moves the money and staying longer barely adds to it. I then argued about that for three days using correlations, got it wrong, corrected it, got the correction wrong, corrected that, and wrote “I do not currently know”. An editor pointed out that this is not an open question. It is a two-line regression? on data that had been frozen for a week: a correlation is one number and cannot separate two effects, so I should have fitted the two. Here it is.

Corporate and trade PAC share regressed on having served and on years served
EffectEstimate 95% intervalp
The step — arriving in Congress at all 9.41 points 4.46 to 14.36 0.0003
The slope — each further year served 0.348 points a year -0.437 to 1.133 0.39

Standard errors are heteroskedasticity-robust (HC3), which lets each observation carry its own residual spread instead of assuming one spread for all. The classical formula gives a narrower slope interval (0.145 to 0.551) at p = 0.0011. The two disagree about the slope for a reason that is stated in the paragraph below, because it is a person, not a formula. The step is significant on either (robust p = 0.0003).

The answer is both, and the step is much the larger of the two for almost everybody. Arriving in Congress is associated with 9.41 percentage points? of corporate and trade PAC share. Each further year adds about 0.348 of a point — small, and how sure you can be of it depends almost entirely on one person. On the robust formula the slope’s interval (-0.437 to 1.133) includes zero, at p = 0.39. But decompose that uncertainty by candidate and 89% of it is Marcy Kaptur alone — the longest career on the chart, sitting far below the line — while the 64 never-served candidates, all at zero years, contribute exactly 0%, because a person with no tenure carries no information about a per-year rate. Refit without her, on the same robust formula, and the slope is 0.727 a year, interval 0.283 to 1.170, p = 0.0017. A pairs bootstrap over everyone, her included, gives 0.050 to 1.061 and lands above zero 99% of the time. So the honest reading is probably positive; whether it is “significant” is decided by one 44-year career, and she stays in, because this page keeps her in everywhere else and does not get to drop her here. Corrected 2026-09-01, and again 2026-09-02: an earlier build called the slope “measurable” on the classical p = 0.0011; the first correction replaced that with the robust figures but blamed the wide interval on the “floor of zeros” among the never-served, which a reviewer showed is wrong — the zeros contribute nothing to it; she does. This is the second correction. The step was never in doubt on any formula. The seniority effect only catches up with the arrival effect after 27 years of service, and 1 of the 112 candidates in the frozen file has served that long. So for 111 of 112 people here, “the pattern” means the step. My August headline was about the slope. Together the two explain 55% of the variation.

What this does not say. No causal claim is made or will be made from it. There are no controls for committee position, majority status, seat safety, district industry mix or total receipts, and a cross-section of 112 candidates in one cycle cannot separate “serving attracts corporate money” from “candidates who attract corporate money win and stay”. It describes a shape. It does not explain it. Removing Marcy Kaptur, the highest-leverage point?, moves the step to 7.37 and the slope to 0.727 — she is holding the slope down, and this is printed for the same reason her point is circled and kept in. This model is not registered and is not scored. It was fitted on 30 August, after the register froze, so it cannot be a prediction and is not being presented as one. A5 below still stands, still retracted, still scored on the test I actually wrote.

The correlation is low because the scatter? is 8.64 points, not because the gradient is small. That is what a correlation could never have told me, and what two coefficients say in one line.

How that number fell from 0.82 to what it is now — and how I spent a day claiming the cause could not be determined when the answer sat in my own records — is the subject of a companion piece: The Headline I Never Re-Ran. It is the follow-up, not the introduction. Everything you need in order to judge these predictions is on this page.

Why the bets are still here

A prediction made after you have seen the data is a description wearing a costume, and from the outside the two look identical. The only fix is to write the claim down first, in public, with a date and a number on it. So each card below carries a threshold, an exact condition that would make it false, and a stated confidence. On 31 October each one gets marked hit or miss on this page, with the actual value beside it, and nothing above gets rewritten.

Which brings us to the four I have already lost. The 0.82 was never a forecast — it was a claim about data I already had, so it gets corrected. And the honest word for what was wrong with it is stale, not false. The reconstruction on the companion page recovers all seven August figures exactly by restricting to the 92 candidates coded when they were computed — so 0.82 was true of the people it was computed over, and stopped being true as twenty more arrived and nobody re-ran it. Declared 2026-08-30 because the distinction cuts against me: calling it “false” sounds worse and lets the four dead bets read as bad forecasts. They are not bad forecasts. They are one vintage decision, counted four times, and the register is harsher on me than the facts are. The predictions built on top of it are forecasts, and a forecast you rewrite after learning it loses is not a forecast. So A1a, A1b, A3 and A4 stand word for word at 92%, 88%, 65% and 70%, and they will be marked missed in the same table and the same denominator as everything else. The corrected versions are registered separately and scored too. The real penalty is specific and computed: freezing the four dead bets rather than deleting them costs 0.0871 Brier — and no more than that: it is the cost if every remaining prediction hits, and it shrinks as they miss, so it is a ceiling, and the two scores that follow are ceilings on the same assumption. With the base, which I had not printed: delete the four and the score is 0.0887; freeze them and it is 0.1759. So the penalty is not a rounding item, it is roughly double. Lower is better on this scale. An earlier version of this paragraph claimed the mistake was “charged twice”; v4 retracted that as arithmetically backwards, and it stays retracted.

The four I have already lost

Priced on 22 August against figures — 0.82 across everybody, 0.70 among those who had served — that the data does not support. Frozen here at the confidence I gave them. Two of the four are measured; two are conceded. A1a, A1b are computed against the frozen data and land outside the bands they registered, and the figures are on their cards. A3 and A4 are not computed and cannot be from this data — the index codes no chair or ranking-member field, and the frozen finance extract is the 2026 cycle only. I concede them because the figure they were priced against was wrong and I do not expect them to survive, which is an argument, not a result. They are badged “conceded, not measured” and they count as misses in every total on this page, because conceding a forecast I would otherwise have to score is not a way to avoid scoring it — a reviewer noticed the page reported an argument in the same words as a measurement.

The ruling that decides 3 of the verdicts

Registered 24 August, before October can see the answer

Three of the 59 districts — OH-09, TX-28, TX-34 — sit outside the selection rule the frozen specification prints. That rule, stated here because it is the most loaded choice in the whole index and it had been left in a spreadsheet: a district qualifies if the 2024 presidential margin was 10 points or less, or the 2024 House margin was 10 points or less — or less, not under: two districts, NY-01 and OH-15, sit at exactly 10.0 and are in on that boundary. These three clear neither — presidential margins 10.5, 10.3 and 10.1 points, with House margins not comparable after redistricting, so neither leg of the rule is satisfied. They were added as marquee races; the flag that was supposed to accompany them was never published. That matters because the choice of universe? now chooses verdicts: under the strict rule the pooled correlation is 0.712 instead of 0.576 and the incumbents-only figure is 0.499 instead of 0.311 — enough to flip A1b from a near-certain miss to a hit, A1a′ from a hit to a miss, and A1b′ from a hit to a miss, its band running to 0.6941 against a rule-compliant 0.712. Corrected 2026-08-30: until today this paragraph said the ruling decided two verdicts and named only the first two. A1b′ was computed by no script and named in no document. Counting it, the universe I chose is one prediction better for me than the alternative rather than the wash it was presented as — which is the same failure the glossary already confesses about the Fisher-z bands: computing some and printing fewer.

So the ruling is made now, blind: every money prediction is scored on the as-published 59-district universe — the one the bets were priced against — and the rule-compliant figures are computed by the same script and published beside every verdict, unscored and labelled. Nobody, me included, gets to pick a universe in October after seeing which one flatters.

Who scores this — the thing all of the above does not fix

Everything on this page secures the inputs. The register is frozen, the data is frozen, the digests are printed, the record is timestamped by a third party, and five build gates refuse to ship a page that drifts from any of it. None of that touches the verdicts. On 31 October I mark 25 predictions hit or miss, on my own site, with my own scripts. There is no external adjudicator, no pre-committed referee, and no arbitration if someone disagrees with a call. The scoring maps are written down and frozen precisely to make that judgment as small as possible — but small is not none, and the person checking whether I applied them is me.

The honest mitigation available to a reader is that every scored figure recomputes from the published bundle, so a disputed verdict can be checked by running the same script against the same rule. The honest limitation is that nothing compels me to accept the answer. Stated here, in advance, because a reviewer asked who scores this and the page had no answer anywhere.

The specification moved after the register froze — declared 2026-08-30

The register froze on 24 August. The document that decides every verdict — the analysis specification — was reissued on 30 August, six days later, and it changed five scoring rules that its own summary does not mention. An independent reviewer found this by diffing the two versions. I did not disclose it because I had not noticed it, which is worse. Four of the five run in my favour or cost me nothing; the fifth, D1’s narrowed event list, runs against me — fewer things count as a change, so the threshold is harder to reach — and the spec says so. None is reverted: reverting after seeing the consequences is the move this page exists to refuse. They are printed instead, and priced.

G1 is the serious one, and it turns a bet into a foregone conclusion. The registered claim says the veteran-share difference stays non-significant “when extended beyond the battleground set.” Version 1 of the specification scored it on “the extension set.” Version 2 scores it on “the roster” — which is the battleground set, the very thing the claim says to go beyond. On the roster the answer is already in hand: p = 0.1223, a hit. Under any real extension it misses; at twice the sample the same rates give p ≈ 0.03. So a prediction I priced at 62% became a near-certain hit by a definition written after the freeze. It is scored as the specification now says, and the hit is worth nothing. Read it as a defect in my drafting, not as a forecast that came in.

The other four. D1: the qualifying events dropped from six to four (death and disqualification removed) and the window changed from “dated between 2026-08-22 and 2026-11-03” to “after 2026-08-21” — the same start day, worded differently, no end date, and “dated” deleted, so it no longer says whether the date is when the change happened or when I found out. E1: “an age meeting rule G4 of the evidentiary standard” became “meeting the frozen evidentiary standard, not appearing on a data broker” — a named rule replaced by a gesture, on the card I most control. F2: the commitment to publish the Murphy decomposition was dropped. That is the only listed statistic that separates whether my confidences are unbiased from whether they carry any information, and with four dead bets at high confidence it is the one that would look worst. I am reinstating it: the decomposition will be published in October whatever it says. §5.5: “no prediction is withdrawn after 2026-08-22” became “after registration,” which retroactively makes an earlier retirement a violation of a rule written later.

One thing to know when you read the specification: its thresholds are frozen literals — they are the claim — but its basis figures are measurements labelled “as at issue”, meaning as at 30 August. Where a coding correction has moved one since, the specification keeps the older number and this page carries the newer one. G2 is the live example: the binding v3 spec prints the basis it had at issue, 6 Republicans to 7 Democrats, and the card now prints 6 to 6, because on 2 September the spec’s own enumeration of what counts was applied to the count and it excludes a pending suit the broad count had included; the verdict is the same at p = 1.00. That is the design, not a contradiction — but two independent reviewers read it as one, so it is worth saying here rather than leaving in a header.

Nothing here is being corrected in place. Both specification versions are published, hashed, and anchored in the timestamp above, so the diff that produced this paragraph can be run by anyone.

The three that carry the most weight

The replacement finding, a bet against something else I have already published, and a bet about how often I expect to be wrong.

The full register

The remaining 18 predictions in the register?, grouped by what they test. Everything registered on 22 August was written to the same rules on the same day; everything added on 23 August was written against the corrected baseline and is marked as such on the card.

The other 18 betsopen — click to collapse

A · The tenure–money relationship

  • A1a′ revisedPendingStated confidence 85%

    Among candidates who have served in Congress, the link between years served and corporate money should still be there in October — weak, but there.

    Registered as Incumbents-only Pearson r stays within 0.20 Fisher-z ? units of z = 0.3214 — the z form governs — which is r in [0.1208, 0.4788].

    Fails if: |zOct − 0.3214| > 0.20.

    Registered 2026-08-23. Basis r = 0.3108, n = 48, p = 0.032, bootstrap 95% CI ? [0.058, 0.683]. Corrected 2026-08-23: this card first printed the band as r in [0.12, 0.48]; both of those rounded edges fall outside the rule they were meant to state, so the four-decimal figures are printed instead. Declared so it cannot be claimed as a triumph later: this is an easy prediction and I am pricing it as one. October is the same 48 people with one more quarter of receipts, not an independent draw. It tests continuity, not the finding.

  • A1b′ revisedPendingStated confidence 88%

    The overall figure — the one that mixes challengers in — should also hold roughly steady, and stay positive in both parties.

    Registered as Pooled r stays within 0.20 Fisher-z of z = 0.6559 — the z form governs — and the sign stays positive within both parties.

    Fails if: the band is missed, or either party’s r drops to zero or below.

    Registered 2026-08-23. Pooled r = 0.5756 (n = 112); DEM 0.5183, REP 0.6830. Permanently labelled: this number contains the challenger/incumbent contrast. It is context for the card above, not a finding of its own.

  • A3′ revisedPendingStated confidence 78%

    Long-serving members pull more corporate money even when you set the committee chairs aside — but only a little more. Stripping out the gavels should barely move the number.

    Registered as Among incumbents, controlling for chair / ranking-member status, tenure’s partial correlation ? falls in [0.10, 0.45].

    Fails if: the partial correlation lands outside [0.10, 0.45].

    Registered 2026-08-23. One binary control is not an identification strategy?. Majority-party status, committee assignment, seat safety, district industry mix and total receipts are uncontrolled. No causal claim is made or will be made from this.

  • A4′ revisedPendingStated confidence 65%

    The same weak pattern should show up in the 2024 cycle. If it only exists in one cycle it isn’t a pattern.

    Registered as Rebuilt on the 2024 cycle, the incumbents-only correlation falls in [0.10, 0.55].

    Fails if: outside [0.10, 0.55].

    Registered 2026-08-23. Largely the same members appear in both cycles, so this is not an independent replication and will not be described as one. The band is wide because I do not know; pretending to a tighter one would be the same mistake twice.

  • A2PendingStated confidence 65%

    This should be a Congress pattern, not a swing-seat one. Check every district in the country and it should look about the same.

    Registered as Extended to all 435 districts, the incumbents-only correlation lies within 0.20 Fisher-z units of the battleground incumbents-only figure computed on the same October data.

    Fails if: |Δz| > 0.20.

    The two samples are nested?, so this is not a two-independent-sample comparison and no p-value? is claimed. It matters even if it holds, because it changes the headline: “battleground incumbents” and “Congress” are different stories.

B · How politicians write

  • B2PendingStated confidence 70%

    A press release talks about a politician; the quote inside it talks as them. Nobody refers to themselves in the third person inside their own quotation.

    Registered as Among the first 100 written attributed quotations collected in the sweep, not one contains a third-person self-reference.

    Fails if: any single instance, in any quotation, of any length.

    A black-box Beta(1,1) posterior on 23-for-23? gives only 19% for 100 more. I am departing from it because press releases quote their principal in the first person by construction — a mechanism, not a streak. If this misses, the mechanism argument was wrong, and that is the more interesting failure.

  • B3Against my own workStated confidence 65%

    I said younger politicians shift their voice more between formal statements and social posts. I now think that was noise.

    Registered as The under-55 / 55-and-over register-gap? split fails to replicate, and the difference is smaller than 1.5 reading grades (an equivalence test ?, margin 1.5 reading grades).

    Fails if: anything other than: the split fails at p ≥ 0.05 two-sided AND equivalence is established at margin 1.5. Inconclusive is a MISS. (Corrected in v4 — the earlier falsifier? omitted the equivalence conjunct? its own claim carried, so an underpowered test could make the claim false while scoring a hit.)

    The second bet against my own published work. It has already been downgraded once from a smooth gradient to a coarse split, it collapses when the oldest member is removed, and it disappears under an alternative measure of the same thing.

  • B4PendingStated confidence 55%

    Cabinet secretaries and senators change how they write more sharply than House candidates do — more formal in a written statement, plainer in a social post. The gap between the two should be at least three and a half school grades of reading difficulty for them, and no more than two and a half for the candidates.

    Registered as National office-holders’ written-quote-to-social gap has a median of at least 3.5 reading grades, against a candidate median at or below 2.5.

    Fails if: office-holder median < 3.5, or candidate median > 2.5.

    The weakest prediction here and flagged as such: n? = 4 at registration. Instrument? defect declared 2026-08-24 (v5, Ruling 2): the frozen measurement script pooled samples with a bare newline join, and 72 of 468 samples — 66 of them social — end without terminal punctuation, so sentences merged across boundaries and the candidate register-gap median came out 2.38. Corrected, it is 3.54 — through the 2.5 threshold. Second instrument defect, declared 2026-08-30: that median pools 19 candidates, of whom only 3 meet the frozen protocol's own corpus minimums (8 samples per register; 600 words official, 300 social). The other 16 are cells the protocol says must be reported as insufficient corpus, never as a number — and here they bound a comparison the protocol scores. Restricted to compliant cells the median is 3.40: the verdict does not change and the number does. (An earlier wording of this note applied the social floor to both registers and counted 7 compliant — understating the problem by half; the wording is in the source log.) B4 is scored on the corrected pipeline with both values published, which on today’s data makes it an expected miss. The confidence stays 55%: the instrument was wrong, the bet was the bet.

C · What can and cannot be collected

  • C1PendingStated confidence 78%

    Facebook won’t let anyone collect posts at scale, so candidates who mainly use Facebook can’t be measured at all. Not a gap I can close — a limit I have to declare.

    Registered as At least 15% of the 115 candidates have no collectable social register? under the frozen protocol.

    Fails if: under 15%.

  • C2Against my own workStated confidence 60%

    The platform I can collect from most easily leans left, so my own data probably over-samples Democrats. If that’s true it’s a flaw in my work, and it gets the same billing as any other finding.

    Registered as Among candidates whose social cells? meet the minimum, the odds ratio ? of collectability, Democrat versus Republican, is at least 1.30. The roster baseline is 57 D to 56 R.

    Fails if: OR < 1.30, or it runs the other way.

    v1 stated a bare ratio against no baseline. Corrected 2026-08-24 (v5, Ruling 3): the baseline first published was 57:57, “exactly 1.00”. Kevin Kiley left the Republican Party on 9 March 2026 and filed as an independent; the roster now reads 57 D / 56 R / 2 I, baseline odds 1.018. The party-status flag was in the master coding all along — the roster column had fossilised the Phase 1 source. Claim, threshold and confidence unchanged.

D · The index itself

  • D1PendingStated confidence 88%

    At least two more of these 59 races will change shape before Election Day — someone withdraws, gets replaced, switches party or suspends. (v1 priced this at 75%; re-priced to 88% in v2 the same day, before any data, with the reason stated there.)

    Registered as At least two more of the 59 races materially change before Election Day.

    Fails if: fewer than two.

    The index’s own record already holds four material changes absorbed while it was being built — a March party switch, two July suspensions, an August replacement nominee. If this holds it is not a finding, it is a maintenance requirement: any published index must timestamp its roster and re-verify at publication.

E · The missing ages

  • E1PendingStated confidence 55%

    Seven candidates have no age in the index because I couldn’t establish one. I think general-election coverage will print two to four of them.

    Registered as Between two and four of the seven unestablished ages become establishable by Election Day, under the published evidentiary standard.

    Fails if: fewer than two, or more than four.

    Those seven have been searched four times over with different tools, all converging on the same names. At least one of the seven has an age listed on data-broker sites. Those aren’t sources, so the cells stay blank.

  • E2PendingStated confidence 57%

    One Ohio candidate’s birthday falls either just before or just after Election Day, and I cannot tell which, so I cannot say whether the right age is the older or the younger one. The range of possible birth dates splits 182 days to 139 in favour of older, so that is what I am betting.

    Registered as If the flagged Ohio case resolves, it resolves to the older value.

    Fails if: it resolves to the younger value. If it does not resolve, this is unscoreable — v1 left that outcome unhandled.

    The 182-to-139 split assumes birth dates are uniform across the year. US births are seasonal by roughly ±5–8%, which moves this to somewhere in 53–61%. Publishing 57% to the point without that caveat was overprecision. Named 2026-08-30: the flagged case is Carey Coleman (OH-13). Register v1 named him in its own heading (“Coleman’s asterisk resolves to 67, not 66”); the naming was lost in the v2 card rebuild, which left the subject — and which four candidates E3 is about — to be chosen in October, after the answer was visible. Restored here: E2’s subject is fixed in the binding spec; E3’s four are named in the v1 register and now on the E3 card (the spec carries only E3’s scoring map), so both are scoreable by someone other than their author. (An earlier wording of this note blamed “every published version”, which the timestamped v1 disproves; the wording is in the source log.) The claim, threshold and confidence are unchanged.

  • E3PendingStated confidence 70%

    The other four ages carrying a one-year question mark — Schultz, Vo, Knickerbocker and Flint, named in the v1 register so the set cannot be chosen after the fact — stay uncertain. Election coverage tends to reprint the age it already printed.

    Registered as The other four uncertainty markers do not resolve.

    Fails if: two or more resolve.

    Declared 2026-08-30, after an independent review: this falsifier is looser than the claim it polices?. The claim says none of the four resolve; the fail condition needs two. If exactly one resolves, the claim is false and this card is scored a hit anyway. It is registered and I do not get to rewrite it, so the defect is printed instead — and if that is how October lands, the hit should be read as worth nothing.

F · Scoring the scorer

  • F1PendingStated confidence 88%

    Something I’ve already published gets retracted or downgraded in October. This isn’t modesty, it’s a base rate — every round of this project so far has retracted something. (v1 priced this at 80%; re-priced to 88% in v2 the same day, before any data.)

    Registered as The October publication contains at least one item explicitly labelled RETRACTED or DOWNGRADED, referencing a claim published before 2026-09-01.

    Fails if: no such labelled item.

    Declared conflict: this outcome is partly under my control. It is scored on the published artifact rather than on my judgment, which is the most externally checkable form available — but it is not a clean forecast and should not be read as one. The Laplace rule ? on 6-of-6 prior rounds gives 87.5%.

  • F2′ revisedFrozen, underpricedStated confidence 60%

    The same claim as the card above — that I run overconfident — but priced at 60% before I knew four predictions were already lost. It will almost certainly be scored a hit, and cheaply. Left standing anyway.

    Registered as The log-odds calibration shift â comes out negative.

    Fails if: â ≥ 0.

    Registered 2026-08-22 at 60%. It will almost certainly hit, and it will hit cheaply. Frozen rather than deleted, because deleting the easy version and keeping the honest one would still leave me choosing which of my own bets counts.

G · Service record and legal record

  • G1PendingStated confidence 62%

    Republicans in these races include more veterans than Democrats do — 31.5% against 17.9%, or 17 of 54 against 10 of 56 — and this bets the gap does not clear the usual test of significance?.

    Registered as The difference in veteran share between the two parties does not reach significance? when extended beyond the battleground set (Fisher exact, two-sided, p > 0.05).

    Fails if: p ≤ 0.05.

    Currently Fisher exact ? p = 0.1223, odds ratio? 2.114 — in plain English 31.5% against 17.9%, a gap of 13.6 percentage points, not the “twice as likely” the odds ratio sounds like — on 17 of 54 Republicans against 10 of 56 Democrats. Corrected 2026-08-23: the basis first published was p = 0.186 on denominators of 57 and 57, which counted the three candidates whose service could not be established as non-veterans — the opposite of what the frozen specification says to do with them. The verdict does not change; the basis figure was wrong, and a wrong basis figure is what started all of this. Registered because a reviewer noticed that military service and legal record — the two most politically explosive variables the index codes — carried no predictions at all. It could not prove that was deliberate. Neither can I. Closing it beat arguing about it. Declared 2026-08-30: the specification's population for this card was changed after the register froze, from “the extension set” to “the roster” — which is the set this claim says to go beyond. That makes it a hit on data already in hand. The full declaration is above; the short version is that a hit here is a drafting failure, not a forecast. Declared 2026-08-30, after an independent review: this card scores a hit when the test simply fails to find a difference? — which an underpowered test does for free. B1 and B3 were rebuilt with an equivalence conjunct? to close exactly this hole; this one was registered before that fix and cannot be rewritten now. Read a hit here as weak evidence, and note that Fisher exact is the most conservative of the exact tests available, which pushes the p-value in the direction I am betting on.

  • G2PendingStated confidence 80%

    Recorded legal matters? are split about evenly between the parties in these races — 6 Republicans, 6 Democrats — and should stay that way.

    Registered as Through the October refresh, the share of candidates carrying any recorded legal matter does not differ significantly by party (Fisher exact, two-sided, p > 0.05).

    Fails if: p ≤ 0.05.

    Currently Fisher exact? p = 1.0000, odds ratio? 1.020 — in plain English 10.7% against 10.5% — on 6 of 56 Republicans against 6 of 57 Democrats. Basis corrected 2026-08-30: one Republican's legal cell read “None found” while its own detail documented a settled personal-capacity civil suit — inconsistent with how comparable civil matters are coded elsewhere in the index. Correcting it made the parties more even and pushed the p-value further into the range this card bets on. A correction that helps me needs saying louder, not quieter. Basis questioned 2026-08-30 on legal review, settled by verification. Two rows — one from each party — carried an adverse legal matter cited only to a wiki; both were briefly coded “not established” and dropped from this basis, then re-sourced to named contemporary outlets and court records the same day, so they are established, carry their findings, and stand in this basis at 6 Republicans and 6 Democrats. The rule survives them: an adverse coded value the index cannot stand behind is downgraded and excluded — the answer to an unsupported claim is to source it or drop it, and here it was sourced (to named contemporary outlets; the court records themselves are not linked, and those rows now carry the outlet tier). The full arc is in the source log. Basis corrected again 2026-09-02, twice, and the verdict does not move. First: the binding spec says what counts here is one of four enumerated values — Conviction, Charge filed, Settlement, Civil judgment — and the code had been counting anything that was not “None found”, which swept in one Democrat’s pending civil suit. Under the spec’s own rule the basis is 6 Republicans of 56 against 6 Democrats of 57, not the 7 Democrats the broad count gave and the spec’s own printed basis repeated; p = 1.00 either way. Second: one Democrat whose court record reads “waiver guilty, conviction date 2026-04-14” on a noise-ordinance count had been coded “Charge filed” while two Republicans’ guilty pleas to comparably minor matters were coded “Conviction”; he is now coded as the record reads, under the same rule applied to them (the obstruction count against him was dismissed by the prosecution and is recorded as such). He was already in the basis, so it does not change the count; it changes a label, in the direction of even-handedness, and it is said here because it is the kind of asymmetry this site exists to catch. Declared 2026-08-30: like G1 above, this card scores a hit when the test merely fails to find a difference, which an underpowered test does for free — G1's card states the full weakness and it applies here unchanged. Read a hit as weak evidence.

What the words meanplain definitions

↩ back to the top of the page · or use your browser’s back button to return to the exact card you came from.

This box is open on purpose: every “?” on the page jumps straight to a definition inside it, and a box you had to open first would break that jump. These are definitions and nothing else. Where the honest definition is less flattering to this page than the usual one, the honest one is here, and it is marked.

Words about this register

register
Used on its own, on this page, it always means one thing: the list of predictions, written down and dated before the October data existed. Some of them, the money ones, were written around figures already measured in August — that is stated on the cards, and it is why several of them are called easy there. Corrected 2026-09-02: this entry used to say the stated confidence had never changed. That is false against the frozen v2 register, which says on its own first page that nine predictions were re-priced on 22 August, most of them upward, after a reviewer found the easy ones priced below their true odds in a pattern that would have flattered the scorecard — A1 from 80% to 92% and 88% on its two halves, D1 from 75% to 88%, F1 from 80% to 88%, among them. That re-pricing happened once, the same day, before any October data existed, with its reasons written in v2; since v2 no stated confidence has changed, the build checks every card against the frozen register, and nothing has ever been deleted. Claims and fail conditions were amended, in versions two through five, and every amendment is on the card that carries it — B1 and B3 both say so. Since v5 froze on 24 August nothing has been rewritten; corrections since are added below the claim, dated, never over it. Five versions exist and all five are published. Where the word instead means how formal someone’s writing is, it never appears alone — it is always social register, register gap or social cells, all defined further down.
universe
Which people or races a number is computed over. Most arguments about a statistic on this page are arguments about its universe rather than about its arithmetic: the same pooled correlation is 0.576 on the 59 districts as published and 0.712 on the districts that actually meet the selection rule. That is why the universe is ruled on in writing, in advance.
pooled
Everybody together — the 48 candidates who have served in Congress and the 64 who never have, in one calculation. Incumbents-only means the 48 alone. Nearly every pair of numbers on this page that looks like a contradiction is one of each: pooled 0.576 against incumbents-only 0.311 is not a disagreement, it is two different populations.
n
How many cases a number was computed from. n = 4 means four. It is the first thing to look at on any card here, because most of the weak predictions on this page are weak for that reason alone.
coding, coder, coded
Nothing to do with programming. Coding is reading the sources for one candidate and recording what they say against a written rulebook — occupation, service, legal record — so that two coders following the same rulebook would write down the same thing. A separate, blind AI coding session — not a second person — was given the rulebook and nothing else. The number, with its base: agreement was 46 of 50 sampled decisions — about ten candidates, not the whole index — and it was measured against version 1.0 of the rulebook, not the 1.1 in force. Faith was the weakest field by a distance, carrying three of the four disagreements. Corrected 2026-09-02: this entry used to say the cases where the two codings differed were not published. They are: all four are listed by candidate and field, with both codings, in VALIDATION_RESULTS.md in the bundle. What is not published is the blind coder’s full 50-decision sheet, so the 46 agreements are the part you take on the summary.
falsifier
The written condition that turns a prediction into a miss, fixed before the answer can be known. Every card carries one, under Fails if.
conjunct
One of the conditions joined by “and” inside a single claim. If a claim has two conjuncts, both have to hold for it to be true — so a falsifier that tests only one of them is easier than the claim it is meant to police. One card here still has it — E3, whose claim is that none of four things happen and whose fail condition needs two — and that card says so.
instrument
Two different things on this page, and the sentence around it tells you which. (1) The measurement script — the code that turns filings or text into a number. (2) This whole scoring exercise, considered as a way of measuring how well calibrated I am; in that second sense it is not a real instrument yet, because two dozen predictions cannot separate a good forecaster from a lucky one.
the null, and failing to reject it
The null is the assumption that there is no difference. When a test does not find one, the only thing it can report is that it failed to reject the null — “no difference showed up”, not “there is no difference”. A small or noisy sample fails to reject almost anything, so treating that failure as proof of sameness pays a reward for weak evidence. The equivalence test below is the fix. Not all of this page is fixed: B1 and B3 carry an equivalence conjunct, and G1 and G2 do not — they score a hit on a test that merely failed to find something. B1 and B3 were fixed while the register was still open; G1 and G2 were not, and the register froze on 24 August, so they cannot be fixed now. The defect is declared on each card instead. (E3 has the same shape of problem for a different reason and no test at all; its card explains itself.)
significance, and two-sided
Calling a result significant here means only that its p-value fell below 0.05. It is not a claim that the difference is large or that it matters. Two-sided means the test counts a difference in either direction, so a prediction cannot quietly be scored on whichever half of the possibilities suits me.
identification strategy
The full argument for why a measured association is the thing you say it is, rather than something else moving both numbers at once. Holding one yes-or-no variable constant is not one, and this page says so on the card where it does exactly that.

Words about scoring my own confidence

Brier score
A score for stated confidences, running from 0 to 1, where lower is better and 0 would be perfect. Say 90% and be right and it costs 0.01; say 90% and be wrong and it costs 0.81. The rule it replaced could be improved by shading every number downward without changing a belief; this one cannot, because it rewards saying what you actually think. The honest qualifier, added 2026-08-30: that only holds if what you think is right. A forecaster who knows he is overconfident — and F2″ on this page bets 93% that I am — would score better by shading toward 50%, because he would be moving toward the truth rather than away from it. So the rule stops a bare gaming move and does not stop that one. It is a reason to read â beside the Brier, not instead of it.
calibration, and the shift â
Calibration is whether “80% sure” comes true about 80% of the time. â is a single number for which way I lean: negative means I was more confident than I was right. It is fitted on 23 of the 25 — F2′ and F2″ are left out, because they are predictions about â, and a statistic that included them would be scoring itself. It measures the lean only, not whether the predictions were any use.
log-odds
Confidence rescaled so that 50% sits at zero and every doubling of the odds is the same-sized step wherever you start from. Confidences are compared this way so that the distance from 90% to 95% is not treated as smaller than the distance from 50% to 55%.
Laplace rule, Beta(1,1), black-box
Standard ways to turn “it has happened every time so far” into a probability for next time without assuming the streak holds. Beta(1,1) is the flat starting assumption that treats every underlying rate as equally likely before you look, and black-box means the number is computed from the bare count alone, ignoring everything I know about why the run happened. Always ask how far ahead. The same 23-out-of-23 record gives 96% for the very next case and 19% for the next hundred in a row, and the second number is the one printed on B2.

Words about the money and the record

PAC, and “corporate and trade”
A political action committee is an organisation that pools donations and gives them to campaigns. Corporate and trade is read off the committee type in the federal filings rather than judged case by case: company committees, industry-association committees. Labour committees, party committees, leadership PACs and independent expenditures are excluded. One judgment call, since the filing codes do not make it for me: the corporate bucket includes committees of corporations without capital stock — mutuals, non-profit hospitals and the like — which are not what most readers picture. The exact codes are in the specification so the choice can be reversed by anyone who disagrees.
FEC
The Federal Election Commission, the agency campaigns file their finance reports with. FEC bulk data is the raw downloadable form of those filings. Where this page says the FEC codes someone as a challenger, that is the filing’s own status field — not my classification, and occasionally not what the person’s career would lead you to expect.
Anything appearing in a public court or regulatory record: a charge, a suit, a finding — and its outcome. It is deliberately broad and deliberately not a synonym for wrongdoing. A dismissal and an acquittal are recorded legal matters too, and wherever a charge is printed here the outcome is printed with it. One boundary the definition does not state and the practice draws: litigation a candidate was party to only in an official capacity — sued as mayor, as county judge, as a state officer — is recorded in the row’s notes, not in the legal column that G2 is scored on, unless it ended in an admission or a settlement by the person. That keeps G2 a count of personal matters; it is a choice, it is applied to both parties, and G2 comes out the same either way at p = 1.00. And the floor under the whole column: “None found” means nothing surfaced in a general-press search. Court dockets were consulted for a handful of rows that already had a matter in the press; 102 of the 115 rows say in their own detail that no docket search was performed. So the count G2 is scored on is a press-surfaced floor, not a court-record census, and press depth is greater for sitting members, who are most of the rows counted.

Words about the writing measure

reading grade
Roughly the US school year at which the writing becomes easy to read. Grade 8 is a newspaper; grade 14 is a legal filing. Watch for levels versus differences, the same trap as percentage points: a margin of “1.0 reading grades” means one school year of difference between two groups, not writing pitched at first grade. It is computed from sentence length and syllable count — syllables set most of the level, but sentence length is what a missing full stop moves — which is why a missing full stop is not a cosmetic problem here. Samples that ran together across a missing full stop are exactly what moved B4’s median from 2.38 to 3.54, through the threshold it was registered against.
social register
How formal a politician’s writing is in their own social posts, measured in reading grades. Nothing to do with the register of predictions above, and nothing to do with any social directory.
register gap
How far a politician’s writing moves between a formal quotation and a social post, in reading grades. A gap of 3 is about three school years of difference between the two.
social cells, and the minimum
A cell is one candidate’s collected sample of one kind of writing. A cell holding fewer samples than the frozen protocol requires is reported as insufficient corpus and that candidate drops out of that measure — because a reading grade computed from four posts is a number, but it is not a measurement. What does not happen is the whole prediction quietly becoming unscoreable: if too few candidates clear the minimum, the writing predictions are scored as misses, so attrition costs me rather than excusing me.

Words about the statistics

correlation (Pearson r)
One number from −1 to +1 for how closely points follow a straight line that tilts. The tilt is not optional: points packed tightly along a flat line score zero, not one. Around 0.3 is a loose tendency; above 0.7 is a tight one. Two things it is not. It is not a measure of size: r is unit-free, so a relationship can move a great distance and still score low if the points scatter widely around the line — which is the case here, and reaching for a correlation to make a claim about size is the specific mistake this project made twice. And it says nothing about cause.
step
A jump between two groups: everyone on one side of a line sits at about one value, everyone on the other side at about another, and the distance between the two is the step. Here it is the gap between candidates who have never served in Congress and candidates who have — 11.89 percentage points. A step says arriving is what matters and the years afterward do not. Step or slope was the open question on this page until it was fitted: the regression section answers it — both, with the step far the larger for almost everyone — and the claim that it was mostly a step is registered below as A5 and retracted, because it was tested with a tool that could not answer it.
regression, and what these two coefficients mean
A correlation is one number and answers one question: how tightly do two things move together. A regression fits several effects at once and gives each its own number, so it can answer “how much of this is arriving, and how much is staying?” — which a correlation structurally cannot. Here two things are fitted: an on/off term for having served in Congress at all, and a count of years served. The first is the step, the second the slope. Each comes with an interval, which is the range the data cannot rule out; when that range excludes zero, the effect is distinguishable from nothing. Two numbers on this page are both called the step, and they are not the same: the plain difference between the two group averages is 11.89 points (12.40% against 0.51%), while the fitted step is 9.41 — the jump at zero years served. The gap between them is the slope working on the 7.1 years the served average: 9.41 plus 0.348 × 7.1 is roughly the raw 11.89. Neither is wrong; they answer different questions, and the raw gap folds some seniority into what it calls arrival. It took an editor rather than a statistician to point out that this was the tool the question needed, after three days of arguing about it with a correlation. Fitting is not explaining. Nothing about which way the arrow points follows from any of it.
slope, and the fitted line
The fitted line is the single straight line drawn through a cloud of points to sit as close to all of them as it can. Its slope is how steeply it rises: here 0.348 percentage points of corporate money for each additional year in Congress, fitted through the 48 who have served rather than through all 112 points on the chart. Multiply it by the span those 48 actually cover — 1.8 to 43.8 years — and you get how far the line climbs. A slope says every extra year adds something. Step or slope was the open question on this page; the regression section settles it as both, the step much the larger, with the slope’s significance hanging on one career.
scatter (residual spread)
How far the points typically sit from the line, in the units of the thing being measured. Here it is 8.64 percentage points, and it is most of the reason the correlation is low: the line does rise, but individual candidates land nowhere near it. Not all of the reason — deleting one point on this page lifts the correlation by half again while barely touching the scatter, so how much tilt there is matters too. A correlation answers “how tight”; this answers “how far off”, and for a reader deciding whether a pattern is any use, the second question is usually the one that matters.
percentage points
The gap between two percentages, as opposed to a percentage of a percentage. Going from 0.51% to 12.40% is a rise of 11.89 percentage points; calling it a rise of 2,320% would be arithmetically true and useless. Every money difference on this page is in percentage points. The levels themselves — “12.40%” — are ordinary percentages, of a candidate’s total receipts.
zero-inflation
What happens when a large share of the cases sit at exactly zero rather than being spread out. Here 37 of 112 candidates take no corporate money at all — 33% — and every one of them is someone who has never served. A correlation computed across that pile plus everybody else can be produced largely by the gap between the pile and the rest while looking like a steady trend, which is why the pooled and incumbents-only figures are always printed as a pair. Corrected 2026-08-30: this entry first printed 64, which is how many candidates have never served — their years in Congress are zero, their money is not. 64 is a real computed figure, correctly used elsewhere on both pages; the error was substituting the count of one thing for the count of another, which overstated the pile by 73%. Not a number typed into a sentence, then — a number computed and then pointed at the wrong noun, which is the harder version to catch and the reason a build check now compares the two.
rank correlation (Spearman)
The same idea as the ordinary correlation but using only the order of the values, not their size. It survives outliers that would drag the ordinary kind around. Where the two disagree, look for one of two causes: a handful of extreme points, or a step in the data. On this page both are at work at once, and the arithmetic splits them. Among those who have served, removing the one far-out career shrinks the gap between the two measures from 0.153 to 0.027 — nearly all of it was that one person. Across everybody the same removal takes it from 0.256 to 0.127 — about half of it was that person and about half is the pile at zero. Corrected 2026-08-30: this entry first said the pooled gap barely moved and put all of it down to the pile. Half of it is one person.
zero-order
A correlation with nothing held constant: the raw two-variable number, before any control is introduced.
partial correlation
The correlation that remains once a third factor is held constant — here, whether someone chairs or is ranking member on a committee.
suppression
The case where holding a third factor constant makes a relationship look stronger instead of weaker. A3 needs the partial figure to land well above the raw 0.311, which is why the card treats the number going up as the surprising outcome. The honest qualifier: a partial correlation is not always smaller than the raw one, and a small rise can fall out of the arithmetic of the formula without anything interesting behind it. So the word is not the alarm it sounds like. It is a rise of the size A3 asks for — from 0.311 to 0.50 — that would need a real third variable doing real work.
Fisher z
A rescaling of a correlation so that the same-sized change means the same thing at 0.3 as it does at 0.9. Bands are set in z because a band of fixed width in plain correlation would mean something different depending on where it was centred, and I did not want to be choosing that. What it costs, and this is the whole of it: a z-band is not symmetric once you convert it back. A1a′’s allows 0.1900 of downward movement against 0.1680 of upward — 1.13 to one. A1b′’s allows 0.1489 against 0.1185. All four bands, 2026-08-30, because two rounds of this went wrong. A3′’s [0.10, 0.45], set by hand rather than by transform, allows 0.2108 down against 0.1392 up — 1.51 to one, the most lopsided here, and in the direction the mechanics say the number should go. A4′’s [0.10, 0.55] allows 0.2108 down against 0.2392 up, the only band on the page leaning the other way. Ranked by plain distance from 1.00, A4′’s band is the least lopsided (0.119 from even) and A3′’s the most (0.515) — and A3′’s is hand-set, in the direction the mechanics favour. (On a log scale A1a′’s comes out narrowly least; the plain distance is the one this sentence names, so it is the one computed.) Corrected 2026-08-30: I first declared only the smallest asymmetry, and two successive wordings of this entry then misdescribed the ranking (both are in the source log). Computing four and printing three is exactly how selective disclosure happens, so all four are printed. Corrected 2026-09-02: a third wording said A1a′ was least lopsided “by distance from 1.00”, which is true on a log scale and false on the plain distance it named; the ranking above is now computed on that plain distance and printed with the distances, not recalled.
nested samples
Two samples where one sits inside the other — battleground incumbents inside all incumbents, or incumbents inside everybody. They are not independent draws, so tests built for two separate groups do not apply and no p-value is claimed from comparing them.
leverage, and influence
Two different things, often confused. Leverage is how far a point sits from the middle of the horizontal axis. Influence is how much the answer changes when you delete the point. High leverage does not by itself mean high influence — a far-out point sitting right on the line leaves the slope alone. It does not leave a correlation alone, though, and correlations are what this page bets on: dropping such a point in would push the incumbents-only figure up, not hold it still. So leverage is a warning to go and look, never a verdict on its own. Marcy Kaptur is the extreme case of both here: leverage 0.46 against 0.11 for the next-highest, and deleting her moves the incumbents-only correlation from 0.311 to 0.481 and the slope from 0.348 to 0.727 points a year — the slope is the bigger cost of the two, and it is the one the step-versus-slope argument turns on. And the part that costs me most, added 2026-08-30: deleting her would push the incumbents-only figure to 0.481 against A1a′’s band top of 0.4788, and the pooled figure to 0.706 against A1b′’s top of 0.6941 — both live bands broken by the removal of one person. That is the number a hostile reader wants and it was the one I had not printed. All of it is published, because naming an outlier and keeping it proves nothing unless you also print what keeping it costs. She stays in; nothing on this page is scored with her removed.
p-value
How often a result at least this extreme would turn up by chance if there were nothing there. The “at least” is not a nicety: the tests used here add up every outcome as far from the middle as the one observed, and a reader who computes the chance of the exact observed result alone will not get the number printed. Below 0.05 is the customary line for “probably not chance”. It is not the chance that the claim is true, and it says nothing about how big the effect is.
Fisher exact test
A p-value for a small two-by-two table, computed exactly rather than approximated — the usual choice when the counts are in the dozens. Declared, because it runs in my favour: this test is conservative. Its p-values come out systematically larger than those of the other exact tests available, and both cards here that use it (G1 and G2) score a hit when the p-value stays above 0.05. I am using the test most likely to hand me the outcome I bet on. It is a defensible choice and it is still a choice.
odds ratio
A ratio of odds, not of chances. If 8 people in a group of 20 have some trait, the odds are 8 to 12; an odds ratio of 2.0 means the other group’s odds are twice that. It is always further from 1.0 than the everyday “X times as likely” figure — not only when a trait is common, though the gap grows as it becomes common. On G1 the odds ratio is 2.11 while the plain-English figure is 1.76×, and the plainest form of all is 31.5% against 17.9%, which is a gap of 13.6 percentage points. Read 2.11 as if it were the everyday multiplier and you overstate that multiplier by 20%. Where it matters most: C2 does not merely report an odds ratio, it is scored on one — at least 1.30. On a quantity as common as whether a candidate’s posts can be collected at all, an odds ratio of 1.30 is a difference in rate of about five points, not thirty.
equivalence test (TOST)
An ordinary test can only fail to find a difference; an equivalence test can go the other way and rule out any difference larger than a stated size — at the usual error rate, not with certainty. It is how a prediction of “no effect” is made falsifiable instead of unfalsifiable, and it needs that size fixed in advance, which is why every such card prints its margin. The awkward part: it answers a different question from the ordinary test, so both can fire at once — a difference can be real and small at the same time. On B1 that combination is a deserved miss, because the card claims both that there is no difference and that any difference is under a grade; a real one falsifies the first half. It is worth knowing that the two halves can disagree, and worth not pretending the disagreement would be unfair.
bootstrap, and the 95% CI
Re-sample the same data thousands of times, recompute the number on each re-sample, and keep the middle 95% of the answers. The honest reading is about the procedure and not about this one interval: run it on fresh samples and 95% of the intervals it produces would contain the true value. It is not a 95% chance that the true value lies inside this particular range — the looser version of that sentence is the one you will usually see, and it is wrong. And the 95% is nominal: this is the plain middle-95% form, whose real coverage runs below its label on skewed, bounded quantities like a correlation. Read the width, which is wide, rather than the label. Not every interval here is a bootstrap, and the slope has four: the one on the correlation is a bootstrap. The tenure slope of 0.348 a year is the SAME coefficient whether it is fitted on incumbents alone or in the two-term regression (never-served candidates have zero tenure, so only the served can estimate it), but its interval depends on which formula is used, and this page prints more than one. Together, so they can be compared: the incumbents-only textbook interval is 0.040 to 0.656; the pooled textbook interval is 0.145 to 0.551; the heteroskedasticity-robust (HC3) interval is -0.437 to 1.133; and a pairs bootstrap gives 0.050 to 1.061, wider and lopsided, as a bootstrap on this data should be. The textbook forms are symmetric by construction and assume one residual spread for everyone; the robust form lets each observation carry its own. They disagree here for one reason, and it is not the challengers at zero (they have no tenure and so say nothing about a per-year rate): it is that 89% of the robust variance comes from Marcy Kaptur, one high-leverage point with a large residual. The robust interval includes zero with her in; the bootstrap, which resamples her in and out, does not; and without her the robust interval is 0.283 to 1.170. None of the four is “the honest one”; printing all four, and saying which single observation moves them, is. That is why the regression section says the slope is probably positive and that one career decides how sure you get to be.

How this gets scored

On 31 October 2026 — a fixed date, not “mid-October” — each card gets a result badge, Hit or Miss, and the actual value printed beside it. The claim, the threshold and the confidence stay exactly as written.

Calibration is scored with a Brier score ?, which is proper: shading your stated confidences downward costs you. The first version of this page used a rule that did the opposite — a reviewer showed that knocking ten points off every number raised the “well calibrated” pass rate sharply, with no change in belief. Declared 2026-08-30: this sentence used to print that pass rate moving from 0.484 to 0.857. Neither figure is computed by any script here — they came from a reviewer's own working — and the build was quietly exempting both from the rule that every printed number must be recomputable. An exemption the reader is not told about is the same defect as a number in prose. The figures are removed rather than dressed up; the argument they supported does not depend on them, and the rule they broke was replaced. That was the one calibration statistic that pays you to understate.

What this round cannot tell you, stated before the data arrives

25 predictions cannot measure calibration, and a good-looking chart would only imply otherwise — which is why the calibration plot that used to sit here has been removed. Working the arithmetic out exactly — and being careful what it is an interval on, because nothing has been scored yet, so this is the spread of hit rates a perfectly calibrated version of me would produce, not a range for my real accuracy — it runs 56.0% to 88.0%. Because four predictions are already known to be misses, those four are set to zero and the other 21 left at the confidences I gave them, which shifts the range down — 44.0% to 76.0% — and the best score available to me is 84.0%.

And that interval is too narrow, because it assumes these 25 predictions are independent — which this page spends five cards proving they are not. A5 holds in 14048 of 14048 resamples in which A1a′ holds — every one. Corrected 2026-08-30: this sentence read “20,000 of 20,000”, which was the number of resamples DRAWN, not the number in which A1a′ held; A1a′ holds in 14048 of 20000. The implication is unchanged at 100%, and the figure was still wrong, and it was wrong because it was typed rather than computed. It is now computed by figures.py like everything else. A2 is nested inside A1a′. All four dead bets trace to one event, the roster completing. B1 to B4 share a single scoreability floor and fail together. C1 and C2 come from one uncollected sweep. F2′ and F2″ are arithmetic on the other 23. Group them into the 8 things they actually measure, let each cluster resolve as one draw carrying its members, and the range becomes 28.0% to 84.0%. The truth is somewhere between that and the figure above; it is not the figure above. Declared 2026-08-30, after three independent reviewers raised it separately: the limitation I had volunteered was the sample size, and the sample size is the weaker of the two problems. Twenty-five dependent predictions are not twenty-five observations, and fifty of them would not fix it.

And the free hits are not counted anywhere, which is the same failure pointed the other way. Four predictions are pre-declared dead and they are in the headline, the dateline, the lede and a section of their own. 5 more are pre-declared worthless wins — A5 (“it will almost certainly be scored a hit — and that is the problem”), G1 (a definition changed after the freeze), F2′ and F2″ (arithmetic on the other 23), and E3 (a fail condition looser than its own claim) — carrying a mean stated confidence of 74.6%, and they sit in the same denominator. Each is confessed on its own card and the total appears nowhere. Counting the losses loudly and the free wins quietly is exactly what the glossary calls out one level down: computing four and printing three. Declared 2026-08-30, after a reviewer aggregated them and I had not.

B4 is an expected miss too and is not zeroed here, because its outcome still turns on what October collects, where those four are dead against data already in hand — if B4 lands the way I expect, both figures come out lower than the ones printed. This round cannot distinguish a well-calibrated forecaster from a badly overconfident one. The instrument becomes real somewhere north of fifty predictions. This is round one of counting.

Corrected 2026-08-23. This paragraph previously said the interval ran “roughly 47% to 87%”. That figure was computed for an earlier nineteen-prediction version of the register and carried onto a page of 25 without being recomputed — the same failure as the correlation, on the page about the correlation. Both numbers here are now produced by figures.py at build time and cannot go stale.

Two other rules are fixed now rather than later. The writing predictions are scoreable only if at least 90 of the 115 candidates yield 5 or more collectable posts — and failing that floor counts as a miss, not a deletion, because attrition is exactly what C1 and C2 predict, and letting it delete predictions would let the denominator be chosen on a variable correlated with the outcome. Ruled 2026-09-01, the same way: three A-block cards depend on data that still has to be collected before 31 October — A3′ needs a field coded, A2 the all-435 join, A4′ the 2024 extract pulled. The binding specification did not say what they score if that work is never done, and a reviewer pointed out that silence would let me choose the denominator after seeing the outcome. So it is fixed now, against me: a prediction whose data I fail to collect scores a miss, not a deletion. This ruling lives on this page and in the page generator, which is inside the code commitment; it is anchored when that commitment next confirms, and until then it rests on this dated sentence. With that, the only unscoreable outcome anywhere on this page is E2 failing to resolve, which is specified on its card.

Said up front, not in a footnote

The count has changed three times and here is the whole arithmetic. Eighteen predictions were registered on 22 August; two were withheld from scoring, for reasons set out below, so sixteen were published; a rebuild after review added three more, making nineteen; six more were registered on 23 August after the money finding was found to be wrong, giving the 25 scored here. Nothing was deleted at any step. Every version is published unedited with its hash at the foot of this page.

Two of the original eighteen are withheld from scoring. One predicts whether a specific named candidate actively contests his race — that is commentary about a person, not a test of a method. The other depended on access to a government file I may never get, so it could not be fairly scored either way. And here a promise failed: this page used to say both were published only as hashes, “so the count verifies without disclosing the content.” That was false for both — their full texts are printed in the v1 register, a file this page links, hashes and tells you to download. v1 is preserved verbatim, so the disclosure cannot be unmade and will not be scrubbed; what changes is the claim. The withholding failed, the failure is logged, the retired prediction stays retired, and the other is still scored privately — privately meaning the verdict, not the text, which is already public by my own error.

Two other numbers move around this page and both are correct. The roster is 115 candidates, but the money chart plots 112: 3 candidates (Jennifer Balkcom, Kevin Kiley, Matt Little) have no row in the frozen finance file — one has no federal committee at all, and two have committees whose rows the August bulk files misfiled, a mismatch documented in the master coding and expected to self-heal at the October refresh. And 64 candidates have never served in Congress, of whom the FEC codes 48 as challengers and 16 as open-seat candidates. One further thing the chart does not show: at least one plotted candidate had suspended his campaign before the roster froze and is expected to withdraw formally before the general. He stays in every count here, because the roster is frozen and removing people after the fact is how a denominator gets chosen — but a reader counting live candidates should know the number is a frozen roster, not a current one. That 48 is not the 48 this page uses everywhere else for the candidates who have served: the two are equal by coincidence and count different people.

What anchors the dates, and what does not. The frozen record — every register version, the specifications, the rulebooks, the frozen data files — is hashed into commitment files and timestamped onto the Bitcoin blockchain through four independent OpenTimestamps calendars. One thing is NOT in the frozen record and should be said plainly: the coded index itself (the master coding and the roster it produces, which G1, G2, C2, E1 and E3 are scored on) lives in the code commitment, which re-anchors at each release, and it was amended after the register froze, in the scored legal column four times: two on 30 August (Tijerina’s cell corrected, declared on the G2 card; Brooks’s recoded from a judgment to a pending suit on a re-read of the outlet’s own wording, declared in the source log and here), two on 2 September, both on the G2 card (Leonard recoded to match the rule applied to two Republicans; the G2 count made to follow the spec’s enumeration, which drops Brooks’s pending suit from the basis) — G2’s verdict unmoved by any of them at p = 1.00; and in labels and free text no prediction is scored on (the 30 August re-tiering of 30 rows, the Costa/Perry same-day round-trip through “not established,” three tier relabels on 2 September, notes struck from 8 rows — all in the sourcing paragraph below). All 4 commitments are confirmed on the Bitcoin blockchain. The frozen-record commitments — the registers, the specifications, the data — are anchored in Bitcoin blocks 964701, 964709 and 964789 (all 2026-08-30), and none of them may ever anchor to a block mined after the freeze; the code commitment re-anchors at each release, and a build whose code is not yet anchored is a labelled DEV build, never a published one. The binding specification is on its third version — v2 froze a figure its own register had corrected, misdated its own precedent, and changed five rules from v1 with no changelog; v3 corrects the false statements, changes no scoring map, threshold, population or confidence, and carries the changelogs, with v1 and v2 standing verbatim beside it. The proof is verified, not read. A hostile reviewer showed that reading the proof’s own header proves nothing — it forged it with one command, rewrote a registered threshold, and passed the build. A second showed the block-pin alone is movable, since the file naming the expected block is editable by the same hand that rewrites the content. So the gate walks each proof from the file’s digest to the merkle root of the named Bitcoin block, checks that root against two independent public explorers, requires the block pinned in ANCHORS.txt, and requires that block to have been mined on or before the freeze — which the movable pin cannot dodge, because rewritten content can only be anchored to a block that postdates the rewrite (a fresh proof over rewritten content is still a valid proof — a timestamp establishes a date, and a fresh one establishes none), and reads every threshold off every card back into the frozen registers, so the number a reader sees cannot drift from the number that was registered. Its attacks, and every other one that ever landed, are replayed on every build. What this proves is narrower than it sounds. It proves the files had exactly this content on 30 August — not that they existed on the earlier dates written inside them. The register was frozen on 24 August and anchored six days later; that gap is my fault. What it establishes is the property a registered prediction actually needs: the record was fixed by a third party before the 31 October scoring date, before any outcome was known. Anyone can verify it without trusting this page — ots verify, or the published ots_verify.py, which does the walk in a hundred lines of standard-library Python.

Five things you cannot check from this page alone, said here rather than left to be noticed. The reviewers were AI sessions, not people, and this page used to let you assume otherwise. Nearly every correction on this page was found by blind review — rounds of four, four and then about ten, across reproducibility, statistics, fact-checking, red-teaming, accessibility, law and editing. Every one of those reviews, and the blind inter-coder replication behind the agreement figure, was performed by an AI model session that I ran, under a written role brief, with no sight of the other sessions’ work. Where this page says “a media lawyer,” “a statistician,” “a red-teamer,” “an independent coder,” it names the brief the session was given, not a credential anyone holds; no lawyer, statistician or second human coder has reviewed this page. Blind, adversarial and separate they were; independent of me they were not. Corrected 2026-09-02: a sixth blind review pointed out that the page’s own source log records these as “adversarial subagent reviews” while the page called them “an independent media lawyer” and “an independent coder” — which reads as human, and was false as read. Where a correction on this page names no finder, a review session found it, not me. Their transcripts are not published and nothing in the manifest timestamps them; everything they found is checkable — the scripts run, the figures recompute — but the account of how it was found rests on my word. The source ladder had 30 rows resting on a wiki, and closing them is what this update did. Every recorded claim carries a tier: 1 for a court or government record, 2 for a named independent outlet — national or regional, the label reads “major outlet” and several adverse rows rest on a regional or local paper, a local broadcaster, a university paper or a named contributor at a national outlet, which the ladder’s own written wording would call tier 3; the label is looser than the rulebook and this sentence is the disclosure — down to 4 for a single or partisan source. A hostile red-teamer, and then a media lawyer, established that 30 of the 115 rows claimed tier 1 or 2 while the only source actually cited was a wiki — a tertiary source that may itself cite good ones, but does not show them. The worst of them claimed tier 1, a court record, on an encyclopedia; one recorded a criminal conviction against a named sitting member beside the index’s own note that it could not support the claim. That is fixed, three ways, and the counts below are computed from the roster so this account cannot itself go stale. The 2 rows that carried a finding were verified, not softened: both were re-sourced off Wikipedia to named contemporary outlets — a forty-year-old misdemeanor plea to three named papers, a pair of court matters to the Patriot-News, CNN and the Washington Post (the outlets’ reporting of the rulings, not the court records themselves, which is why those rows now carry the outlet tier rather than the court-record tier) — and they now carry their findings and sit in the legal statistic, where an unsupported claim never should have. 5 rows whose only independent source is the member’s own official House biography — a government record, legitimate for a non-adverse biographical fact and for nothing adverse — are labelled as exactly that. The remaining 23 rows, cited to nothing but a wiki, were downgraded to a tertiary tier: no adverse coded value rests on any of them — a reported matter noted in a row’s free text with only a wiki behind it is labelled tertiary and enters no statistic, and on 2 September such notes were struck from 8 rows under that rule (the master coding, phase2.py in the bundle, marks which) — and the ladder now says tertiary rather than claiming a rung it cannot show. That is the 30 accounted for (2 verified, 5 to House-bio records, 23 downgraded), and the tighter gate then caught 1 further row it had been passing — a court record, reclassified. Then, on 2 September, the final pre-publication review found three rows labelled “court record” whose citations were news outlets or the member’s own page and no court record at all: two adverse rows (Perry, Van Orden) now carry the outlet tier their reporting supports, and 1 non-adverse row (Fitzpatrick) now says what it rests on, his own House.gov page — and the gate now refuses the court-record label to any row that cites no court or government record. That makes 7 rows labelled “Government record” (10 tier-1 rows in all, counting the 3 that cite a court or enforcement record for an adverse matter). The result is 0 rows outstanding, and the tier gate now enforces the rule by claim type — a bare biographical fact may rest on a government record, an adverse finding may not, it needs a court record or a named outlet on an accepted list — so a row cannot reintroduce the defect without failing the build. What that gate does and does not do: it checks that an adverse row cites an accepted host; it does not read the article, so a citation to the right outlet about a different matter would pass it. The articles were read by hand for the rows that carry a finding; the gate is a tripwire, not that reading. The earlier version of this paragraph, which called all of it “open work,” is preserved in the source log. A third party comparison was measured and never published. The index codes three variables that can be split by party: military service, recorded legal record, and occupational category. Two of them carry predictions and appear on this page. The third does not. Under a rubric frozen before it was applied, care-profession careers split 10 Democrats to 3 Republicans — Fisher odds ratio? 3.76, p = 0.074, the largest party asymmetry in any coded variable in this project by odds ratio, larger than the veteran gap that opens G1 on that measure — though in plain percentage points, the form the glossary calls the plainest, it is 17.5% against 5.4%, and the veteran gap of 13.6 points is the wider of the two. It sits at the project’s own “suggestive, not established” tier and appears on no page. The same run refuted the Republican-leaning half of the same rubric that I had expected to hold: uniformed-service careers split 9 Republicans to 6 Democrats, odds ratio 1.63 at p = 0.42 — nothing. Neither is scored and neither changes a verdict; both are printed here because a page whose rule is print-what-you-compute had computed three party splits and printed two. And those figures were typed literals until this build, and stale. The cross-tab was coded when the roster read 57 Democrats and 57 Republicans; Kevin Kiley was afterwards corrected to independent, so the Republican denominator is 56, not 57. Recomputed: the care odds ratio moves from 3.83 to 3.76, and the uniformed half — the larger of the two moves, and the one an earlier wording of this item omitted — from 1.59 (p = 0.58) to 1.63 (p = 0.42). Neither move changes a finding: the care lean stays suggestive-not-established, the uniformed claim stays refuted. The coding itself is not redone — re-coding 113 candidates after seeing the result is the failure the blind freeze exists to prevent, and it is what wrecked the uniformed half of this same rubric. Frozen coding, scripted arithmetic, declared correction, all four figures published. And a third of the money data is measured over a different window. The finance file is frozen, but only 77 of its 112 committees report through 30 June 2026; the rest run anywhere from 16 June 2026 to 18 August 2026. Corporate share is a share of receipts, and receipts accumulate, so a candidate measured over a longer window is not being measured on the same basis as one measured over a shorter. I do not know how much that moves the numbers. It is a live limitation, not a resolved one, and it is the most likely place for the next defect.

Nobody in this index was approached for comment, and that is a limitation rather than a policy I am pleased with. The index records only matters already documented in public sources and does not originate an allegation about anyone; every adverse entry carries its own bound and its own outcome in the same cell, and where a subject’s own response exists in the cited reporting it is recorded there too. That is an explanation, not a substitute for asking. Two things follow from saying it plainly. A row whose sourcing this index cannot stand behind no longer carries an adverse coded value at all: it reads “not established”, keeps its full detail, and is excluded from every computed statistic — which removed two rows, one from each party, on 30 August, for the hours it took to re-source them; both were verified to named outlets the same day and stand in the basis now, so at this build no row reads “not established.” And any subject who disputes an entry can say so; a sustained challenge is published as an erratum beside the original, never as a silent edit, which is the same procedure this project already runs for its own numbers. Raised by a review session run under a media-law brief, which found that neither page said any of this.

Check it yourself

Everything this page asserts about its own rigour is checkable, because the underlying material is published. Download the folder and run sha256sum -c MANIFEST.sha256; every digest below should match. What that does and does not prove: a matching digest shows a file is unchanged since I published it, which is worth something and is not the same as showing it is the file the numbers came from. For the frozen record — the register versions, the specifications, the rulebooks and the frozen data — the stronger check is ots verify COMMITMENT.txt.ots, which is attested by a third party rather than by me.

The table below lists the files this page's claims rest on. It is not the whole bundle: the manifest also covers this page and its companion, the chart, the stylesheet and the reproduction instructions, none of which can usefully print its own digest here. The manifest is the complete list.

The files this page's claims rest on, with their SHA-256 digests. The manifest covers the whole bundle.
FileSHA-256
PREDICTIONS_registered_2026-08-22.md
v1 as first registered. 18 predictions, 16 published. Never edited.
342130288fc65b9ad282fa3299a9593e7058eca63318310b9a856396d9a2268c
PREDICTIONS_v2_2026-08-22.md
v2, after four blind reviews. 19 scored. Never edited.
d0cebea416e4a0677468a5113a10c378777a8f8ebeaa15266c384da50e81c065
PREDICTIONS_v3_2026-08-23.md
v3, after the A-block defect. 25 scored. Never edited.
de5776eda4bd135a176558e2d922bc7217d6f089381f012d696ac9bc8df2f911
PREDICTIONS_v4_2026-08-23.md
v4, after the second review round. The four bad bets frozen, corrections named. Never edited.
b1822bbe99af265260565b43be8b55c37a4f47289ee4fd932d55c713213642e9
PREDICTIONS_v5_2026-08-24.md
v5, after ten blind reviews. The universe, instrument and baseline rulings. Binding.
b94874b647d62f1ca8f9e3dc07f703cd1d394aa8b527192f108f58cc3aae8aa3
DEFECT_A_BLOCK_2026-08-23.md
The defect report: what reproduced, what didn't, and why. ERRATUM (2026-09-02, the file is frozen and cannot be edited): its table headed 'reproduce to two decimals' contains two rows that do not — the Democratic incumbent mean, published 13.2%, is shown against 12.22%, which is the prior-service figure, not the FEC-code figure of 13.21% that matches; and the challenger mean 0.3% against 0.26%. The challenger figure v4 corrected; the incumbent figure v5 corrected, in its vintage caveat; the table's heading overstated both.
5d76d977b93643bf977d9d6ebd70bc4822a29458c541f1beb3a239bfd15cf820
ERRATUM_EVIDENTIARY_STANDARD_2026-08-24.md
Erratum: the rulebook's Tijerina example went stale; the rule it illustrates did not.
300314e7cf54e170c25566519f2d4584f7bc985d5cbd890fbefc2bea4cb1d75a
LICENSE.md
Code MIT, data and text CC BY 4.0, FEC source data public domain.
e67f3b07ad461be00b5acfda824131a28d4428f823c92bdd73a347c67613e4a8
ANALYSIS_SPEC_v1.md
The first analysis specification. Bound nineteen predictions; superseded. Never edited.
35e03eceb02a23500ae6c8576a31323b275ae627495371cae4aeb22f78f9614b
ANALYSIS_SPEC_v2_2026-08-30.md
The second specification; superseded the same day it was issued, and never edited. Independent review found it froze one figure its own register had corrected, misdated its own case-law precedent, and changed five rules from v1 without a changelog. All three failures are enumerated in v3.
f5f6d8830cc0061ad98c216ddd93ca35b7347b4e0429a616f5a6e28cb8ed4e98
ANALYSIS_SPEC_v3_2026-08-30.md
THE BINDING SPECIFICATION: population, estimator, exclusions and a total scoring map for every one of the 25. Changes no scoring map, threshold, population or confidence from v2 — it corrects v2's false statements and carries the two changelogs this series always owed (what v2 silently changed from v1, and every v2-to-v3 delta). Frozen and separately timestamped (COMMITMENT-4).
ac429f2f655969ffaf3effb159b9af91bf335e1d33dadb5c6b1ea06184c10266
EVIDENTIARY_STANDARD_v1.1.md
The coding rulebook the index was built under.
ac0cae0df4fae687c952e70186dbdc076f36fe43e10a3840ecb15a44a7fa9751
STYLE_PROTOCOL.md
The writing-measurement protocol, with its own changelog.
9ee3692c32acf0e55b7760063db92f7eef8bd59eab42ccb6566c25b6fe6ba123
districts_59.csv
The 59 districts and the selection rule. The most politically loaded choice here.
9edff9660e4baf747c9a8121c2dcefcb02ac5eed211f2d688a9bea5fd92fc476
roster_115.csv
All 115 candidates.
d5e8502a825ed62fbfdc48090ea896a52e3cfc9b73a8b0427b09bf2412ae7e21
money_analysis.py
The A-block estimator. Every number in section A comes from running this.
607ba3e32573b64c53b583b1c786d2912cf6785133f81157e5bd5acb16ee4f43
make_figures.py
The chart above, generated from the same data.
2c9c6b8840cfa2d763618f0ed1c85abf0bf4f98fdc61a1e567904bff15aaef84
COMMITMENT.txt
The frozen record and its digests, as submitted for timestamping. Anchored 2026-08-30 — later than the dates inside the documents, which is stated in the file.
174b228a149cc4cda08586b58df02e7271b42789e59c7b89665f06b53b264fcc
COMMITMENT-2.txt
A second commitment, for a record document published after the first. Editing a timestamped file destroys its attestation, so a later addition gets a new commitment rather than an amended one.
8c17794387eed6a8460461ee886543d3e6251d8ca336f59ad813928801e200e0
COMMITMENT-2.txt.ots
Its attestation — anchored in Bitcoin block 964709, mined 2026-08-30 08:08:56 UTC. Verify with: ots verify COMMITMENT-2.txt.ots
0caf245666edfaaf3649685ccc0fe8b3b47d54e8ca4ed6e28650445ba8f2cb77
COMMITMENT-3.txt
The CODE commitment: every script that computes or generates anything here, frozen as at this release. Added 2026-08-30 after a red-teamer overrode one figure inside figures.py — four lines — and turned a declared loss into a hit with every other check still passing, because the scripts were excluded from the first commitment as living code.
643418c9d26a04760692ffae4ea1c9bf712f0d24239e8a00982cd38b985de7b6
COMMITMENT-3.txt.ots
Its attestation — anchored in Bitcoin block 965121, mined 2026-09-02 04:58:21 UTC. Verify with: ots verify COMMITMENT-3.txt.ots The code anchor re-pins at every release.
e04b75d23ce77c2822696cf19f01f8375f91e4fe3b25d5a27360546b8cfa3dd2
COMMITMENT-4.txt
The spec-v3 commitment: the binding scoring specification, frozen the day it was issued. A new commitment rather than an edit to an old one, because a frozen file is never edited and a commitment is never regenerated — either act destroys the attestation.
c71be01e5e41430afd55353be7f30911315036ce9b60dfcb63b45dee5aaaabf8
COMMITMENT-4.txt.ots
Its attestation — anchored in Bitcoin block 964789, mined 2026-08-30 21:35:13 UTC. Verify with: ots verify COMMITMENT-4.txt.ots
7f6d6e7446c1e9f1b5923e049f9f786e47609c62d0560658d89022dc6c6850cd
ANCHORS.txt
What each proof is REQUIRED to prove: the exact Bitcoin block, its hash and its time. Without it the gate accepts any valid proof, including one made moments ago over rewritten content.
322464f8b4b1eb0a62a33bfefac7970d80000843efc00eb7829f7e872da9c7a3
ots_verify.py
Walks an OpenTimestamps proof from a file's digest to the merkle root of the Bitcoin block it names, and checks that root against two independent explorers. About a hundred lines, standard library only. This is what makes the commitment gate more than a reader of its own input.
cb3177abb26beca38d6ed204848f33d7e5efe4d1722e1a57d70df956d304de62
COMMITMENT.txt.ots
Its attestation — anchored in Bitcoin block 964701, mined 2026-08-30 06:59:54 UTC. Verify with: ots verify COMMITMENT.txt.ots Submitted to four independent Bitcoin calendars.
09851041dcbabb9b8f9a126339f038ae01b00b38e1189dde6567e160a229cbe3
VALIDATION_RESULTS.md
The blind inter-coder replication (46 of 50, alpha 0.91) and the occupational split, including the Republican-leaning claim it refuted. Published 2026-08-30 after a reviewer found both were cited on the page and available nowhere.
8036da0d60e012e00d3dfc4b6a41ae06a71ec5db5ad72759dc6fd6e13af9c5d5
privacy_gate.py
A tripwire for one pattern: a two-word name near private-life language in the Source Log and the page generator. NOT a guarantee — it names five classes it cannot see, in its own text. Added 2026-08-30 after a reviewer found four private individuals in the Source Log.
159eacad372c84114cdd1ea2296abd04c898ae37ad9f57ed99c061231c102243
faith_gate.py
Refuses a faith cell filled on grounds C4 excludes. It found one the reviewers had only called soft, and a second when a red-teamer removed the blanket veto that had been letting one admissible-looking word disarm all four exclusion rules.
1199b4b06fd18bf4a3220c59a3e9d2350d77a95719dd6c5bdf68ce9070c72f1a
tier_gate.py
Refuses a tier 1 or 2 claim whose citations are all tertiary, self-published, or absent. Rewritten 2026-08-30 after a red-teamer walked past it with a semicolon: it used to split the citation field on ';' and count every fragment as a source, which certified six rows on the candidate's own campaign site. The declared count is computed here, not typed.
1f0cc1f62044475d6037491b18e18ae4a437554a67ca10324f563774e9737d2c
register_gate.py
Cross-checks all 25 confidences AND every threshold on every card against the five frozen register versions. Extended 2026-08-30: it used to check confidences only, so a threshold could be rewritten freely; and a card it failed to parse counted as agreement, which one leading space was enough to arrange.
6e60cd092d61df3c98918d9a820aab891e0fa1fdefcbc14b8da37ebeb79b3023
commitment_gate.py
Walks each OpenTimestamps proof from the file's own digest to the Bitcoin block it names, and checks that block's merkle root against two independent explorers that must agree. The anchor it must reach is pinned in ANCHORS.txt, so a proof re-made over rewritten content lands elsewhere and fails. Rewritten 2026-08-30: it used to READ the digest out of the proof's header, and a red-teamer forged that with one sed.
810a005f8348e11ad4f40c6bd744d3e19f5c7c4c8f0e7ef7a2978fcab10ce2db
style_analysis.py
The writing measure: Flesch-Kincaid as the B-block predictions are scored on, including the terminal-punctuation fix declared on B4.
bf71a8a8f2462ff16e1d8465bea9d7b27b0fd41c3670b9b85fb6773638f7934f
style_corpus_all.csv
The frozen writing corpus every reading-grade figure is computed from. Added 2026-08-30 after a reviewer found the bundle would not run without it.
359e36ec288e948b2122891f7a0e11332538d8f17f6c4340a1460fea934c1344
figures.py
Every number on both pages except the A-block estimator's own: the reconstruction table, the calibration intervals, the leverage diagnostics. Named in the text; now published with it.
308d4cac311eba3eacfeb171226300e9c5e6b5f803a19c95d7c9dd1acf703484
make_pages.py
The page generator. It refuses to build if a figure is missing, a digest is stale, an anchor is dead, a retracted diagnostic reappears in any rendering, a decimal appears in prose that no figure resolves to, a confidence stated in a sentence is not in the register, or a headline changes without its URL. It does NOT catch every typed integer: a bare number under 100 reads as a count. That gap is real and is why the figures cited to published records are required to be wired to them.
97de6262eda0db2438e608761e94a27370326a44b03b80b4d39a27adde0cfdd6
nominees.py
The 59 districts and their nominees, as coded.
dfe41eb3a45cf5668ba37b5ecee10f0f2f7baec8f19394c73df6e3b187565a85
phase2.py
Ages, tenure and the Source Log.
9cd9471dada3d145033ad178b1909853006fa793cd4fcb29f4db5aa0f64a9e06
finance_nominees.csv
The frozen FEC extract every money figure is computed from.
971e490a34b8ce528bac8ae311a1fca2046bbf6ca01f9d69cd77c61093c36ec3
a11y_audit.py
The contrast audit. It resolves every colour rule through the site tokens and measures it.
a7f49f13ab9eaf35d70af3d31549fa5d508d3b9741de0ea9f347fe642906436c
A_BLOCK_OUTPUT_2026-08-23.txt
What that script prints on the frozen data. Corrected 2026-08-30: this row used to say "what it printed on 2026-08-23". The file is regenerated at every build, so it is today's run on August's frozen inputs — which proves the estimator is deterministic, not that a past state was preserved.
530158a00cf3333b6215291090a18fb26b78038c3f696cb2d793aff866d6d009
withheld_hashes.txt
Hashes of the two withheld predictions. The withholding itself failed — both texts appear in the published v1; see the correction above.
2a3299ea55d4f4e344afc5130db03d9362b4478c54fa6d58761e6013c1879500

The two withheld predictions are hashed in withheld_hashes.txt: D2 bddfe9698eda4d3674a7267aba4da6ce1657b23f04bdcb9602a3383f365822fa and E4 ea4e874b6fe95a8fb8d55886b2a478f5058a32a701d5ad1b0a231c19669d1132.

Method. 115 candidates across 59 competitive districts — of which 2 (NH-01 and NH-02) carry no candidate, because their primaries fall after the 2026-08-13 roster freeze, so the coded set spans 57 districts. Every count on these pages that reads “of 115” is against the 115 actually coded. Coded under a frozen written standard and checked by a separate blind AI coding session given only the rulebook — whose four disagreements are published by name in VALIDATION_RESULTS.md, though not the full 50-decision sheet — who reproduced 46 of 50 sampled coding decisions (92%), against rulebook v1.0 — a sample of about ten candidates, not the whole index, and faith carried three of the four disagreements. Campaign finance from FEC bulk filings frozen 2026-08-21, most coverage through 30 June 2026 (full range 16 June 2026 to 18 August 2026); every figure in section A is produced by money_analysis.py, and any section-A number published without a run of that file is a defect and gets logged as one. Reading grade is Flesch-Kincaid. Nothing on this page evaluates any person’s competence, character or fitness for office — the communication-style measures describe prose, and nothing more.

25 predictions, mean stated confidence 73.80%, computed rather than estimated. Scored 2026-10-31. This page will be updated in place; the predictions will not be edited.