Method & corrections · 2026 House Battleground Index

The Headline I Never Re-Ran

I sat down to write one sentence. The strongest finding this project has produced — corporate money rising with years in Congress, at a correlation? of 0.82 — turned out to be 0.576 when I computed it from the frozen data for the first time — or 0.712 if I score it on the districts that actually meet my own printed selection rule, which three of the 59 do not. I am confessing against 0.576, because that is the universe the bets were priced on and I ruled on that in writing on 24 August, before October could move the figures. Both belong in the first paragraph: leading with the worse number alone would be its own kind of dishonesty. Among those with congressional service, 0.311 on the published universe and 0.499 rule-compliant (I published 0.70 for this in August; both of these are what the frozen data actually gives). Here is what went wrong, what it cost, and — after three days of asking the question with the wrong instrument — what the replacement finding turned out to be.

Written 2026-08-23 Corrected 2026-08-30 — the step-versus-slope passage, and its own correction; corrected 2026-09-01 and 2026-09-02 — the slope’s inference, twice Every current figure below is computed by a script published with this piece; numbers quoted as history are labelled as history

The sentence was supposed to open a page about predictions. Four review sessions had reviewed that page, and the complaint that stuck was that hundreds of words on the philosophy of pre-registration passed before the first fact appeared. Lead with the finding, the editor said. (The round-one review texts were not preserved — a record-keeping failure in its own right; later review rounds are archived in full — so that sentence is memory, and marked as memory.) So the plan was to open with the finding and put the philosophy underneath it, where it belongs.

The finding was this. Across the 2026 House battleground, the share of a candidate’s money that comes from corporate and trade PACs rises with the years they have spent in Congress, at a correlation of 0.82. Among candidates with prior congressional service — restricting to those who have served, which is not the same as removing the candidates at zero: 37 of the 112 take no corporate money at all, and all of them are people who have never served — it still held at 0.70. I had published that figure in five separate documents since August. It was the headline of the whole project.

To write the sentence properly I needed the number to two decimals. So I computed it from the frozen data files, under the frozen specification, which — and this is the part that matters — I had never actually done.

It came back 0.576. And among those with congressional service, 0.311.

What I checked before I believed it

The first assumption when a number moves is that you broke something. So I checked everything the original claim had published alongside the correlation. Corrected 2026-08-30: this paragraph used to say those figures “come from the same rows” as the correlation. They do not. The correlations below are recomputed on the 92 candidates coded in August; these four averages are computed on all 112 under August's grouping. So they show the money and the coding are intact today — a real check, and a weaker one than the sentence claimed. A reviewer found it, the code comment recorded it, and nobody carried it to the page, which is this project's signature failure happening inside its own correction.

The four averages published in August beside the correlation, recomputed under the definitions August used
Published in AugustRecomputed, August’s own definitions
Challenger average corporate share0.3%0.30%
Open-seat average1.3%1.28%
Republican incumbent average12.6%12.58%
Democratic incumbent average13.2%13.21%

All four reproduce — under the definitions August used, which group candidates by their FEC incumbency code. The frozen specification later switched to grouping by prior service, which moves two candidates between groups and shifts these means slightly. Corrected 2026-08-30: this sentence used to say both sets of figures were published. They were not. Only the August grouping was, and it is the one in which Democrats look worse. Both are computed by the published script, so here they are. Under the August FEC-code grouping the incumbent averages are 12.58% Republican against 13.21% Democratic. Under the definition that actually binds every scored prediction — prior service — they are 12.58% Republican against 12.22% Democratic, and the ordering reverses. Neither gap is large and no prediction turns on either; what matters is that the page was printing one of two computed figures while claiming to print both, on the comparison most likely to be quoted by someone with a side. Then I checked the tenure figures by hand against the year each member was first elected — Kaptur 1982, Turner 2002, Costa 2004, Cuellar 2004, Perry 2012, Fitzpatrick 2016. All correct.

So the money is right and the years are right. Only the number connecting them is wrong.

The published correlation figures against what the frozen data gives
PublishedActual
All candidatesr = 0.82 (n = 92)r = 0.576 (n = 112)
Members already in Congress — and the two columns do not define that the same wayr = 0.70 (n = 38), on the FEC incumbency coder = 0.311 (n = 48), on prior service (tenure > 0)
The same, within party0.81 D / 0.63 R, on the FEC code0.255 D / 0.430 R, on prior service
Three longest careers removed0.540.502 or 0.355 — two members tie for third-longest, so “the three longest” is ambiguous; both readings shown
Nine mid-decade-redistricting states removed— (never published)0.684 (n = 82) — the script has always computed this and the page had not printed it. It is higher than the headline figure, which is the reason it appears here: a sensitivity that flatters me is the one I am least entitled to leave in the code

Why it happened

Look at the sample sizes. Ninety-two candidates when the figure was computed; a hundred and twelve now. Twenty people arrived after the number was calculated, and the number was never calculated again.

That is possible for one reason: the figure was never in a saved script. It was computed once, in a working session, written down in prose, and then quoted from that prose in every document that followed. Prose cannot be re-run. So while the coding of the roster went on — batch after batch of new candidates verified and added — nothing recomputed, nothing failed, and nothing warned me. The number just sat there going quietly out of date while I built four predictions on top of it.

I know that is what happened because it is the second time. A day earlier I had caught the same thing with a different number — a measure of how often Donald Trump writes in capital letters, the single most striking statistic in the project, computed by a script nobody saved. Rebuilt under a written-down definition it came out 14.29 per hundred words, not the 11.62 I had published. Both of those are history and neither is reproducible here: that measure belongs to an earlier piece, its script is not in this bundle, and I am asking you to take the pair on my word — which is precisely the thing the rest of this page says not to do. I wrote at the time that the lesson was to save the script. I saved that one. I did not go looking for the others.

The first version of this piece stopped here and offered two hypotheses — the new people, or a mix-up between the two standard ways of measuring the relationship — and claimed I could not distinguish between them because the raw federal filings were gone. That claim was false, and reviewers proved it with my own records. The coding history names exactly which candidates were verified in the last two batches. Set those aside, restrict today’s published data to the 92 candidates that existed when the figure was computed, and every August number comes back:

The seven August correlation figures, reproduced by restricting to the batches one through eight coding universe
Published in AugustBatches 1–8, recomputed today
All candidates, Pearson0.820.8231
All candidates, Spearman0.860.8587
Served in Congress, Pearson0.700.6963
Served in Congress, Spearman0.600.6004
Three longest careers removed0.540.5412
Within party, incumbents by FEC code0.81 D / 0.63 R0.8095 D / 0.6263 R

Seven for seven, at n = 92 and 38 incumbents by FEC code — the exact counts the August record states. So the mechanism is settled: the district list never expanded; the coding roster completed. The figure was true of the 92 candidates coded when it was computed and was never recomputed as the last twenty arrived — among them Marcy Kaptur, 44 years of service and 8.4% corporate money, the single point that pulls the correlation down hardest. And the second hypothesis, the Pearson/Spearman? mix-up, is retracted by name: both statistics were recorded separately at the time, and both reproduce.

Which leaves the part that stings. On a page about not asserting things you have not checked, I asserted an uncertainty I had not checked. The answer was sitting in my own handoff record, one restriction away, the whole time I was writing “I cannot distinguish.” The reviewers did not have access to anything I lacked. They just looked.

What it costs

Four of the nineteen predictions the 22 August register ended with were priced against a number that was wrong.

All four will miss. I am not changing them.

That distinction is the only interesting thing I have to say here, so let me be precise about it. The 0.82 was never a prediction. It was a claim about data I already had, and a claim about data in hand that no longer holds — stale, as the predictions page puts it, rather than false when made — is an error; you fix an error. The predictions built on top of it are forecasts, and a forecast you rewrite after learning it loses is not a forecast at all. It is the exact move this entire project was set up to make impossible.

So A1a, A1b, A3 and A4 stand word for word at 92%, 88%, 65% and 70%, and in October they will be marked missed, in the same table and the same denominator as everything else. No asterisk, no adjusted score printed helpfully alongside the real one. Then, separately, I have registered the corrected versions against the number that actually reproduces, and those are scored too. The honest arithmetic of what that costs me: freezing the four dead bets rather than deleting them is worth 0.0871 on the Brier score — a real penalty, and at most that large — it is the figure if every one of the remaining predictions hits, and it falls as they do not, so it is a ceiling and not a point. And it needs its base to mean anything: deleting the four would give a score of 0.0887, freezing them gives 0.1759, so this roughly doubles it. Lower is better. An earlier version of this sentence called it “exactly that large,” which was wrong in the direction that made the penalty sound bigger. An earlier version also claimed the mistake was charged twice; v4 retracted that claim as arithmetically backwards, and it stays retracted.

One more consequence. Among those nineteen was a prediction that I would turn out overconfident on average. I priced it at 60%. Knowing that four predictions are guaranteed misses before any data arrives, overconfidence is now close to arithmetically certain — that 60% has become free money. So I have frozen it and registered the honest price beside it: 93%. That price is fixed now whatever happens; for what it is worth, the simulation currently puts the chance at 93.6%, which is not a revision and does not replace it. Both predictions are scored. Corrected 2026-08-30: I used to call the gap below an advantage I was declining to take. I am not declining it — both predictions are scored, and the cheap one still lands in my denominator. What I declined was hiding it. The thirty-three-point gap is the size of the cheap hit I chose to disclose rather than delete.

The part I got wrong that I would rather have got wrong

One of the four review sessions was briefed as a statistician (an AI session, like every reviewer here — said plainly on the predictions page), and its main objection was something called zero-inflation?. 64 of the candidates have never held the seat, so their years in Congress are zero and their corporate PAC share is very close to zero too. It argued that a large overall correlation could be produced almost entirely by the gap between those people and those who have served — that what looked like a smooth climb might be a single step with a flat line on either side. It showed it in a simulation. I accepted the point in principle and made prior-service-only the primary measure.

It was not arguing in principle. Here is what the data does:

Corporate and trade PAC share against years in Congress Scatter plot of 112 battleground House candidates. The horizontal axis is years of congressional service, 0 to 45. The vertical axis is the share of campaign receipts coming from corporate and trade PACs, 0 to 40 percent. 64 candidates who have never served sit in a flat band at zero years, averaging 0.51 percent. 48 candidates with congressional service spread across the rest of the chart, averaging 12.40 percent, with an upward trend of correlation 0.31. The gap between the two group averages is 11.89 points; the fitted line climbs 14.62 points across the observed range; the points scatter widely around it. A two-term regression in the text finds both a step at arrival and a small, uncertain per-year slope, with the step much the larger. 0% 10% 20% 30% 40% 0 10 20 30 40 the step: 0.51% → 12.40%, never served vs served the slope: r = 0.31 among those who have served Matt Schultz (AK-00) — never served, 0.00% Bill Hill (AK-00) — never served, 0.00% Rhett Marques (AL-02) — never served, 6.45% Amish Shah (AZ-01) — never served, 1.45% Jay Feely (AZ-01) — never served, 0.83% Jonathan Nez (AZ-02) — never served, 0.00% JoAnna Mendoza (AZ-06) — never served, 0.07% Richard Pan (CA-06) — never served, 6.10% Kevin Lincoln (CA-13) — never served, 0.31% Kyle Kirkland (CA-21) — never served, 0.00% Randy Villegas (CA-22) — never served, 0.00% Chuong Vo (CA-45) — never served, 0.00% Marni Von Wilpert (CA-48) — never served, 0.32% Jim Desmond (CA-48) — never served, 0.33% Dwayne Romero (CO-03) — never served, 0.00% Jessica Killin (CO-05) — never served, 0.00% Manny Rutinel (CO-08) — never served, 0.00% Scott Singer (FL-25) — never served, 0.00% Christina Bohannan (IA-01) — never served, 0.07% Lindsay James (IA-02) — never served, 0.00% Joe Mitchell (IA-02) — never served, 2.48% Sarah Trone Garriott (IA-03) — never served, 0.00% Barb Regnitz (IN-01) — never served, 0.00% Matthew Dunlap (ME-02) — never served, 0.00% Paul LePage (ME-02) — never served, 0.10% Sean McCann (MI-04) — never served, 0.19% William Lawrence (MI-07) — never served, 0.00% Thomas Smith (MI-08) — never served, 0.00% Christina Hines (MI-10) — never served, 0.00% Michael Bouchard (MI-10) — never served, 0.00% Eric Pratt (MN-02) — never served, 2.45% Sam Forstag (MT-01) — never served, 0.00% Aaron Flint (MT-01) — never served, 4.57% Laurie Buckhout (NC-01) — never served, 0.71% Jamie Ager (NC-11) — never served, 0.02% Denise Powell (NE-02) — never served, 0.00% Brinker Harding (NE-02) — never served, 1.43% Rebecca Bennett (NJ-07) — never served, 0.00% Rosie Pino (NJ-09) — never served, 0.00% Gregory Cunningham (NM-02) — never served, 0.00% Carrie Buck (NV-01) — never served, 0.00% Martin O'Donnell (NV-03) — never served, 0.00% Cody Whipple (NV-04) — never served, 0.79% Christopher Gallant (NY-01) — never served, 0.00% Michael LiPetri (NY-03) — never served, 0.00% Jeanine Driscoll (NY-04) — never served, 0.00% Cait Conley (NY-17) — never served, 0.02% Peter Oberacker (NY-19) — never served, 0.00% Eric Conroy (OH-01) — never served, 0.94% Derek Merrin (OH-09) — never served, 0.33% Kristina Knickerbocker (OH-10) — never served, 0.00% Carey Coleman (OH-13) — never served, 0.00% Don Leonard (OH-15) — never served, 0.00% Bob Harvie (PA-01) — never served, 0.00% Bob Brooks (PA-07) — never served, 0.84% Paige Cognetti (PA-08) — never served, 0.02% Janelle Stelson (PA-10) — never served, 0.58% Tano Tijerina (TX-28) — never served, 0.00% Eric Flores (TX-34) — never served, 0.23% Shannon Taylor (VA-01) — never served, 0.00% Doug Ollivant (VA-07) — never served, 0.00% John Braun (WA-03) — never served, 1.11% Mitchell Berman (WI-01) — never served, 0.00% Rebecca Cooke (WI-03) — never served, 0.05% Nick Begich (AK-00) — 1.8 years, 5.81% Shomari Figures (AL-02) — 1.8 years, 32.53% Eli Crane (AZ-02) — 3.8 years, 0.13% Juan Ciscomani (AZ-06) — 3.8 years, 10.28% Adam Gray (CA-13) — 1.8 years, 12.97% Jim Costa (CA-21) — 21.8 years, 36.79% David Valadao (CA-22) — 11.8 years, 19.07% Derek Tran (CA-45) — 1.8 years, 5.43% Jeff Hurd (CO-03) — 1.8 years, 9.11% Jeff Crank (CO-05) — 1.8 years, 11.83% Gabe Evans (CO-08) — 1.8 years, 10.63% Jared Moskowitz (FL-25) — 3.8 years, 8.02% Mariannette Miller-Meeks (IA-01) — 5.8 years, 14.68% Zach Nunn (IA-03) — 3.8 years, 17.99% Frank Mrvan (IN-01) — 5.8 years, 15.86% Bill Huizenga (MI-04) — 15.8 years, 24.99% Tom Barrett (MI-07) — 1.8 years, 5.90% Kristen McDonald Rivet (MI-08) — 1.8 years, 10.06% Don Davis (NC-01) — 3.8 years, 20.02% Tom Kean Jr. (NJ-07) — 3.8 years, 9.70% Nellie Pou (NJ-09) — 1.8 years, 11.54% Gabriel Vasquez (NM-02) — 3.8 years, 8.97% Dina Titus (NV-01) — 15.8 years, 14.39% Susie Lee (NV-03) — 7.8 years, 11.52% Steven Horsford (NV-04) — 9.8 years, 27.06% Nicholas LaLota (NY-01) — 3.8 years, 7.93% Thomas Suozzi (NY-03) — 8.7 years, 13.79% Laura Gillen (NY-04) — 1.8 years, 2.47% Mike Lawler (NY-17) — 3.8 years, 8.70% Josh Riley (NY-19) — 1.8 years, 0.60% Greg Landsman (OH-01) — 3.8 years, 4.57% Marcy Kaptur (OH-09) — 43.8 years, 8.41% Mike Turner (OH-10) — 23.8 years, 21.24% Emilia Sykes (OH-13) — 3.8 years, 9.00% Mike Carey (OH-15) — 5.0 years, 39.09% Brian Fitzpatrick (PA-01) — 9.8 years, 16.25% Ryan Mackenzie (PA-07) — 1.8 years, 7.33% Rob Bresnahan (PA-08) — 1.8 years, 9.59% Scott Perry (PA-10) — 13.8 years, 1.88% Henry Cuellar (TX-28) — 21.8 years, 14.81% Vicente Gonzalez (TX-34) — 9.8 years, 19.56% Rob Wittman (VA-01) — 18.9 years, 19.86% Elaine Luria (VA-02) — 4.0 years, 0.11% Jen Kiggans (VA-02) — 3.8 years, 8.51% Eugene Vindman (VA-07) — 1.8 years, 2.02% Marie Gluesenkamp Perez (WA-03) — 3.8 years, 2.69% Bryan Steil (WI-01) — 7.8 years, 18.68% Derrick Van Orden (WI-03) — 3.8 years, 2.81% Marcy Kaptur, 44 yrs — kept in NEVER SERVED YEARS IN CONGRESS CORPORATE / TRADE PAC SHARE
112 of the 115 candidates in the index — 3 have no row in the frozen finance file — showing corporate and trade PAC share of receipts against years of congressional service, FEC bulk data frozen 2026-08-21 (77 of 112 committees reporting through 30 June 2026; coverage runs 16 June 2026 to 18 August 2026). The 64 who have never served average 0.51%; the 48 candidates who have served average 12.40%. Among them the upward trend is r = 0.31. Marcy Kaptur, 44 years and 8.4%, is the single highest-leverage point? in the chart; deleting her moves the incumbents-only correlation from 0.311 to 0.481, and the slope from 0.348 to 0.727. She is circled and kept in, because removing an inconvenient point is how you get a number that does not reproduce.

I wrote at this point, in the first version of this piece: the relationship is mostly a step?, not a slope?. That was my replacement headline, and it lasted about a day. Reviewers pointed out that it is a claim about size tested with a correlation, and that the sizes say the opposite. The step is 11.89 percentage points?. The fitted line climbs 14.62 points across the observed range, which makes the slope look like the bigger of the two. Checked properly on 2026-08-30, and the check went against me twice. Only 1 of the 48 sits above 23.8 years, so I wrote that the line only outruns the step by being stretched out to where one person stands. That was the wrong test: it held the slope fixed and shortened the range, which measures nothing. The right test is the one this page already runs on that point — delete it and refit. Do that and the slope rises from 0.348 to 0.727 points a year and climbs 15.98 points across the shorter range: more than the step, not less. Marcy Kaptur was holding the slope down. So the first arithmetic was right, my correction of it was wrong, and both are printed here rather than the surviving one alone.

The correlation is low because the points scatter? 8.64 points either side of the line, not because the gradient is small.

The universe number is in the first paragraph now, and it belongs there. Three of the 59 districts — OH-09, TX-28, TX-34 — sit outside the selection rule? the frozen specification prints, and they were added as marquee races without the flag that should have accompanied them. Score the same data on the districts that actually satisfy the rule and the pooled correlation is 0.712 rather than 0.576, with 0.499 rather than 0.311 among those who have served — and under the register’s own ruling that flips A1b from a miss to a hit. So the honest size of this correction is 0.82 to 0.712 as much as it is 0.82 to 0.576. That universe is defensible, it is closer to the 0.82 I published, and it is not the one I am confessing against — because the bets were priced on the 59 as published, and the ruling fixing that was made in writing on 24 August, before October could move the figures. The full ruling is on the register.

So I have now been wrong about the shape of this relationship twice, and not in the same way. The first time I reached for a correlation to describe a magnitude. The second time, on 30 August, I tried to show that the slope only beat the step because the line ran out to one long career — and did it by shortening the range without refitting the line, which tests nothing. Refitted, the slope beats the step by more. My original arithmetic survived and my correction of it did not.

And then a third reviewer pointed out that I had spent three days arguing about a question I could have answered in two lines. Step or slope is not a matter of opinion to be settled by comparing correlations, which cannot separate two effects because they are one number. It is a regression? with two terms, on data that had been frozen for a week. Fitted: arriving in Congress at all is worth 9.41 percentage points (95% interval 4.46 to 14.36, robust p = 0.0003); each further year served is worth about 0.348 of a point, and whether that is “significant” turns on one person: on robust standard errors the slope’s interval (-0.437 to 1.133) includes zero (p = 0.39), and 89% of that uncertainty is Marcy Kaptur alone; without her it is 0.727 at p = 0.0017, and a bootstrap over everyone puts it above zero 99% of the time. The step is real and significant on every formula; the slope is probably positive, and one 44-year career decides how sure you get to be. Corrected 2026-09-01 and 2026-09-02: this passage first carried the classical p = 0.0011 and called the slope measurable; a first correction blamed the wider robust interval on the never-served “floor of zeros,” which was wrong — those candidates contribute nothing to it. The full account is on the predictions page. The step is far the larger for almost everyone: seniority only catches up after 27 years, and 1 of the 112 candidates in the file has served that long.

So there is a replacement finding, and it is not the one I published in August. The money arrives with the seat, not with the seniority. Years in Congress may add a little to it, at a rate that is small, uncertain, and nothing like the story the 0.82 told. To be plain about what is and is not new here: that corporate and trade PACs give overwhelmingly to sitting members is a decades-old, well-documented pattern, and a reader who knows the literature will rightly say the tilt itself is the baseline, not a finding. What this adds is the shape of it inside these 59 races — a step at arrival rather than a climb with seniority — and the fact that my August headline had the shape wrong. No causal claim is made from any of it: this is a cross-section of 112 candidates in one cycle with no controls, and it cannot tell whether serving attracts corporate money or whether candidates who attract corporate money are the ones who win and stay. The model is published, unregistered and unscored, because it was fitted on 30 August after the register froze. That took an editor, not a statistician, and it is the most useful thing anyone said to me about this project. Three days of careful argument with the wrong instrument is a worse failure than the original mistake, and no amount of timestamping repairs it.

The retracted claim is still on the register as A5 and is still scored, because my own rules forbid deleting a prediction once it is made. It will almost certainly be marked a hit. That is a fact about how weak a test I built, not about whether the claim was true, and the card says so.

What is actually fixed

Not the number — the number was only ever a symptom. What is fixed is that every current figure in this piece is computed by scripts published with it — the two historical figures above are not, and say so where they appear: money_analysis.py for the core estimator — the correlations, the confidence interval?, the party breakdown, and a hit-or-miss verdict against the two thresholds it was written for, A1a and A1b — the other twenty-three are scored by hand against the register in October — and figures.py for everything else on this page, the reconstruction table above included. Run them and you get these numbers. Run them in October and you get October’s.

And the rule that goes with it, which is the actual output of this week: a number that lives only in prose cannot be re-run, and a number that cannot be re-run will go stale without telling you. Any figure from this section that appears in anything I publish without a corresponding run of that file is a defect, and it gets logged as one.

Twice now I have found a headline statistic that no code could reproduce. Both times I found it by trying to use the number for something else, not by auditing. That is luck, and luck is not a method.

Putting the numbers in files is necessary and it is not sufficient, and I know that because it has already failed here. One of the writing predictions on the register was measured by a script that was written down, frozen, published and wrong: it joined text samples on a bare line break, so wherever a sample ended without a full stop two sentences merged into one, and the median it produced — 2.38 — was carried through a threshold the corrected figure of 3.54 sits the other side of. A file makes an error findable and repeatable. It does not make it false.

So the honest version of the rule is narrower than the one I wanted to end on. A number that lives only in prose cannot be checked at all. A number that lives in a file can be, by someone who runs it against something else and finds it does not fit. Both times, that someone was me, by accident. Neither time was it the method.

Published with this piece. The full verification, including everything that did reproduce, is the defect report. The binding register is v5; the four superseded versions are published unedited beside it, with their hashes, so every correction can be checked rather than taken on trust. The estimator is money_analysis.py and the chart above is generated by make_figures.py. Run sha256sum -c MANIFEST.sha256 against the published digests.

Campaign finance from FEC bulk filings frozen 2026-08-21, most coverage through 30 June 2026 (full range 16 June 2026 to 18 August 2026). Nothing here evaluates any person’s competence, character or fitness for office.