Method & corrections · 2026 House Battleground Index
The Headline I Never Re-Ran
I sat down to write one sentence. The strongest finding this project has produced — corporate money rising with years in Congress, at a correlation? of 0.82 — turned out to be 0.576 when I computed it from the frozen data for the first time — or 0.712 if I score it on the districts that actually meet my own printed selection rule, which three of the 59 do not. I am confessing against 0.576, because that is the universe the bets were priced on and I ruled on that in writing on 24 August, before October could move the figures. Both belong in the first paragraph: leading with the worse number alone would be its own kind of dishonesty. Among those with congressional service, 0.311 on the published universe and 0.499 rule-compliant (I published 0.70 for this in August; both of these are what the frozen data actually gives). Here is what went wrong, what it cost, and — after three days of asking the question with the wrong instrument — what the replacement finding turned out to be.
Before any of it: the reviewers in this piece are anonymous. Their reports are not published and nothing timestamps them. Every turn in the story below is performed by them — they found the error, they checked the replacement, they caught the correction of the correction — and while everything they found is checkable, because the scripts run and the figures recompute, the account of how it was found rests on my word. The companion page says the same thing; it belongs here too, because this is the page that needs them.
This is the follow-up to Twenty-Five Bets, Four Lost and One Retracted, which holds the predictions themselves and does not need this piece to make sense. If you only read one, read that one. Every statistical word used here — correlation, Spearman, universe, confidence interval — is defined in plain English in the glossary on that page, and the “?” marks below link straight to the right entry.
The sentence was supposed to open a page about predictions. Four review sessions had reviewed that page, and the complaint that stuck was that hundreds of words on the philosophy of pre-registration passed before the first fact appeared. Lead with the finding, the editor said. (The round-one review texts were not preserved — a record-keeping failure in its own right; later review rounds are archived in full — so that sentence is memory, and marked as memory.) So the plan was to open with the finding and put the philosophy underneath it, where it belongs.
The finding was this. Across the 2026 House battleground, the share of a candidate’s money that comes from corporate and trade PACs rises with the years they have spent in Congress, at a correlation of 0.82. Among candidates with prior congressional service — restricting to those who have served, which is not the same as removing the candidates at zero: 37 of the 112 take no corporate money at all, and all of them are people who have never served — it still held at 0.70. I had published that figure in five separate documents since August. It was the headline of the whole project.
To write the sentence properly I needed the number to two decimals. So I computed it from the frozen data files, under the frozen specification, which — and this is the part that matters — I had never actually done.
It came back 0.576. And among those with congressional service, 0.311.
What I checked before I believed it
The first assumption when a number moves is that you broke something. So I checked everything the original claim had published alongside the correlation. Corrected 2026-08-30: this paragraph used to say those figures “come from the same rows” as the correlation. They do not. The correlations below are recomputed on the 92 candidates coded in August; these four averages are computed on all 112 under August's grouping. So they show the money and the coding are intact today — a real check, and a weaker one than the sentence claimed. A reviewer found it, the code comment recorded it, and nobody carried it to the page, which is this project's signature failure happening inside its own correction.
| Published in August | Recomputed, August’s own definitions | |
|---|---|---|
| Challenger average corporate share | 0.3% | 0.30% |
| Open-seat average | 1.3% | 1.28% |
| Republican incumbent average | 12.6% | 12.58% |
| Democratic incumbent average | 13.2% | 13.21% |
All four reproduce — under the definitions August used, which group candidates by their FEC incumbency code. The frozen specification later switched to grouping by prior service, which moves two candidates between groups and shifts these means slightly. Corrected 2026-08-30: this sentence used to say both sets of figures were published. They were not. Only the August grouping was, and it is the one in which Democrats look worse. Both are computed by the published script, so here they are. Under the August FEC-code grouping the incumbent averages are 12.58% Republican against 13.21% Democratic. Under the definition that actually binds every scored prediction — prior service — they are 12.58% Republican against 12.22% Democratic, and the ordering reverses. Neither gap is large and no prediction turns on either; what matters is that the page was printing one of two computed figures while claiming to print both, on the comparison most likely to be quoted by someone with a side. Then I checked the tenure figures by hand against the year each member was first elected — Kaptur 1982, Turner 2002, Costa 2004, Cuellar 2004, Perry 2012, Fitzpatrick 2016. All correct.
So the money is right and the years are right. Only the number connecting them is wrong.
| Published | Actual | |
|---|---|---|
| All candidates | r = 0.82 (n = 92) | r = 0.576 (n = 112) |
| Members already in Congress — and the two columns do not define that the same way | r = 0.70 (n = 38), on the FEC incumbency code | r = 0.311 (n = 48), on prior service (tenure > 0) |
| The same, within party | 0.81 D / 0.63 R, on the FEC code | 0.255 D / 0.430 R, on prior service |
| Three longest careers removed | 0.54 | 0.502 or 0.355 — two members tie for third-longest, so “the three longest” is ambiguous; both readings shown |
| Nine mid-decade-redistricting states removed | — (never published) | 0.684 (n = 82) — the script has always computed this and the page had not printed it. It is higher than the headline figure, which is the reason it appears here: a sensitivity that flatters me is the one I am least entitled to leave in the code |
Read the middle two rows carefully, because until 30 August they were mislabelled. They compared two different populations under one heading: the August figures were computed on the FEC’s incumbency code, which is how n came out at 38, and the frozen figures are computed on prior congressional service, which gives 48. Both halves are correct; putting them in one row called “prior congressional service only” was not, and it obscured part of the gap the table exists to show. The population change is one of the reasons the number moved, not a detail beside it — it is the same distinction that reverses which party’s incumbents average more corporate money, three paragraphs above. The like-for-like comparison, August grouping against August grouping, is the reconstruction table below, and it reproduces all seven August figures exactly.
Why it happened
Look at the sample sizes. Ninety-two candidates when the figure was computed; a hundred and twelve now. Twenty people arrived after the number was calculated, and the number was never calculated again.
That is possible for one reason: the figure was never in a saved script. It was computed once, in a working session, written down in prose, and then quoted from that prose in every document that followed. Prose cannot be re-run. So while the coding of the roster went on — batch after batch of new candidates verified and added — nothing recomputed, nothing failed, and nothing warned me. The number just sat there going quietly out of date while I built four predictions on top of it.
I know that is what happened because it is the second time. A day earlier I had caught the same thing with a different number — a measure of how often Donald Trump writes in capital letters, the single most striking statistic in the project, computed by a script nobody saved. Rebuilt under a written-down definition it came out 14.29 per hundred words, not the 11.62 I had published. Both of those are history and neither is reproducible here: that measure belongs to an earlier piece, its script is not in this bundle, and I am asking you to take the pair on my word — which is precisely the thing the rest of this page says not to do. I wrote at the time that the lesson was to save the script. I saved that one. I did not go looking for the others.
The first version of this piece stopped here and offered two hypotheses — the new people, or a mix-up between the two standard ways of measuring the relationship — and claimed I could not distinguish between them because the raw federal filings were gone. That claim was false, and reviewers proved it with my own records. The coding history names exactly which candidates were verified in the last two batches. Set those aside, restrict today’s published data to the 92 candidates that existed when the figure was computed, and every August number comes back:
| Published in August | Batches 1–8, recomputed today | |
|---|---|---|
| All candidates, Pearson | 0.82 | 0.8231 |
| All candidates, Spearman | 0.86 | 0.8587 |
| Served in Congress, Pearson | 0.70 | 0.6963 |
| Served in Congress, Spearman | 0.60 | 0.6004 |
| Three longest careers removed | 0.54 | 0.5412 |
| Within party, incumbents by FEC code | 0.81 D / 0.63 R | 0.8095 D / 0.6263 R |
Seven for seven, at n = 92 and 38 incumbents by FEC code — the exact counts the August record states. So the mechanism is settled: the district list never expanded; the coding roster completed. The figure was true of the 92 candidates coded when it was computed and was never recomputed as the last twenty arrived — among them Marcy Kaptur, 44 years of service and 8.4% corporate money, the single point that pulls the correlation down hardest. And the second hypothesis, the Pearson/Spearman? mix-up, is retracted by name: both statistics were recorded separately at the time, and both reproduce.
Which leaves the part that stings. On a page about not asserting things you have not checked, I asserted an uncertainty I had not checked. The answer was sitting in my own handoff record, one restriction away, the whole time I was writing “I cannot distinguish.” The reviewers did not have access to anything I lacked. They just looked.
What it costs
Four of the nineteen predictions the 22 August register ended with were priced against a number that was wrong.
- A1a predicted the prior-service correlation would hold at 0.60 or above. I put 92% on it. The data it was registered against says 0.311.
- A1b predicted the overall figure would land between 0.70 and 0.90, at 88%. It is 0.576.
- A3 predicted that controlling for committee chairmanships the relationship would survive at 0.50 or above, at 65%. You cannot get 0.50 out of a relationship that starts at 0.311 by controlling for something, except under a specific and unlikely statistical quirk.
- A4 predicted the same pattern in the 2024 cycle between 0.50 and 0.85, at 70% — mostly the same members, so mostly the same problem.
All four will miss. I am not changing them.
That distinction is the only interesting thing I have to say here, so let me be precise about it. The 0.82 was never a prediction. It was a claim about data I already had, and a claim about data in hand that no longer holds — stale, as the predictions page puts it, rather than false when made — is an error; you fix an error. The predictions built on top of it are forecasts, and a forecast you rewrite after learning it loses is not a forecast at all. It is the exact move this entire project was set up to make impossible.
So A1a, A1b, A3 and A4 stand word for word at 92%, 88%, 65% and 70%, and in October they will be marked missed, in the same table and the same denominator as everything else. No asterisk, no adjusted score printed helpfully alongside the real one. Then, separately, I have registered the corrected versions against the number that actually reproduces, and those are scored too. The honest arithmetic of what that costs me: freezing the four dead bets rather than deleting them is worth 0.0871 on the Brier score — a real penalty, and at most that large — it is the figure if every one of the remaining predictions hits, and it falls as they do not, so it is a ceiling and not a point. And it needs its base to mean anything: deleting the four would give a score of 0.0887, freezing them gives 0.1759, so this roughly doubles it. Lower is better. An earlier version of this sentence called it “exactly that large,” which was wrong in the direction that made the penalty sound bigger. An earlier version also claimed the mistake was charged twice; v4 retracted that claim as arithmetically backwards, and it stays retracted.
One more consequence. Among those nineteen was a prediction that I would turn out overconfident on average. I priced it at 60%. Knowing that four predictions are guaranteed misses before any data arrives, overconfidence is now close to arithmetically certain — that 60% has become free money. So I have frozen it and registered the honest price beside it: 93%. That price is fixed now whatever happens; for what it is worth, the simulation currently puts the chance at 93.6%, which is not a revision and does not replace it. Both predictions are scored. Corrected 2026-08-30: I used to call the gap below an advantage I was declining to take. I am not declining it — both predictions are scored, and the cheap one still lands in my denominator. What I declined was hiding it. The thirty-three-point gap is the size of the cheap hit I chose to disclose rather than delete.
The part I got wrong that I would rather have got wrong
One of the four review sessions was briefed as a statistician (an AI session, like every reviewer here — said plainly on the predictions page), and its main objection was something called zero-inflation?. 64 of the candidates have never held the seat, so their years in Congress are zero and their corporate PAC share is very close to zero too. It argued that a large overall correlation could be produced almost entirely by the gap between those people and those who have served — that what looked like a smooth climb might be a single step with a flat line on either side. It showed it in a simulation. I accepted the point in principle and made prior-service-only the primary measure.
It was not arguing in principle. Here is what the data does:
- 64 candidates who have never served: average corporate share 0.51%
- 48 with prior congressional service: average corporate share 12.40%
- Overall correlation 0.576; among those with prior service 0.311
I wrote at this point, in the first version of this piece: the relationship is mostly a step?, not a slope?. That was my replacement headline, and it lasted about a day. Reviewers pointed out that it is a claim about size tested with a correlation, and that the sizes say the opposite. The step is 11.89 percentage points?. The fitted line climbs 14.62 points across the observed range, which makes the slope look like the bigger of the two. Checked properly on 2026-08-30, and the check went against me twice. Only 1 of the 48 sits above 23.8 years, so I wrote that the line only outruns the step by being stretched out to where one person stands. That was the wrong test: it held the slope fixed and shortened the range, which measures nothing. The right test is the one this page already runs on that point — delete it and refit. Do that and the slope rises from 0.348 to 0.727 points a year and climbs 15.98 points across the shorter range: more than the step, not less. Marcy Kaptur was holding the slope down. So the first arithmetic was right, my correction of it was wrong, and both are printed here rather than the surviving one alone.
The correlation is low because the points scatter? 8.64 points either side of the line, not because the gradient is small.
The universe number is in the first paragraph now, and it belongs there. Three of the 59 districts — OH-09, TX-28, TX-34 — sit outside the selection rule? the frozen specification prints, and they were added as marquee races without the flag that should have accompanied them. Score the same data on the districts that actually satisfy the rule and the pooled correlation is 0.712 rather than 0.576, with 0.499 rather than 0.311 among those who have served — and under the register’s own ruling that flips A1b from a miss to a hit. So the honest size of this correction is 0.82 to 0.712 as much as it is 0.82 to 0.576. That universe is defensible, it is closer to the 0.82 I published, and it is not the one I am confessing against — because the bets were priced on the 59 as published, and the ruling fixing that was made in writing on 24 August, before October could move the figures. The full ruling is on the register.
So I have now been wrong about the shape of this relationship twice, and not in the same way. The first time I reached for a correlation to describe a magnitude. The second time, on 30 August, I tried to show that the slope only beat the step because the line ran out to one long career — and did it by shortening the range without refitting the line, which tests nothing. Refitted, the slope beats the step by more. My original arithmetic survived and my correction of it did not.
And then a third reviewer pointed out that I had spent three days arguing about a question I could have answered in two lines. Step or slope is not a matter of opinion to be settled by comparing correlations, which cannot separate two effects because they are one number. It is a regression? with two terms, on data that had been frozen for a week. Fitted: arriving in Congress at all is worth 9.41 percentage points (95% interval 4.46 to 14.36, robust p = 0.0003); each further year served is worth about 0.348 of a point, and whether that is “significant” turns on one person: on robust standard errors the slope’s interval (-0.437 to 1.133) includes zero (p = 0.39), and 89% of that uncertainty is Marcy Kaptur alone; without her it is 0.727 at p = 0.0017, and a bootstrap over everyone puts it above zero 99% of the time. The step is real and significant on every formula; the slope is probably positive, and one 44-year career decides how sure you get to be. Corrected 2026-09-01 and 2026-09-02: this passage first carried the classical p = 0.0011 and called the slope measurable; a first correction blamed the wider robust interval on the never-served “floor of zeros,” which was wrong — those candidates contribute nothing to it. The full account is on the predictions page. The step is far the larger for almost everyone: seniority only catches up after 27 years, and 1 of the 112 candidates in the file has served that long.
So there is a replacement finding, and it is not the one I published in August. The money arrives with the seat, not with the seniority. Years in Congress may add a little to it, at a rate that is small, uncertain, and nothing like the story the 0.82 told. To be plain about what is and is not new here: that corporate and trade PACs give overwhelmingly to sitting members is a decades-old, well-documented pattern, and a reader who knows the literature will rightly say the tilt itself is the baseline, not a finding. What this adds is the shape of it inside these 59 races — a step at arrival rather than a climb with seniority — and the fact that my August headline had the shape wrong. No causal claim is made from any of it: this is a cross-section of 112 candidates in one cycle with no controls, and it cannot tell whether serving attracts corporate money or whether candidates who attract corporate money are the ones who win and stay. The model is published, unregistered and unscored, because it was fitted on 30 August after the register froze. That took an editor, not a statistician, and it is the most useful thing anyone said to me about this project. Three days of careful argument with the wrong instrument is a worse failure than the original mistake, and no amount of timestamping repairs it.
The retracted claim is still on the register as A5 and is still scored, because my own rules forbid deleting a prediction once it is made. It will almost certainly be marked a hit. That is a fact about how weak a test I built, not about whether the claim was true, and the card says so.
What is actually fixed
Not the number — the number was only ever a symptom. What is fixed is that every
current figure in this piece is computed by scripts published with it — the
two historical figures above are not, and say so where they appear: money_analysis.py for the core
estimator — the correlations, the confidence interval?, the party breakdown, and a
hit-or-miss verdict against the two thresholds it was written for, A1a and A1b — the other twenty-three are scored by hand against the register in October — and figures.py for
everything else on this page, the reconstruction table above included. Run them and you get these
numbers. Run them in October and you get October’s.
And the rule that goes with it, which is the actual output of this week: a number that lives only in prose cannot be re-run, and a number that cannot be re-run will go stale without telling you. Any figure from this section that appears in anything I publish without a corresponding run of that file is a defect, and it gets logged as one.
Twice now I have found a headline statistic that no code could reproduce. Both times I found it by trying to use the number for something else, not by auditing. That is luck, and luck is not a method.
Putting the numbers in files is necessary and it is not sufficient, and I know that because it has already failed here. One of the writing predictions on the register was measured by a script that was written down, frozen, published and wrong: it joined text samples on a bare line break, so wherever a sample ended without a full stop two sentences merged into one, and the median it produced — 2.38 — was carried through a threshold the corrected figure of 3.54 sits the other side of. A file makes an error findable and repeatable. It does not make it false.
So the honest version of the rule is narrower than the one I wanted to end on. A number that lives only in prose cannot be checked at all. A number that lives in a file can be, by someone who runs it against something else and finds it does not fit. Both times, that someone was me, by accident. Neither time was it the method.
Published with this piece. The full verification, including everything that did
reproduce, is the defect report. The
binding register is v5; the four superseded
versions are published unedited beside it, with their hashes, so every correction can be checked
rather than taken on trust. The estimator is
money_analysis.py and the chart above is generated by
make_figures.py. Run
sha256sum -c MANIFEST.sha256 against
the published digests.
Campaign finance from FEC bulk filings frozen 2026-08-21, most coverage through 30 June 2026 (full range 16 June 2026 to 18 August 2026). Nothing here evaluates any
person’s competence, character or fitness for office.