# Validation Results — E3, E4, V2, V3
**Run 2026-08-21, on the completed 115-candidate sample · tests specified in CONCEPT_VALIDATION_PLAN.md · standard applied: EVIDENTIARY_STANDARD v1.0 (frozen)**

Four tests, four verdicts — two of them against our own concepts. That is the point: a validation plan that only confirms is not a validation plan.

---

## E3 — Stock-conduct cluster vs. chamber base rate: **NOT a battleground excess**

**Sub-coding first (per E3-b).** The six-member cluster is not one phenomenon. Member-era STOCK Act late-disclosure matters: **Lee, Suozzi** (Suozzi's dismissed — the dismissal travels). Member-era trading-conduct controversies with no located complaint: **Bresnahan, Moskowitz**. Candidate-era financial-disclosure timing (a different statute, resolved by filing): **Barrett, Harding**. Publishing "six-member stock cluster" without this split would overstate the phenomenon; the workbook notes already carry the distinctions.

**The base-rate test.** The comparable class — sitting members with STOCK-Act-type late-disclosure matters — is 2 of the index's 38 incumbents = **5.3%**. The chamber-wide benchmark: Insider's *Conflicted Congress* project identified **78 members (≈14.6% of 535)** with STOCK Act violations through late 2022 ([Insider, Levinthal & Hall, via Yahoo Finance](https://finance.yahoo.com/news/75-members-congress-violated-law-160918370.html)); typical penalty $200 or waived. Binomial test of 2/38 against 14.6%: two-sided p = 0.16 (expected count at base rate: 5.5; one-sided p for *below* base rate: 0.07).

**Verdict (the pre-committed honest outcome):** battleground incumbents are **not above** the chamber base rate — if anything the point estimate runs below it. "Battlegrounds have a stock problem" is not supported; **Congress is the story**, and the index's contribution is candidate-level documentation with outcomes attached. **Bound on this test (rule B2 applied to ourselves):** our 5.3% comes from general-press searches, not a systematic audit of every member's periodic transaction reports as Insider ran; our figure is a floor. The test therefore rules out "battlegrounds are worse" but cannot distinguish "same" from "better."

## E4 — Occupational asymmetry under the frozen rubric: **half refuted, half suggestive**

All 114 D/R candidates coded under frozen H1–H4 (most-recent-career-defining governs; H2 uniformed = career military + sworn civilian police/sheriff/fire/corrections/border, FBI/CIA excluded; H3 care = nursing/allied practice, teaching, social work, chaplaincy, direct-care nonprofit; prosecutors and physicians excluded as pre-committed):

| | DEM | REP |
|---|---|---|
| Uniformed-service career | 6 | 9 |
| Care-professions career | 10 | 3 |
| Other | 41 | 45 |

- **The "R uniformed-service cluster" claim FAILS.** 9R vs 6D, Fisher OR = 1.59, **p = 0.58**. Once the rubric was applied blind, the asymmetry dissolved: Democrats run career military too (Conley, Mendoza, Luria, Vindman) plus two career fire-service candidates (Brooks, career Bethlehem firefighter; Forstag, smokejumper) that H2's own "fire" clause pulls in. Several impressionistic "R uniformed" cases fell out under H1's most-recent-career rule (Crane → business; Flint → media; Fitzpatrick → FBI, excluded by H2's own pre-commitment). The original observation was partly an artifact of coding categories around the impression — exactly the failure mode the freeze existed to catch. **This claim should not be published.**
- **The care-professions lean survives as suggestive.** 10D vs 3R, Fisher OR = 3.83, **p = 0.074** — the index's own middle tier applies: *suggestive, not established*. Printable only with the p-value and the coding log shown. (Notable under honest coding: Kiggans, R, codes CARE — her most recent career is nurse practitioner.)
- Overall 2×3 chi-square: p = 0.10.

**Coding-log decisions carried with the result:** Buck (charter-school principal) coded care — school practice leadership, logged as borderline vs. H3's education-executive exclusion; Bohannan and Leonard (professors) coded care under H3's "higher-ed teaching" text; Landsman and McDonald Rivet (education *executives*) excluded per H3.

## V2 — Blind inter-coder replication: **α = 0.91 — "replicable standard" is earned, with one weak zone**

A fresh, blind coder (separate session; given ONLY the v1.0 rulebook and ten names; instructed to use only its own web research, no access to our entries) coded 5 fields × 10 stratified candidates — deliberately including the hard cases (O'Donnell, Kean, Perry, Suozzi, Wittman, the common-name Smith, Shah, Gonzalez, Buckhout, Mendoza).

**Result: 46/50 decisions agree (92%). Krippendorff's α, all 50 decisions pooled: 0.91** — above the pre-registered 0.80 bar. Per field: Military **1.00** (10/10, including Buckhout=Yes, Mendoza=Yes, and the blind coder independently reproducing Perry's Guard service from a Tier 1 source), Legal **0.77**, Harassment and Clergy 10/10 (α undefined — no variance, all agree), Faith **0.24** (3 of the 4 disagreements).

**The four disagreements, classified — this taxonomy is the real finding:**

1. **Perry / legal (research depth, our favor):** the blind coder coded None found, missing the 2002 PA DEP falsified-reports charge resolved by ARD. The record is documentable to Tier 1; deeper search resolves it. Not a rule failure.
2. **Wittman / faith (research depth, our favor):** blind coder didn't locate the congregation-membership documentation behind our C3 middle-tier entry.
3. **O'Donnell / faith (research depth, UNRESOLVED):** the blind coder recorded faith citing "own statements on Catholic conversion" but its report carried no URL, and a follow-up search here did not verify such statements. Our BLANK stands unchanged; the claim is logged as a lead to check (locate-vs-establish applied to our own validation test). If verified, our entry gets a superseding correction — which would make the blind coder right.
4. **Smith / faith (RULE GAP — the one genuine v1.1 item from V2):** the blind coder read a campaign slogan invoking "Faith" as C4 "campaign's own generic faith language" → RECORDED; we read C4 as requiring personal faith language *about the candidate* ("a man of faith") → BLANK. Both readings are defensible under v1.0 as written. **C4 needs a clarifying sentence.**

**Verdict:** rule-attributable disagreement is 1 in 50 decisions (2%); the standard replicates. The faith cell is the weak zone — precisely the field with the most graduated rules — and gets the v1.1 attention. The blind coder also independently dodged the planted traps: it separated Tony Gonzales (TX-23) from Vicente Gonzalez, put Suozzi's dismissed complaint and Kean's health absence in notes, applied B3 to Smith's common name, and refused the Free-Beacon-type bait.

## V3 — Adversarial stress test: **10 of 12 cases deterministic; two real gaps found**

A second blind session applied the rulebook cold (no tools) to 12 synthetic edge cases built to hurt. Ten produced exactly one verdict with correct rule citations — including the arbitration-win-plus-contempt case, diversion ≠ conviction, admitted-affair-to-notes, the infobox religion, the "proud veteran" with no records (→ Unclear), and the fact-check-bounded attack ad. Two cases exposed genuine gaps:

1. **Common-name POSITIVE hits (Case 1):** B3 flags weak *negatives*, but no rule governs an unverified *positive* record hit on a common name — "exclude entirely" and "record as flagged lead" are both defensible. **v1.1: add B4** — an identity-unverified record hit is never recorded; it becomes a required DOB/address-keyed docket check, logged as a lead.
2. **Partisan filings that are also official records (Case 12):** a party committee's ethics complaint is Tier-4 *content* inside what A1 could call a Tier-1 *official filing* — the tier logic collides with itself. **v1.1: add an A3 sentence** — the *existence* of a filing may be recorded as Tier 1 fact once located in official records; the filing's *claims* remain Tier 4 partisan-origin until independently established.

---

## The scoreboard, and what it means for publication

| Test | Concept/claim | Verdict |
|---|---|---|
| E1 (Aug) | Tenure-money gradient | Survived internal robustness; October prediction registered; F&H cliff-tension on record |
| E2 (n=115) | Bimodal battleground | **Confirmed** (dip p=0.0008; officeholders p<0.001) |
| E3 | Stock-conduct cluster | **Not a battleground excess** — reframe as "Congress is the story" + documentation value |
| E4 | Occupational asymmetry | **Uniformed half refuted (p=0.58); care half suggestive only (p=0.074)** |
| E5 | Roster volatility | Baseline test still pending (prior-cycle hazard data) |
| V1 | Vocabulary prior art | Convergence with five fields; synthesis unclaimed; no tracker publishes an equivalent |
| V2 | Replicability | **α=0.91 — earned**; faith cell weak zone; 1 rule gap |
| V3 | Determinism | 10/12; 2 rule gaps |

**v1.1 changelog queue (v1.0 stays frozen and untouched):** clarify C4 (personal faith language, not slogan/values words); add B4 (identity-unverified positive hits are leads, never entries); add A3 clarification (existence of a partisan filing = Tier 1 fact; its claims stay Tier 4). Plus one open verification task: O'Donnell faith-statement lead from V2.

**Publication logic after today:** Story 1 remains the gradient (pending October). Story 2 — the standard — is now *stronger* than before the tests: "we tested our own rulebook with a blind coder and an adversarial case set; it scored α=0.91; here are the two rules it was missing and the one claim we retracted because our own test failed it" is the most credible version of that story that could exist. The retired uniformed-asymmetry claim should be mentioned there, not buried: retracting it in public is the brand.
