# Communication-Style Analysis Protocol — v0.1 (FROZEN 2026-08-21, before any measurement)
**Companion to EVIDENTIARY_STANDARD v1.0 · pilot race: AZ-02 (Crane / Nez) · frozen before the first sample is collected, per the E4 lesson: categories defined after seeing data are categories drawn around the finding.**

## What this is and is not

This protocol measures **documented, quantifiable properties of a candidate's public writing** — never personality. The output vocabulary is style features and pattern descriptions with receipts. The words "aggressive," "authentic," "narcissistic," "folksy," or any trait label never appear as findings. If a pattern invites a personality interpretation, the pattern is printed and the interpretation is left to the reader.

## Corpus rules

1. **Official public content only:** the candidate's official House/campaign websites, official press releases, and posts from verified/official public accounts. No private, deleted, or leaked content; no constituent replies (other people's words); no paywall circumvention.
2. **Every sample archived** with source URL, capture date, register tag, and word count in `style_corpus_az02.csv`. A sample that cannot be tied to a URL does not enter the corpus.
3. **Quoted-post rule:** where platforms block fetching, a post quoted VERBATIM in a named-outlet story may enter the corpus with the story as source — tagged `social-quoted`, and the selection bias this introduces (outlets quote the spiciest posts) is stated in every use of that register. This is the corpus's biggest bound.
4. **Two registers per candidate:** `official` (press releases, site bios, official statements) and `social` (posts, incl. social-quoted). Minimum 8 samples and 600 total words per register to compute anything; below that, the cell reports "insufficient corpus," never a number.
5. **Time window:** 2023-01-01 through capture date, stated per sample.

## Frozen feature set (computed per sample; aggregated per register)

Lexical/structural: mean sentence length + variance; type-token ratio (vocabulary richness, on 100-word windows); readability (Flesch-Kincaid grade). Voice: first-person-singular density (I/me/my per 100 words); first-person-plural (we/our); third-person-self (own name/title per 100 words). Intensity: exclamation rate; ALL-CAPS word rate (excluding acronyms); question rate. Content orientation: attack-lexicon rate (opponent/party names + a fixed list: corrupt, radical, extreme, lie(s/ing), disgrace, failed, weak, crooked) vs policy-lexicon rate (fixed list: bill, act, vote(d), fund(ing), veterans, border, water, healthcare, jobs, tribal, district). Platform furniture: hashtag and emoji rates (social register only).

**The consistency measures (the "same writer?" question):** within-register dispersion of the voice + intensity features across samples (coefficient of variation); register gap = absolute difference of each feature between official and social registers. High within-register dispersion or a large register gap is reported as a **measured style break**. A style break is a pattern, NOT a ghostwriting finding — authorship explanations are recorded only if the campaign or candidate has stated them (locate-vs-establish).

## Never-do list

No trait or diagnosis labels. No authorship attribution stated as fact. No AI-detection claims (unreliable; not in scope). No inference from writing style to policy positions or fitness. No comparison to non-candidates. Results below corpus minimums are "insufficient corpus." All limitations printed with the results, prominently: the social-quoted selection bias, the platform-fetch bound, and pilot-scale n.

## Pilot success criteria (decided now)

The pilot succeeds if it produces (a) a per-candidate feature table with every number traceable to archived samples, (b) at least one register-gap or consistency measurement per candidate, and (c) a scaling decision: whether the method is worth running on all 115 post-October, and what corpus source would be required to do it fairly. The pilot does NOT need to find anything interesting to succeed — "these two candidates write the way their press shops write" is a valid result.

*v0.1 frozen 2026-08-21. Changes → v0.2 with changelog, never retroactive to collected samples.*

---

## v0.2 changelog (2026-08-22, append-only; v0.1 text above unchanged)

1. **Social-register word minimum recalibrated: 600 → 300 words** (sample minimum stays 8). The v0.1 minimum was set without accounting for platform length caps — eight posts of ~40 words cannot reach 600. **Erratum disclosed:** the AZ-02 pilot report published social-cell numbers (Crane 409w, Nez 468w after dedup) that were below the v0.1 minimum as frozen — a protocol breach caught in the four-race review. Under v0.2 those cells are compliant; the breach is recorded here rather than erased.
2. **Corpus descriptors exempted from the insufficiency gag:** sample counts, pooled word counts, and median post length are corpus facts and are always reportable; style features remain barred for insufficient cells.
3. Named-opponent/epithet detection added to the attack-lexicon roadmap (from pilot); not yet implemented — lexicon results remain labeled crude.

---

## v0.3 changelog (2026-08-22, append-only; v0.1/v0.2 text above unchanged)

1. **New subtype `spoken`** for quotations sourced from interviews, debates, town halls, gaggles, and rally remarks. Spoken samples are collected and archived but EXCLUDED from the written-register comparisons (register gap, quote-to-post distance), which now use written attributed quotes only.
2. **Stated rationale for #1 was WRONG and is corrected.** The rule was introduced on the assumption that spoken quotes read simpler than written statements and would compress measured gaps. Both candidates who could test it showed the opposite or nothing: Dunlap written FK 11.1 vs spoken 14.3 (spoken 3.2 grades *harder*); Taylor 12.7 vs 13.1. The separation rule is retained as sound practice — different production processes should not be pooled — but the simplicity claim is withdrawn.
3. **Named-staffer content excluded from the candidate's registers.** A campaign-account post carrying a statement authored by a named consultant is that staffer's writing, not the candidate's, and does not enter the corpus (applied to one LePage post).

---

## v0.4 changelog (2026-08-22, append-only; earlier text unchanged)

1. **Labeled out-of-frame benchmarks are now permitted, under conditions.** v0.1's never-do list barred "comparison to non-candidates." That bar is relaxed ONLY as follows: a non-candidate public figure may be measured on the frozen feature set as a **labeled reference benchmark** when (a) their corpus is collected under the same rules, (b) the samples are stored in a SEPARATE file (`style_corpus_benchmark.csv`) and **never pooled into candidate statistics**, (c) every published number is labeled "benchmark, out of frame," and (d) the genre mismatch is stated. Authorized by the operator 2026-08-22; first applied to Donald Trump.
2. **The rest of the never-do list applies to benchmark subjects without exception** — no trait labels, no authorship attribution, no AI-detection claims, no inference from style to fitness or character.
3. **New reported metric: release-body → social FK gap**, computed identically for all subjects. Added because some subjects (Trump, LePage) have no written-attributed-quote cell, making the primary register-gap metric uncomputable for them. NOTE: this metric shows **no age correlation** among candidates (r = +0.05) where the written-quote metric showed the under-55/55+ split — the age finding is operationalization-sensitive and must be reported with that caveat.

---

## v0.5 queue (opened 2026-08-22 from the Hegseth benchmark; not yet applied)

1. **The attack lexicon is invalid for national-security subjects and must be reworked.** It cannot distinguish opponent-directed attack rhetoric from portfolio vocabulary — "lethal," "terrorists," "kill" are a Defense official's job description and a candidate's escalation. Restrict the lexicon to opponent-directed and characterological terms; freeze before use; declare it inapplicable to national-security portfolios.
2. **Logged error:** in the Hegseth run, war-domain words were added to the attack lexicon mid-analysis — an after-the-fact feature adjustment of exactly the kind the frozen-feature rule exists to prevent. The resulting attack figure was withheld as invalid rather than published. Recorded, not erased.
3. **Executive-branch subjects have no attributed-quote register** (confirmed for both Trump and Hegseth): congressional releases quote the member in first person, executive releases speak institutionally. The primary written-quote→social register-gap metric is therefore uncomputable for executive officials; only the release→social alternate applies. Treat as a structural property, not missing data.
4. **Account pooling must be declared.** Where a subject posts from multiple accounts (institutional + personal), state that they were pooled; ideally separate them in future runs.

---

## v0.5 changelog (2026-08-22, append-only; earlier text unchanged) — applied from the media-register round

1. **Queue item 1 APPLIED.** The attack lexicon is restored byte-identical to the frozen v0.1 set in `scripts/bench_analysis.py`; the war-domain words added mid-analysis during the Hegseth run are **reverted**. The lexicon is now **declared inapplicable to national-security portfolios**, and attack figures for such subjects are withheld rather than published (applied to Hegseth and Gabbard). The narrower opponent-directed/characterological rework remains open for v0.6.
2. **Queue item 3 is REFUTED and withdrawn.** "Executive-branch subjects have no attributed-quote register" was generalized from n=2 (Trump, Hegseth). Three of four further executive subjects **do** have one: Duffy (transportation.gov, k=7), McMahon (ed.gov, k=11), Gabbard (odni.gov, k=5). The absence is a **department house-style choice, not a structural property of the branch**, and the primary written-quote→social metric IS computable for most executive officials. Missing quote cells are recorded as absent-for-this-office, not as a branch rule.
3. **Queue item 4 APPLIED.** Account pooling is declared in-report wherever it occurs (Hegseth ×3 accounts, Gabbard ×2).
4. **ALL-CAPS feature definition PINNED.** The feature was never present in `style_analysis.py`; the figures in the Trump and Hegseth reports came from an unpreserved ad-hoc computation. Definition as of now, in `scripts/bench_analysis.py`: *purely alphabetic tokens of length ≥ 2 that are entirely uppercase, over all word tokens* — hashtags, handles and alphanumerics excluded as formatting rather than emphasis. Figures published before this date are **superseded**, not corrected in place (Trump 11.62 → 14.29 under the pinned rule).
5. **New rule — withheld cells may never serve as comparison bounds.** The Trump and Hegseth reports quoted "candidate max 3.12 (LePage)" and "1.55 (Buck)" as field maxima while both cells were withheld as insufficient in the same project. A cell too thin to report is too thin to bound a comparison. Field maxima are now drawn only from cells meeting the frozen minimums.
6. **Benchmark subjects may serve as pre-registered controls.** The Kelly control condition ("if he lands at candidate median, the broadcaster claim holds; if he lands at 8.6, it is public-communication experience generally") was stated to the operator before collection and decided the result. Pre-registering the control's interpretation *before* collection is adopted as required practice for any future background-effect test.

### v0.6 queue (opened 2026-08-22)

1. Narrow opponent-directed/characterological attack lexicon; freeze before use (carried from v0.5 item 1).
2. Named-opponent/epithet detection — carried unimplemented since v0.2.
3. Separate, rather than pool, multi-account subjects.
4. Solve or formally bound the Facebook-invisibility problem before the 59-race sweep (carried from the 10-race round).
5. Test the executive-vs-candidate register-gap difference (Finding 5, n=4) on a larger set of national-office holders.

---

## v0.6 changelog (2026-08-22, append-only; earlier text unchanged) — the platform bound, and a correction to our own bias claim

### 1. THE FACEBOOK BOUND IS PERMANENT AND IS NOW DECLARED

A methods test run 2026-08-22 establishes that organic Facebook Page text **cannot be collected by any legitimate route available to this project**, and the reason is not effort:

- Every unauthenticated route terminates at a login wall — `facebook.com/<page>`, `mbasic.facebook.com`, `m.facebook.com`, and direct post permalinks (HTTP 400). Facebook's `robots.txt` prohibits automated collection without written permission.
- **CrowdTangle was shut down on 2024-08-14.** The standard method the field used for exactly this measurement no longer exists; studies built on 2022 data are not replicable on 2026 data. This is a citable, externally verifiable reason for the bound.
- The one legitimate successor, **Meta Content Library**, requires affiliation with an accredited degree-granting nonprofit institution (journalists and independent researchers are ineligible), takes 4–6 weeks, and confines analysis to a secure enclave whose raw-text export policy would likely make a two-register comparison impossible anyway.
- The **Meta Ad Library** covers paid ads only. Using it would be actively harmful here: ad creative is committee-written and legally reviewed, so measuring it as "the candidate's writing style" would introduce a worse error than the missing data.

**The decisive point is methodological, not logistical.** Every partial channel that does leak text — search-result slugs, indexed snippets — **truncates at roughly 200 characters and strips casing and punctuation**. A slug renders `RESULTS MATTER. ACTION MATTERS.` as `results-matter-action-matters`. Those are precisely the features this protocol measures. A workaround here would corrupt the style measure while appearing to succeed, which is the worst available failure mode.

**Rules adopted:**

**F1.** Facebook-primary candidates are recorded as **structurally uncollectable** in the social register — not as missing data, and never as a thin cell padded to the minimum.
**F2.** No user-agent spoofing, headless-browser login, or third-party scraping service is used, on ethics, terms, and data-quality grounds together.
**F3.** Register substitution is barred. Press-conference quotes, ad copy, and slug fragments may never be used to reach the social minimum — each is a different register and would manufacture a spurious style signal.
**F4.** A candidate with an uncollectable social register may still be carried **official-register-only**, with the social cell marked structurally missing rather than withheld-as-insufficient. The two are different findings and are labeled differently.

**LePage (ME-02) is the standing case and stays unmeasured.** He is the project's single most load-bearing observation — 78 years old with zero congressional tenure, the point that separates age from tenure — and he cannot be bought with contaminated data. One manual check remains worth doing: `x.com/PaulRLePage1` exists and its volume could not be verified from here; the existing 3-post yield suggests it is dormant.

### 2. CORRECTION — our stated bias claim was wrong twice, and the real asymmetry is on the platform we DO collect

The 10-race report called the Facebook gap "a sampling bias with a partisan tilt risk" and "the single biggest threat to a fair 59-race sweep," on the reasoning that Facebook skews older and more rural. Against Pew's 2025 social-media data (American Trends Panel Wave 164) and Lukito et al. 2025, **two parts of that are not supported**:

- **"Facebook skews older" is wrong as stated.** Daily Facebook use peaks at ages 30–49 (58%) and is *lowest* among 65+ (45%). What is genuinely age-graded is **relative dependence**: for a 65+ audience Facebook is 11.3× more available than X (45% vs 4%), against 2.7× for 18–29. That is the real mechanism — an older candidate concentrates on the only channel his electorate uses — and it is what the claim should have said.
- **The partisan tilt claim is thin and should be dropped.** Candidate-level Facebook adoption in 2022 ran 87% D vs 84% R — a 3-point gap. Pew's daily-use numbers show Republicans higher on Facebook (56% vs 49%) *and* higher on X (12% vs 9%), which cuts against a clean "Republicans are Facebook-first" story.
- **The rural gap survives** and is claimable: Facebook daily use rural 57% vs urban 49%; X rural 8% vs urban 10%.

**The asymmetry we should have been worried about is Bluesky — the platform this project collects most successfully.** Bluesky carries a documented heavy left tilt (Pew on news influencers; arXiv:2506.03443 finds minority political stances at ≤1% of users on major topics). Our richest single social corpus in the whole project — Dunlap at 29 posts, via Bluesky's open API — comes from that platform. **A design that collects Bluesky fully and Facebook not at all is structurally more likely to over-sample Democratic written output than the Facebook gap alone implies.** This is now the primary stated limitation of the sweep, ahead of the Facebook bound, and it must appear in any published style result.

### 3. Carried forward, still unimplemented

1. Narrow opponent-directed / characterological attack lexicon; freeze before use (from v0.5).
2. Named-opponent / epithet detection — carried unimplemented since v0.2.
3. Separate, rather than pool, multi-account subjects.
4. Test the executive-vs-candidate register-gap difference (media-register benchmark, Finding 5, n=4) on a larger set of national-office holders.

*v0.6 issued 2026-08-22. Changes → v0.7 with a changelog.*
