# Framing the News: India 2026 — v1 · Methods & Deviations Log

**Study window:** 18 June – 2 July 2026 · **Coded:** 4 July 2026 · **Design:** constructed week within a single fortnight
**Version note:** This is the first public release. Two internal drafts preceded it; both failed audit (the first on data provenance, the second on exclusion-narrative accuracy and reproducibility gaps identified in an adversarial review) and are superseded in full. Findings of that adversarial review and their dispositions are logged in §8.

## 1. Sampling design

**Day selection (seeded, reproducible — run it yourself).** From the 15-day window, one date per weekday was drawn at random with seed `20260618`. The shipped script `sampling.py` regenerates the draw from scratch and self-checks against the shipped artifacts; it also regenerates the 115-ID validation draw (§4). Both draws use the same literal seed value (the window's start date) as two independent re-seeded invocations; two distinct seeds would have been cleaner practice, and the value is kept only because it is what actually generated the published sample. Result: **Sat Jun 20, Sun Jun 21, Thu Jun 25, Fri Jun 26, Mon Jun 29, Tue Jun 30, Wed Jul 1** — each weekday exactly once. Honest label: a *constructed week within one fortnight*; it controls weekday composition but not seasonality, and single-event domination is expected and handled at event level (§5).

**Snapshot rule.** For each outlet-day: the Wayback Machine capture closest to 12:00 IST (06:30 UTC) that calendar day; else the nearest capture within ±1 day, flagged `ADJACENT`. A capture may serve at most one cell; adjacent fallbacks that resolved to an already-used capture were coded missing. Result: **62 usable cells: 55 same-day, 7 adjacent.**

**Story selection rule.** First **10** editorial story links in the main content area of the archived homepage, in DOM/JSON order, deduplicated by URL and headline. Where the archived page embeds its homepage list as ordered structured data (FirstPost, News18, India Today), the outlet's own declared ordering was used. Pre-specified exclusions: navigation/section links, web-stories/video/podcast program pages, non-English (Devanagari) items, anchor text < 30 characters unless a full headline was recoverable from `aria-label`/JSON.

**Provenance.** Each of the 575 records carries the capture timestamp and a resolvable `https://web.archive.org/web/<ts>/<url>` link. Every headline is auditable against the archive.

## 2. Outlets and the exclusion of The Wire and Times of India

Ten outlets. **Digital native (4):** Scroll.in, The Print, Newslaundry, FirstPost. **Legacy digital (6):** NDTV, India Today, Hindustan Times, Indian Express, News18, The Hindu.

**Inclusion criterion:** an outlet needed usable captures for **at least 4 of the 7 sampled days**. This threshold was formulated after seeing archive coverage (post-hoc, and disclosed as such) but is applied uniformly: all ten retained outlets have 5–7 usable cells; the two exclusions have 1 and 0.

**The Wire (0 usable cells) — corrected justification.** The homepage (`thewire.in`) is a ~9 KB client-rendered shell; archived homepage HTML contains no stories, and no homepage feed endpoint was ever archived. However — correcting this study's earlier draft, which wrongly said the CMS API was "never archived" — the WordPress REST API at `cms.thewire.in/wp-json/thewire/v2/` **was** archived: 604 captures of category-section feeds in/around the window, touching 5 of the 7 sampled days. A reconstruction from these feeds was attempted and rejected on three grounds, documented with the evidence in `appendix_wire_reconstruction.json`: (1) the feeds are section pages ordered by recency (a 15-item reverse-chronological `recent-stories` block plus small curated slots) — a different measurement unit from homepage editorial prominence, which is what this study codes; (2) coverage fails the design — zero captures on sampled Jun 20 and Jun 26, a single category (culture) on Jun 21, and the politics feed captured on only one sampled day; (3) mixing one reconstructed recency-feed outlet into a nine-outlet homepage-prominence sample would reintroduce exactly the undocumented-heterogeneous-selection flaw this study exists to correct. The archived feeds would support a separate, differently-framed section-level analysis of The Wire; they cannot support this one.

**Times of India (1 usable cell) — corrected justification.** The CDX index for `timesofindia.indiatimes.com/` shows 72 captures in the window: **71 are contentless HTTP 301 redirect stubs (~915–970 bytes — no page, no edition)** and **one (2026-06-27 10:33 UTC) is a content capture**. Correcting the earlier draft's claim that captures were "US-edition pages": the redirect stubs are edition-less, and the single content capture is, on inspection, the **Indian edition** — generic TOI title, 86 `/india/` section links, and an Indian top-story list (extracted top-15 published in `appendix_toi_capture.json`). Following the archived 301s through Wayback's nearest-capture resolution does land on off-index US-edition material, which misled the earlier draft. The exclusion therefore rests solely on coverage: one usable cell out of seven (and for Jun 27, adjacent to the sampled Jun 26 — not the sampled Saturday), against the ≥4-cell criterion. TOI is recoverable for future waves only via prospective capture.

## 3. Unit counts

575 stories across 62 cells (10 outlets × 7 days = 70 possible). **Exact reconciliation:** 62 populated + 4 dropped as duplicate captures (Newslaundry 6/21, Scroll 6/30, The Hindu 7/1, The Print 7/1 — listed in `dropped_cells.json`) + 4 with no capture within ±1 day (Scroll 6/20 and 6/21, News18 6/30 and 7/1) = 70. (An earlier draft said "eleven cells" missing; that figure silently included the excluded outlets' cells and could not be reconciled against the dataset. This one can.) Newslaundry cells yield 2–4 stories each (its archived pages server-render only the top items before an infinite-scroll boundary), contributing 15 stories total.

## 4. Coding

Four variables per story, using PEJ *Framing the News* (1998) definitions plus the India-specific **Institutional Critique** frame: Frame (14), Topic (16), Trigger (13), Underlying Message (9). Low-inference default: no clearly dominant device in the headline → Straight News / No Message.

**Coding unit = headline as displayed** (plus URL slug) — a disclosed departure from PEJ's full-text coding. Frame percentages are homepage-presentation framing, not full-text framing.

**Coder & confidence.** Single AI coder. Following audit, a systematic self-audit pass flagged **44 of 575 codes (7.7%) as `low_confidence: true`** under stated criteria: digest/roundup items coded by lead item (e.g. "Rush Hour", "Globe at a glance"); interrogative headlines carrying inferential frames (Conjecture/Reality Check/Consensus); and judgment calls where two or more frames were defensible from the headline alone. Notably, **the study's single Consensus-framed story is itself a low-confidence call.** Sensitivity check: excluding all 44 moves straight news from 45.0% to 47.5% and combative from 11.7% to 12.6% — no headline finding changes. (An earlier draft shipped this field wired to nothing — `false` on all rows; the audit was right to call that a dead mechanism.)

**Human validation.** A seeded 20% subsample (115 stories, regenerable via `sampling.py`) is packaged in `validation_sample.csv` with codes blanked and per-story archive links; `validation_key.csv` holds the AI codes (open only after coding). Cohen's κ per variable will be published on return. **All findings are preliminary until κ is published.**

**Event IDs.** 326 unique events among 575 stories; all headline findings reported at story and event level because constructed-week sampling across 10 outlets makes story-level rows non-independent.

## 5. Known artifacts (disclosed)

India Today's "Punjab civic polls" story appears at position 5 on all 7 days (stale cached widget; kept at story level, one event at event level, flagged). Magazine-style homepages (Scroll, Print, NL) persist stories across days — genuine prominence, inflates story-level counts, corrected at event level. One capture ≈ one moment of a changing homepage. Homepages carry sports/entertainment verticals that 1998 print front pages did not; a hard-news subset (n=455, excluding Sports and Culture/Entertainment) accompanies every headline figure.

## 6. Extended layer: Rosen Transparency Index

The ten outlet scores are now **published** in `rosen_scores.json`, with the five scoring criteria and this status flag: the totals are carried over from a pre-release draft, were assigned by a scorer aware of earlier framing results (circularity risk), and per-criterion breakdowns were never contemporaneously documented — so they are not retroactively invented. The reported association (Spearman ρ, with and without the Newslaundry outlier) is descriptive only. Blind independent re-scoring is a pre-condition for any stronger claim. The Deuze typology and C:E ratio remain dropped (uninformative and under-powered respectively).

## 7. Complete deviations log

1. The Wire excluded — homepage unarchived; section-feed reconstruction attempted and rejected with evidence (§2, appendix). 2. TOI excluded — 71/72 captures are contentless redirect stubs; single Indian-edition content capture fails the ≥4-cell criterion (§2, appendix). 3. Four duplicate-capture cells dropped (`dropped_cells.json`). 4. Four no-capture cells (§3). 5. Seven cells use adjacent-day captures (flagged per row). 6. Newslaundry SSR limit: 2–4 stories/cell. 7. ≥4-cell inclusion criterion formulated post-hoc, applied uniformly. 8. Same seed literal reused across the two independent randomizations (§1).

## 8. Audit trail

An adversarial review of the preceding draft (recomputation of all statistics plus cross-referencing of prose against shipped pipeline files) found: all recomputable statistics exact (frames, topics, cross-tabs, CIs, event counts, validation split); B1 Wire justification contradicted by shipped CDX logs — **fixed, §2, with reconstruction attempt**; B2 TOI "71 of 72 US-edition" unsupported — **fixed, §2, with capture extraction**; B3 "6 of 8" media-concentration slip — **fixed: 8 of 8 (4 Newslaundry, 4 The Print)**; B4 Rosen scores unpublished — **fixed, `rosen_scores.json`**; B5 sampling reproducibility asserted not demonstrated — **fixed, `sampling.py`, verified round-trip**; B6 "eleven cells" unreconcilable — **fixed, §3 sums to 70**; B7 `low_confidence` never firing — **fixed, 44 flags + sensitivity check, §4**.

## 9. Package contents

`framing-india-2026-v1-report.md` · `METHODS_v1.md` (this file) · `dataset_v1.json` (575 records) · `sampling.py` · `rosen_scores.json` · `appendix_wire_reconstruction.json` · `appendix_toi_capture.json` · `dropped_cells.json` · `validation_sample.csv` · `validation_key.csv`
