# Framing the News: India 2026 — v1 (Preliminary)

**Constructed week, 18 June – 2 July 2026 · 575 homepage stories · 10 outlets · 62 outlet-days · every story archive-verifiable**

*First public release. Supersedes all prior drafts. Status: PRELIMINARY pending human validation of a 20% subsample (κ to be published on return of `validation_sample.csv`). All percentages carry 95% Wilson confidence intervals. Full methods, deviations log, and audit trail: `METHODS_v1.md`. Data: `dataset_v1.json`. Reproduce the sampling yourself: `sampling.py`.*

---

## Why this version is defensible

Every story was extracted from a dated Wayback Machine capture and carries a resolvable archive URL. Days were drawn with a documented seed, and the shipped `sampling.py` regenerates both the day draw and the validation draw from scratch. The story-selection rule is fixed and stated; duplicate captures were rejected; the deviations log reconciles exactly (62 + 4 + 4 = 70 cells). Two outlets could not be included, and — following an adversarial audit of the previous draft — their exclusions are now documented with evidence rather than asserted: The Wire's archived CMS section feeds were examined and a reconstruction attempted before rejection (`appendix_wire_reconstruction.json`); Times of India's single usable capture was parsed and its Indian-edition top-15 published (`appendix_toi_capture.json`). The dataset self-reports its uncertainty: 44 of 575 codes (7.7%) are flagged low-confidence, with a sensitivity check showing no headline finding depends on them.

## Headline findings

**1. Straight news dominates Indian homepages.**
45.0% [41.0–49.1] of stories; 42.6% [38.2–47.2] in the hard-news subset (n=455, excluding sports and entertainment); 47.5% among high-confidence codes only. Against PEJ's 16% for 1998 US front pages, Indian digital homepages are roughly three times as fact-report-framed. Caveats: headline-level coding skews toward inverted-pyramid readings, and the window was saturated with fast-moving spot news (US–Iran deal, Hormuz, Venezuela earthquake, World Cup).

**2. Combative framing is low — half of what the earlier draft claimed, a third of PEJ's US figure.**
Conflict + horse race + wrongdoing = 11.7% [9.3–14.5] (14.1% hard-news; 12.6% high-confidence). Horse race is 1.2% in a window with no major Indian election — evidence that combativeness findings are heavily election-conditional.

**3. Consensus framing is vanishingly rare — with an honesty asterisk.**
One story in 575 (0.2% [0.0–1.0]): India Today's "US, Iran agree to halt strikes, hold talks in Doha." A fortnight dominated by a peace process — a story literally *about* adversaries agreeing — produced almost no consensus-framed coverage. Asterisk: that single story is itself one of the 44 low-confidence calls, so the defensible claim is "zero to one story in 575." Either way, the frame is functionally absent from Indian homepage journalism in this window.

**4. The digital-native/legacy divide mostly does not replicate.**
Straight news: native 39.5% [32.7–46.6] vs legacy 47.7% [42.8–52.6] — overlapping intervals. FirstPost (digital-native, corporate-owned, 65.7% straight news) breaks the "natives interpret, legacy reports" story on its own.

**5. "Corporate outlets produce zero accountability journalism" is falsified.**
Accountability frames (wrongdoing + reality check + institutional critique): NDTV 7.1% [3.1–15.7], India Today 4.3%, News18 6.0%, HT 4.3% — small, not zero. The gradient survives in compressed form: Newslaundry 66.7% [41.7–84.8] (n=15, watchdog by design), Indian Express 21.4%, The Print 15.0%, Scroll 10.0%, FirstPost 2.9%. Native 13.5% vs legacy 7.9%, intervals overlapping.

**6. Transparency–accountability correlation: modest, outlier-sensitive, now fully auditable.**
Spearman ρ = 0.71 across 10 outlets (Pearson 0.85); ρ = 0.60 without Newslaundry. The Rosen Index inputs are published (`rosen_scores.json`) with their limitations stated: unverified carryover scores, circularity risk, no contemporaneous per-criterion documentation. Read as a suggestive descriptive gradient on ten data points awaiting blind re-scoring — not a law, and nothing like the r = 0.92 once claimed.

**7. Enterprise journalism is one-third accountability journalism.**
31.0% [22.3–41.4] of 87 enterprise-triggered stories carry accountability frames; the modal enterprise frames are historical outlook (17) and trend (16). Indian enterprise reporting is *substantially* accountability work — not *predominantly*, as the earlier draft said.

**8. Government triggers produce compliant framing.**
68.8% [59.7–76.6] of government-triggered stories are straight-news framed; 13.4% combative. PEJ found US journalists most combative precisely on government triggers. Indian homepage journalism largely relays official action rather than contesting it — the sharpest accountability finding here, and it implicates native and legacy outlets alike.

**9. What fills Indian homepages.**
Foreign affairs 24.0%, sports 13.7%, politics 12.7%, crime 8.0%. One murder case (Pune's Lohagad Fort case) generated 27 stories — 4.7% of the national English homepage sample — more than education (3.5%), health (1.7%), or environment (3.7%). Media/press freedom is 1.4% of the sample, and **all 8 of those stories come from just two outlets — Newslaundry (4) and The Print (4)**. No legacy outlet put a single press-freedom story in its homepage top-10 across seven sampled days.

**10. Event-level correction.**
575 stories collapse to 326 events (biggest clusters: Pune murder 27, US–Iran deal 23, Ram temple donation scandal 20, Venezuela earthquake 14, Hormuz 12, NEET re-exam 11). Modal event-level frame is still straight news (144/326). Outlet fingerprints built on story-level counts partly measure which events fell in the window; both levels ship in the dataset.

## Comparison discipline

The PEJ benchmark confounds country × era × medium and is used directionally only: more straight news than late-90s US front pages, less combative framing, consensus near-absent in both. Claims about "Indian digital journalism" in general are not supported by one fortnight, ten outlets, and headline-level coding — the report deliberately does not make them.

## Limitations

Headline-only coding (frames inside story bodies invisible). Single AI coder; 7.7% of codes self-flagged low-confidence; findings preliminary until human κ is published (validation package shipped). One capture per day; 7 of 62 cells adjacent-day. Newslaundry contributes 15 stories (archive renders only its top items). No TOI or Wire — exclusions evidenced in appendices; the native camp (n=185) is half the legacy camp (n=390). Two-week window; no seasonal or election-period claims. English-only. Rosen scores published but unverified. India Today's stale poll widget appears 7× (flagged; one event).

## Next steps, in order

Human-code the 115-story validation sample; publish κ. Blind re-scoring of the Rosen Index by an independent rater. A second constructed week in a different news season, and one in an election window. Prospective daily archiving of all twelve target homepages (including TOI and The Wire) so no future wave depends on crawl luck.

---

*Package: `dataset_v1.json` · `METHODS_v1.md` · `sampling.py` · `rosen_scores.json` · `appendix_wire_reconstruction.json` · `appendix_toi_capture.json` · `dropped_cells.json` · `validation_sample.csv` · `validation_key.csv`*
