Framing the News: India 2026

How Indian digital journalism frames the news — a replication of PEJ's landmark 1999 US study

v1 · First public release · Constructed week, 18 June – 2 July 2026 · 575 homepage stories · 10 outlets · 62 outlet-days · every story archive-verifiable · A flagship report by OnlineJournalism.in

Status: PRELIMINARY. Single AI coder; findings hold pending human validation of a seeded 20% subsample (115 stories, shipped as validation_sample.csv). Cohen's κ will be published on its return. All percentages carry 95% Wilson confidence intervals, shown in [brackets]. This release supersedes all prior drafts, including the 170-story instrument previously published at this address (preserved under /superseded/).

Why this version is defensible

Every story was extracted from a dated Wayback Machine capture and carries a resolvable archive URL. Days were drawn with a documented seed, and the shipped sampling.py regenerates both the day draw and the validation draw from scratch. The story-selection rule is fixed and stated; duplicate captures were rejected; the deviations log reconciles exactly (62 + 4 + 4 = 70 cells). Two outlets could not be included, and — following an adversarial audit of the previous draft — their exclusions are documented with evidence rather than asserted: The Wire's archived CMS section feeds were examined and a reconstruction attempted before rejection (appendix); Times of India's single usable capture was parsed and its Indian-edition top-15 published (appendix). The dataset self-reports its uncertainty: 44 of 575 codes (7.7%) are flagged low-confidence, with a sensitivity check showing no headline finding depends on them. Full methods and deviations log: METHODS_v1.md.

575
Homepage stories
10
Outlets
62
Outlet-days
326
Unique events
7.7%
Codes flagged low-confidence
100%
Archive-verifiable

Headline Findings

Frame Distribution

Narrative Frames — All Outlets (N=575)

Story level vs Event level (%): does non-independence change the picture?

India 2026 vs. PEJ USA 1999

Frame Distribution: India 2026 (N=575) vs United States 1999 (PEJ) Directional only

Comparison discipline

The PEJ benchmark confounds country × era × medium (1998 US print front pages vs 2026 Indian digital homepages, full-text vs headline coding) and is used directionally only: more straight news than late-90s US front pages, less combative framing, consensus near-absent in both. Claims about "Indian digital journalism" in general are not supported by one fortnight, ten outlets, and headline-level coding — this report deliberately does not make them.

Topics, Triggers & Messages

Topic Distribution

Story Triggers: What Makes News

Underlying Messages 71% No Message

Trigger × Frame (top 6 triggers, % within trigger)

Outlet Fingerprints

Straight-News Share by Outlet (%)

Accountability Frames by Outlet (%)

The Ten Outlets

OutletTypeStories (n)Straight News %Accountability %Rosen Score Unverified
NewslaundryDigital native150.0%66.7%10/10
Indian ExpressLegacy digital7024.3%21.4%7/10
The PrintDigital native6026.7%15.0%6/10
Scroll.inDigital native4027.5%10.0%7/10
NDTVLegacy digital7050.0%7.1%5/10
News18Legacy digital5060.0%6.0%3/10
Hindustan TimesLegacy digital7040.0%4.3%5/10
India TodayLegacy digital7058.6%4.3%4/10
The HinduLegacy digital6058.3%3.3%6/10
FirstPostDigital native7065.7%2.9%4/10

Newslaundry contributes only 15 stories (its archived pages server-render just the top items) and is a media-watchdog by design — its 66.7% carries a wide interval [41.7–84.8]. Rosen scores are published carryovers awaiting blind re-scoring: see rosen_scores.json for the criteria and the caveats.

Transparency & Accountability

Rosen Transparency Index vs Accountability Framing Descriptive only

What the correlation can and cannot say

Spearman ρ = 0.71 across ten outlets (Pearson 0.85). Remove the Newslaundry outlier and ρ = 0.60. That is a suggestive descriptive gradient on ten data points — not a law.

The Rosen scores were assigned before this release by a scorer aware of earlier framing results (circularity risk), per-criterion breakdowns were never contemporaneously documented, and they have not been independently verified. They are published anyway — auditable inputs beat asserted ones. Blind re-scoring by an independent rater is a pre-condition for any stronger claim.

The earlier draft's r = 0.92 and its "corporate outlets produce zero accountability journalism" claim did not survive the audit and are withdrawn.

The Event Level: What Actually Happened This Fortnight

Largest Event Clusters (stories per event)

EventStoriesShare of sampleNote
Pune murder case (Lohagad Fort)274.7%One crime story outweighs education (3.5%), health (1.7%) and environment (3.7%)
US–Iran deal (Doha talks)234.0%The window's dominant diplomacy story — and source of the sample's single Consensus frame
Ram temple donation scandal203.5%
Venezuela earthquake142.4%
Strait of Hormuz tensions122.1%
NEET re-exam111.9%
India Today stale poll widget71.2%Cached "Punjab civic polls" widget at position 5 on all 7 days — flagged, one event

The Excluded Outlets: Evidence, Not Assertion

The Wire — 0 usable cells

The homepage is a ~9 KB client-rendered shell; archived homepage HTML contains no stories. The WordPress CMS section feeds were archived (604 captures in/around the window) and a reconstruction was attempted — then rejected: the feeds are recency-ordered section pages, a different measurement unit from homepage editorial prominence; coverage fails the sampling design (zero captures on two sampled days); and mixing one recency-feed outlet into a nine-outlet prominence sample would reintroduce the undocumented-heterogeneous-selection flaw this study exists to correct.

Evidence: appendix_wire_reconstruction.json

Times of India — 1 usable cell

72 captures in the window: 71 are contentless HTTP 301 redirect stubs (~915–970 bytes — no page, no edition) and one (27 June) is a genuine Indian-edition content capture, whose top-15 stories are published in the appendix. The exclusion rests solely on coverage: one usable cell of seven against the ≥4-cell criterion. The earlier draft's "US-edition pages" claim was wrong and is corrected in METHODS §2.

Evidence: appendix_toi_capture.json

Limitations & Next Steps

Limitations

Headline-only coding (frames inside story bodies invisible). Single AI coder; 7.7% of codes self-flagged low-confidence; findings preliminary until human κ is published (validation package shipped). One capture per day; 7 of 62 cells adjacent-day. Newslaundry contributes 15 stories. No TOI or Wire — exclusions evidenced in appendices; the native camp (n=185) is half the legacy camp (n=390). Two-week window; no seasonal or election-period claims. English-only. Rosen scores published but unverified. India Today's stale poll widget appears 7× (flagged; one event).

Next steps, in order

1. Human-code the 115-story validation sample; publish κ.
2. Blind re-scoring of the Rosen Index by an independent rater.
3. A second constructed week in a different news season, and one in an election window.
4. Prospective daily archiving of all twelve target homepages (including TOI and The Wire) so no future wave depends on crawl luck.

Data & Reproducibility Package

Everything needed to audit, recompute, or extend this study. Every record in the dataset carries a resolvable Wayback Machine URL.

framing-india-2026-v1-report.md

The v1 report — findings with confidence intervals

METHODS_v1.md

Full methods, deviations log, and audit trail

dataset_v1.json

575 coded records with archive URLs, event IDs, low-confidence flags

sampling.py

Regenerates the day draw and the validation draw from seed 20260618

rosen_scores.json

Rosen Index inputs, criteria, and caveats

appendix_wire_reconstruction.json

The Wire exclusion evidence — attempted feed reconstruction

appendix_toi_capture.json

TOI exclusion evidence — the single content capture, parsed

dropped_cells.json

Duplicate-capture cells rejected during sampling

validation_sample.csv

115-story blind human-validation sample (codes blanked)

validation_key.csv

AI codes for the validation sample — open only after coding

Glossary & Reading This Study Honestly

Complete Glossary of Terms

Hover over dotted-underlined terms throughout the dashboard for quick definitions.

Frame
The dominant narrative device a journalist uses to organise a story. Not what the story is about (that's topic), but how it is told. In this release the coding unit is the headline as displayed on the homepage (plus URL slug) — a disclosed departure from PEJ's full-text coding. Frame percentages here are homepage-presentation framing, not full-text framing.
Straight News
The classic inverted pyramid. No dominant narrative device other than presenting who, what, when, where, why, and how. The low-inference default: where no frame is clearly dominant in the headline, the story codes as Straight News.
Conflict
The story is built around the clash or disagreement inherent in a situation, foregrounding opposing sides, tensions, or disputes.
Consensus
The story is built around points of agreement among stakeholders. 0.2% of this sample [0.0–1.0] — and the single instance is itself a low-confidence call. PEJ's already-low US figure was 6%.
Conjecture
Speculation about what will happen next. Forward-looking, predictive, or scenario-based framing.
Process
An explanatory frame: how something works, the mechanics of a system, the steps involved.
Historical Outlook
The story places current events in historical context, using the past to illuminate the present.
Horse Race
Who is winning and who is losing. 1.2% in this window — which contained no major Indian election.
Trend
The news is presented as part of an ongoing, larger pattern.
Policy Explored
The journalist dives into the substance of a policy — mechanisms, impact, trade-offs, implementation.
Reaction
The story is structured around responses from key players to an earlier event.
Reality Check
The journalist examines the veracity of a claim, statement, or widely held belief.
Wrongdoing Exposed
The story uncovers injustice, corruption, malfeasance, or misconduct — revelation with identifiable actors.
Personality Profile
The story is built around an individual — their character, background, motivations.
Institutional Critique
India-specific frame added to the PEJ codebook. Diagnoses systemic failure without necessarily naming individual wrongdoers. "Why does NEET keep failing?" is institutional critique; "CBI arrests teacher in NEET leak" is wrongdoing exposed.
Trigger
What caused the news organisation to cover this story today — a statement, a spontaneous event, a report, or the newsroom's own initiative. Reveals whether journalism is reactive or proactive.
Underlying Message
An often unconscious cultural or moral narrative embedded in the story ("perseverance pays off", "the system is broken"). 71% of headlines carry no detectable message — headline-level coding reads fewer messages than full-text coding would.
Constructed Week
A sampling technique from media research: one Monday, one Tuesday, etc., drawn at random so no single news cycle dominates. This release is honestly labelled a constructed week within one fortnight (18 June – 2 July 2026, seed 20260618): it controls weekday composition but not seasonality, and single-event domination is handled at event level instead.
Event Level
Stories sharing an event ID cover the same underlying occurrence. 575 stories collapse to 326 events. Constructed-week sampling across 10 outlets makes story-level rows non-independent, so every headline finding is reported at both levels.
Low-Confidence Flag
44 of 575 codes (7.7%) self-flagged under stated criteria: digest/roundup items coded by lead item, interrogative headlines carrying inferential frames, and calls where two or more frames were defensible from the headline alone. Sensitivity check: excluding all 44 moves straight news from 45.0% to 47.5% and combative from 11.7% to 12.6% — no headline finding changes.
Digital Native
Born on the internet, no prior print or broadcast edition. In this release: Scroll.in, The Print, Newslaundry, FirstPost (n=185).
Legacy Digital
Traditional print or broadcast organisations that expanded to digital. In this release: NDTV, India Today, Hindustan Times, Indian Express, News18, The Hindu (n=390).
Combative Frames
PEJ grouping: Conflict + Horse Race + Wrongdoing Exposed. PEJ found 30% on US front pages in 1999; this study finds 11.7% [9.3–14.5].
Accountability Frames
Wrongdoing Exposed + Reality Check + Institutional Critique — frames that perform a watchdog or verification function.
Rosen Transparency Index
A 0–10 outlet score after Jay Rosen's scholarship on journalistic authority, on five criteria (each 0–2): funding/ownership disclosure, corrections policy, stated editorial stance, proprietor–editorial separation, reader accountability mechanisms. Published in rosen_scores.json with caveats; awaiting blind re-scoring.
Wilson Interval
The 95% confidence interval used throughout — well-behaved for small samples and extreme proportions.
Cohen's κ
Chance-corrected agreement between two coders. The pending human validation of the 115-story subsample will yield κ per variable; all findings are preliminary until it is published.

Methodology in brief

Sampling: From the 15-day window 18 June – 2 July 2026, one date per weekday drawn at random with seed 20260618 (Sat Jun 20, Sun Jun 21, Thu Jun 25, Fri Jun 26, Mon Jun 29, Tue Jun 30, Wed Jul 1). For each outlet-day, the Wayback Machine capture closest to 12:00 IST, else nearest within ±1 day (flagged). First 10 editorial story links in the main content area, in DOM/JSON order, deduplicated, with pre-specified exclusions. 62 usable cells of 70 possible; the reconciliation is exact (62 populated + 4 duplicate-capture + 4 no-capture = 70).

Coding: Four variables per story — Frame (14), Topic (16), Trigger (13), Underlying Message (9) — using PEJ Framing the News (1998) definitions plus Institutional Critique. Coding unit = headline as displayed. Single AI coder with a systematic self-audit pass; 20% seeded human-validation subsample shipped.

Full detail, deviations log (8 entries), and the audit trail of the adversarial review this release passed through: METHODS_v1.md.

Reading this study honestly

"Headlines aren't stories — you're coding packaging, not journalism."

Correct, and disclosed as the study's first limitation. The coding unit is the headline as displayed on the homepage plus the URL slug. What this measures is homepage-presentation framing — the storytelling promise an outlet makes at its front door — not full-text framing. Headline-level coding also skews toward inverted-pyramid readings, which is stated alongside the straight-news finding it inflates.

"A single AI coder can't be trusted."

Which is why the release ships a seeded, blind 115-story validation sample for human coding, flags its own 44 least-confident calls, publishes a sensitivity check showing no headline finding depends on them, and labels every finding preliminary until κ is published. Distrust is the correct default; the package is built so the distrust can be resolved empirically.

"Two weeks in one news season proves nothing about Indian journalism."

Agreed — and the report deliberately makes no such claim. The window was saturated with fast-moving spot news; horse race at 1.2% in a non-election window shows how season-conditional these figures are. The claim is about this fortnight's homepages, stated with intervals. Wider claims wait for a second constructed week in a different season and one in an election window.

"You dropped outlets until the story fit."

The two exclusions are evidenced, not asserted: The Wire's archived CMS feeds were examined and a reconstruction attempted before rejection on stated grounds; TOI's 72 captures were classified and the single usable one parsed and published. The ≥4-cell inclusion criterion was formulated after seeing archive coverage — disclosed as post-hoc — but applied uniformly: all ten retained outlets have 5–7 usable cells; the exclusions have 1 and 0.

"The earlier version of this page said something different."

It did, and it was wrong in ways an adversarial audit made specific: an unreproducible sample, unpublished Rosen inputs, an exclusion narrative contradicted by the archive's own logs, a dead low-confidence field, and headline claims (r = 0.92; "corporate outlets produce zero accountability journalism") that did not survive recomputation. That version is preserved at /superseded/; the audit findings and their dispositions are logged in METHODS_v1.md §8. This is what correction is supposed to look like.

Academic References

Project for Excellence in Journalism & Princeton Survey Research Associates. (1999). Framing the News: The Triggers, Frames, and Messages in Newspaper Coverage. Pew Research Center.
Rosen, J. (2003–present). PressThink: Ghost of Democracy in the Media Machine. pressthink.org. Key concepts: "The View from Nowhere" (2010), "The People Formerly Known as the Audience" (2006).
Entman, R.M. (1993). "Framing: Toward Clarification of a Fractured Paradigm." Journal of Communication, 43(4), 51-58.
Scheufele, D.A. (1999). "Framing as a Theory of Media Effects." Journal of Communication, 49(1), 103-122.
Wilson, E.B. (1927). "Probable Inference, the Law of Succession, and Statistical Inference." JASA, 22(158), 209-212.