How Indian digital journalism frames the news — a replication of PEJ's landmark 1999 US study
v1 · First public release · Constructed week, 18 June – 2 July 2026 · 575 homepage stories · 10 outlets · 62 outlet-days · every story archive-verifiable · A flagship report by OnlineJournalism.in
Status: PRELIMINARY. Single AI coder; findings hold pending human validation of a seeded 20% subsample (115 stories, shipped as validation_sample.csv). Cohen's κ will be published on its return. All percentages carry 95% Wilson confidence intervals, shown in [brackets]. This release supersedes all prior drafts, including the 170-story instrument previously published at this address (preserved under /superseded/).
Why this version is defensible
Every story was extracted from a dated Wayback Machine capture and carries a resolvable archive URL. Days were drawn with a documented seed, and the shipped sampling.py regenerates both the day draw and the validation draw from scratch. The story-selection rule is fixed and stated; duplicate captures were rejected; the deviations log reconciles exactly (62 + 4 + 4 = 70 cells). Two outlets could not be included, and — following an adversarial audit of the previous draft — their exclusions are documented with evidence rather than asserted: The Wire's archived CMS section feeds were examined and a reconstruction attempted before rejection (appendix); Times of India's single usable capture was parsed and its Indian-edition top-15 published (appendix). The dataset self-reports its uncertainty: 44 of 575 codes (7.7%) are flagged low-confidence, with a sensitivity check showing no headline finding depends on them. Full methods and deviations log: METHODS_v1.md.
575
Homepage stories
10
Outlets
62
Outlet-days
326
Unique events
7.7%
Codes flagged low-confidence
100%
Archive-verifiable
Headline Findings
1. Straight news dominates Indian homepages.Straight news is 45.0% [41.0–49.1] of stories; 42.6% [38.2–47.2] in the hard-news subset (n=455, excluding sports and entertainment); 47.5% among high-confidence codes only. Against PEJ's 16% for 1998 US front pages, Indian digital homepages are roughly three times as fact-report-framed. Caveats: headline-level coding skews toward inverted-pyramid readings, and the window was saturated with fast-moving spot news (US–Iran deal, Hormuz, Venezuela earthquake, World Cup).
2. Combative framing is low — a third of PEJ's US figure.Combative frames total 11.7% [9.3–14.5] (14.1% hard-news; 12.6% high-confidence) — half of what this study's earlier draft claimed, and far below PEJ's 30%. Horse race is 1.2% in a window with no major Indian election — evidence that combativeness findings are heavily election-conditional.
3. Consensus framing is vanishingly rare — with an honesty asterisk. One story in 575 (0.2% [0.0–1.0]): India Today's "US, Iran agree to halt strikes, hold talks in Doha." A fortnight dominated by a peace process — a story literally about adversaries agreeing — produced almost no consensus-framed coverage. Asterisk: that single story is itself one of the 44 low-confidence calls, so the defensible claim is "zero to one story in 575." Either way, the frame is functionally absent from Indian homepage journalism in this window.
4. The digital-native/legacy divide mostly does not replicate. Straight news: native 39.5% [32.7–46.6] vs legacy 47.7% [42.8–52.6] — overlapping intervals. FirstPost (digital-native, corporate-owned, 65.7% straight news) breaks the "natives interpret, legacy reports" story on its own.
5. "Corporate outlets produce zero accountability journalism" is falsified.Accountability frames: NDTV 7.1% [3.1–15.7], India Today 4.3%, News18 6.0%, HT 4.3% — small, not zero. The gradient survives in compressed form: Newslaundry 66.7% [41.7–84.8] (n=15, watchdog by design), Indian Express 21.4%, The Print 15.0%, Scroll 10.0%, FirstPost 2.9%. Native 13.5% vs legacy 7.9%, intervals overlapping.
6. Transparency–accountability correlation: modest, outlier-sensitive, now fully auditable. Spearman ρ = 0.71 across 10 outlets (Pearson 0.85); ρ = 0.60 without Newslaundry. The Rosen Index inputs are published (rosen_scores.json) with their limitations stated: unverified carryover scores, circularity risk, no contemporaneous per-criterion documentation. Read as a suggestive descriptive gradient on ten data points awaiting blind re-scoring — not a law, and nothing like the r = 0.92 once claimed.
7. Enterprise journalism is one-third accountability journalism. 31.0% [22.3–41.4] of 87 enterprise-triggered stories carry accountability frames; the modal enterprise frames are historical outlook (17) and trend (16). Indian enterprise reporting is substantially accountability work — not predominantly, as the earlier draft said.
8. Government triggers produce compliant framing. 68.8% [59.7–76.6] of government-triggered stories are straight-news framed; 13.4% combative. PEJ found US journalists most combative precisely on government triggers. Indian homepage journalism largely relays official action rather than contesting it — the sharpest accountability finding here, and it implicates native and legacy outlets alike.
9. What fills Indian homepages. Foreign affairs 24.0%, sports 13.7%, politics 12.7%, crime 8.0%. One murder case (Pune's Lohagad Fort case) generated 27 stories — 4.7% of the national English homepage sample — more than education (3.5%), health (1.7%), or environment (3.7%). Media/press freedom is 1.4% of the sample, and all 8 of those stories come from just two outlets — Newslaundry (4) and The Print (4). No legacy outlet put a single press-freedom story in its homepage top-10 across seven sampled days.
10. Event-level correction. 575 stories collapse to 326 events (biggest clusters: Pune murder 27, US–Iran deal 23, Ram temple donation scandal 20, Venezuela earthquake 14, Hormuz 12, NEET re-exam 11). Modal event-level frame is still straight news (144/326). Outlet fingerprints built on story-level counts partly measure which events fell in the window; both levels ship in the dataset.
Frame Distribution
Narrative Frames — All Outlets (N=575)
Story level vs Event level (%): does non-independence change the picture?
India 2026 vs. PEJ USA 1999
Frame Distribution: India 2026 (N=575) vs United States 1999 (PEJ) Directional only
Comparison discipline
The PEJ benchmark confounds country × era × medium (1998 US print front pages vs 2026 Indian digital homepages, full-text vs headline coding) and is used directionally only: more straight news than late-90s US front pages, less combative framing, consensus near-absent in both. Claims about "Indian digital journalism" in general are not supported by one fortnight, ten outlets, and headline-level coding — this report deliberately does not make them.
Topics, Triggers & Messages
Topic Distribution
Story Triggers: What Makes News
Underlying Messages 71% No Message
Trigger × Frame (top 6 triggers, % within trigger)
Outlet Fingerprints
Straight-News Share by Outlet (%)
Accountability Frames by Outlet (%)
The Ten Outlets
Outlet
Type
Stories (n)
Straight News %
Accountability %
Rosen Score Unverified
Newslaundry
Digital native
15
0.0%
66.7%
10/10
Indian Express
Legacy digital
70
24.3%
21.4%
7/10
The Print
Digital native
60
26.7%
15.0%
6/10
Scroll.in
Digital native
40
27.5%
10.0%
7/10
NDTV
Legacy digital
70
50.0%
7.1%
5/10
News18
Legacy digital
50
60.0%
6.0%
3/10
Hindustan Times
Legacy digital
70
40.0%
4.3%
5/10
India Today
Legacy digital
70
58.6%
4.3%
4/10
The Hindu
Legacy digital
60
58.3%
3.3%
6/10
FirstPost
Digital native
70
65.7%
2.9%
4/10
Newslaundry contributes only 15 stories (its archived pages server-render just the top items) and is a media-watchdog by design — its 66.7% carries a wide interval [41.7–84.8]. Rosen scores are published carryovers awaiting blind re-scoring: see rosen_scores.json for the criteria and the caveats.
Transparency & Accountability
Rosen Transparency Index vs Accountability Framing Descriptive only
What the correlation can and cannot say
Spearman ρ = 0.71 across ten outlets (Pearson 0.85). Remove the Newslaundry outlier and ρ = 0.60. That is a suggestive descriptive gradient on ten data points — not a law.
The Rosen scores were assigned before this release by a scorer aware of earlier framing results (circularity risk), per-criterion breakdowns were never contemporaneously documented, and they have not been independently verified. They are published anyway — auditable inputs beat asserted ones. Blind re-scoring by an independent rater is a pre-condition for any stronger claim.
The earlier draft's r = 0.92 and its "corporate outlets produce zero accountability journalism" claim did not survive the audit and are withdrawn.
The Event Level: What Actually Happened This Fortnight
Largest Event Clusters (stories per event)
Event
Stories
Share of sample
Note
Pune murder case (Lohagad Fort)
27
4.7%
One crime story outweighs education (3.5%), health (1.7%) and environment (3.7%)
US–Iran deal (Doha talks)
23
4.0%
The window's dominant diplomacy story — and source of the sample's single Consensus frame
Ram temple donation scandal
20
3.5%
Venezuela earthquake
14
2.4%
Strait of Hormuz tensions
12
2.1%
NEET re-exam
11
1.9%
India Today stale poll widget
7
1.2%
Cached "Punjab civic polls" widget at position 5 on all 7 days — flagged, one event
The Excluded Outlets: Evidence, Not Assertion
The Wire — 0 usable cells
The homepage is a ~9 KB client-rendered shell; archived homepage HTML contains no stories. The WordPress CMS section feeds were archived (604 captures in/around the window) and a reconstruction was attempted — then rejected: the feeds are recency-ordered section pages, a different measurement unit from homepage editorial prominence; coverage fails the sampling design (zero captures on two sampled days); and mixing one recency-feed outlet into a nine-outlet prominence sample would reintroduce the undocumented-heterogeneous-selection flaw this study exists to correct.
72 captures in the window: 71 are contentless HTTP 301 redirect stubs (~915–970 bytes — no page, no edition) and one (27 June) is a genuine Indian-edition content capture, whose top-15 stories are published in the appendix. The exclusion rests solely on coverage: one usable cell of seven against the ≥4-cell criterion. The earlier draft's "US-edition pages" claim was wrong and is corrected in METHODS §2.
Headline-only coding (frames inside story bodies invisible). Single AI coder; 7.7% of codes self-flagged low-confidence; findings preliminary until human κ is published (validation package shipped). One capture per day; 7 of 62 cells adjacent-day. Newslaundry contributes 15 stories. No TOI or Wire — exclusions evidenced in appendices; the native camp (n=185) is half the legacy camp (n=390). Two-week window; no seasonal or election-period claims. English-only. Rosen scores published but unverified. India Today's stale poll widget appears 7× (flagged; one event).
Next steps, in order
1. Human-code the 115-story validation sample; publish κ. 2. Blind re-scoring of the Rosen Index by an independent rater. 3. A second constructed week in a different news season, and one in an election window. 4. Prospective daily archiving of all twelve target homepages (including TOI and The Wire) so no future wave depends on crawl luck.
Data & Reproducibility Package
Everything needed to audit, recompute, or extend this study. Every record in the dataset carries a resolvable Wayback Machine URL.
AI codes for the validation sample — open only after coding
Glossary & Reading This Study Honestly
Complete Glossary of Terms
Hover over dotted-underlined terms throughout the dashboard for quick definitions.
Frame
The dominant narrative device a journalist uses to organise a story. Not what the story is about (that's topic), but how it is told. In this release the coding unit is the headline as displayed on the homepage (plus URL slug) — a disclosed departure from PEJ's full-text coding. Frame percentages here are homepage-presentation framing, not full-text framing.
Straight News
The classic inverted pyramid. No dominant narrative device other than presenting who, what, when, where, why, and how. The low-inference default: where no frame is clearly dominant in the headline, the story codes as Straight News.
Conflict
The story is built around the clash or disagreement inherent in a situation, foregrounding opposing sides, tensions, or disputes.
Consensus
The story is built around points of agreement among stakeholders. 0.2% of this sample [0.0–1.0] — and the single instance is itself a low-confidence call. PEJ's already-low US figure was 6%.
Conjecture
Speculation about what will happen next. Forward-looking, predictive, or scenario-based framing.
Process
An explanatory frame: how something works, the mechanics of a system, the steps involved.
Historical Outlook
The story places current events in historical context, using the past to illuminate the present.
Horse Race
Who is winning and who is losing. 1.2% in this window — which contained no major Indian election.
Trend
The news is presented as part of an ongoing, larger pattern.
Policy Explored
The journalist dives into the substance of a policy — mechanisms, impact, trade-offs, implementation.
Reaction
The story is structured around responses from key players to an earlier event.
Reality Check
The journalist examines the veracity of a claim, statement, or widely held belief.
Wrongdoing Exposed
The story uncovers injustice, corruption, malfeasance, or misconduct — revelation with identifiable actors.
Personality Profile
The story is built around an individual — their character, background, motivations.
Institutional Critique
India-specific frame added to the PEJ codebook. Diagnoses systemic failure without necessarily naming individual wrongdoers. "Why does NEET keep failing?" is institutional critique; "CBI arrests teacher in NEET leak" is wrongdoing exposed.
Trigger
What caused the news organisation to cover this story today — a statement, a spontaneous event, a report, or the newsroom's own initiative. Reveals whether journalism is reactive or proactive.
Underlying Message
An often unconscious cultural or moral narrative embedded in the story ("perseverance pays off", "the system is broken"). 71% of headlines carry no detectable message — headline-level coding reads fewer messages than full-text coding would.
Constructed Week
A sampling technique from media research: one Monday, one Tuesday, etc., drawn at random so no single news cycle dominates. This release is honestly labelled a constructed week within one fortnight (18 June – 2 July 2026, seed 20260618): it controls weekday composition but not seasonality, and single-event domination is handled at event level instead.
Event Level
Stories sharing an event ID cover the same underlying occurrence. 575 stories collapse to 326 events. Constructed-week sampling across 10 outlets makes story-level rows non-independent, so every headline finding is reported at both levels.
Low-Confidence Flag
44 of 575 codes (7.7%) self-flagged under stated criteria: digest/roundup items coded by lead item, interrogative headlines carrying inferential frames, and calls where two or more frames were defensible from the headline alone. Sensitivity check: excluding all 44 moves straight news from 45.0% to 47.5% and combative from 11.7% to 12.6% — no headline finding changes.
Digital Native
Born on the internet, no prior print or broadcast edition. In this release: Scroll.in, The Print, Newslaundry, FirstPost (n=185).
Legacy Digital
Traditional print or broadcast organisations that expanded to digital. In this release: NDTV, India Today, Hindustan Times, Indian Express, News18, The Hindu (n=390).
Combative Frames
PEJ grouping: Conflict + Horse Race + Wrongdoing Exposed. PEJ found 30% on US front pages in 1999; this study finds 11.7% [9.3–14.5].
Accountability Frames
Wrongdoing Exposed + Reality Check + Institutional Critique — frames that perform a watchdog or verification function.
Rosen Transparency Index
A 0–10 outlet score after Jay Rosen's scholarship on journalistic authority, on five criteria (each 0–2): funding/ownership disclosure, corrections policy, stated editorial stance, proprietor–editorial separation, reader accountability mechanisms. Published in rosen_scores.json with caveats; awaiting blind re-scoring.
Wilson Interval
The 95% confidence interval used throughout — well-behaved for small samples and extreme proportions.
Cohen's κ
Chance-corrected agreement between two coders. The pending human validation of the 115-story subsample will yield κ per variable; all findings are preliminary until it is published.
Methodology in brief
Sampling: From the 15-day window 18 June – 2 July 2026, one date per weekday drawn at random with seed 20260618 (Sat Jun 20, Sun Jun 21, Thu Jun 25, Fri Jun 26, Mon Jun 29, Tue Jun 30, Wed Jul 1). For each outlet-day, the Wayback Machine capture closest to 12:00 IST, else nearest within ±1 day (flagged). First 10 editorial story links in the main content area, in DOM/JSON order, deduplicated, with pre-specified exclusions. 62 usable cells of 70 possible; the reconciliation is exact (62 populated + 4 duplicate-capture + 4 no-capture = 70).
Coding: Four variables per story — Frame (14), Topic (16), Trigger (13), Underlying Message (9) — using PEJ Framing the News (1998) definitions plus Institutional Critique. Coding unit = headline as displayed. Single AI coder with a systematic self-audit pass; 20% seeded human-validation subsample shipped.
Full detail, deviations log (8 entries), and the audit trail of the adversarial review this release passed through: METHODS_v1.md.
Reading this study honestly
"Headlines aren't stories — you're coding packaging, not journalism."
Correct, and disclosed as the study's first limitation. The coding unit is the headline as displayed on the homepage plus the URL slug. What this measures is homepage-presentation framing — the storytelling promise an outlet makes at its front door — not full-text framing. Headline-level coding also skews toward inverted-pyramid readings, which is stated alongside the straight-news finding it inflates.
"A single AI coder can't be trusted."
Which is why the release ships a seeded, blind 115-story validation sample for human coding, flags its own 44 least-confident calls, publishes a sensitivity check showing no headline finding depends on them, and labels every finding preliminary until κ is published. Distrust is the correct default; the package is built so the distrust can be resolved empirically.
"Two weeks in one news season proves nothing about Indian journalism."
Agreed — and the report deliberately makes no such claim. The window was saturated with fast-moving spot news; horse race at 1.2% in a non-election window shows how season-conditional these figures are. The claim is about this fortnight's homepages, stated with intervals. Wider claims wait for a second constructed week in a different season and one in an election window.
"You dropped outlets until the story fit."
The two exclusions are evidenced, not asserted: The Wire's archived CMS feeds were examined and a reconstruction attempted before rejection on stated grounds; TOI's 72 captures were classified and the single usable one parsed and published. The ≥4-cell inclusion criterion was formulated after seeing archive coverage — disclosed as post-hoc — but applied uniformly: all ten retained outlets have 5–7 usable cells; the exclusions have 1 and 0.
"The earlier version of this page said something different."
It did, and it was wrong in ways an adversarial audit made specific: an unreproducible sample, unpublished Rosen inputs, an exclusion narrative contradicted by the archive's own logs, a dead low-confidence field, and headline claims (r = 0.92; "corporate outlets produce zero accountability journalism") that did not survive recomputation. That version is preserved at /superseded/; the audit findings and their dispositions are logged in METHODS_v1.md §8. This is what correction is supposed to look like.
Academic References
Project for Excellence in Journalism & Princeton Survey Research Associates. (1999). Framing the News: The Triggers, Frames, and Messages in Newspaper Coverage. Pew Research Center.
Rosen, J. (2003–present). PressThink: Ghost of Democracy in the Media Machine. pressthink.org. Key concepts: "The View from Nowhere" (2010), "The People Formerly Known as the Audience" (2006).
Entman, R.M. (1993). "Framing: Toward Clarification of a Fractured Paradigm." Journal of Communication, 43(4), 51-58.
Scheufele, D.A. (1999). "Framing as a Theory of Media Effects." Journal of Communication, 49(1), 103-122.
Wilson, E.B. (1927). "Probable Inference, the Law of Succession, and Statistical Inference." JASA, 22(158), 209-212.