Differences in sensationalism in international news media reporting of COVID-19: An exploratory analysis using the Global Public Health Intelligence Network (GPHIN) system
Bibliographic record
Abstract
Background: The Global Public Health Intelligence Network (GPHIN) is an event-based surveillance platform that collects thousands of pieces of open-source information, including international news media, across multiple languages on a daily basis. Analysts have observed that news media reporting in some languages tended to use more sensational wording to describe major health events. There has been minimal research exploring potential differences in sensationalism in international news media reporting to confirm these observations. Objective: This exploratory study assessed the differences in the level of sensationalism in early international news media reporting of COVID-19 through a mixed-methods analysis. Methods: Relevant news media articles received in GPHIN seven days following the Public Health Emergency of International Concern declaration of COVID-19 by the World Health Organization were extracted for screening and analysis. An adapted tool was used to measure the sensationalism of pandemic-related health news. Deductive thematic analysis was conducted to examine themes of sensationalism. Differences in prevalence of sensationalism in news media reporting by language and country/territory of publication were assessed. Sentiment analysis assessed the sentiment and emotional tone of the news media articles. Results: Of 951 news articles that met the eligibility criteria, 155 contained sensationalism. There were significant differences between languages (French, Russian and Spanish) and various domains of sensationalism. This study also found a more negative emotional tone in news media articles with sensationalism. Conclusion: This exploratory study showed that language has the potential to impact the perception of health events using more sensationalized language.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.043 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.006 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".