Vulnerable voices: using topic modeling to analyze newspaper coverage of climate change in 26 non-Annex I countries (2010–2020)
Bibliographic record
Abstract
Abstract News media influence how climate change is represented, understood, and discussed in the public sphere. To date, media and climate change research has primarily focused on Annex I countries, or treated non-Annex I countries as a homogenous bloc, despite the global nature of climate change and its geographically uneven impacts. This study uses a mixed-method approach, combining machine learning (topic modeling), econometrics, and qualitative analyses, to investigate newspaper coverage of climate change in 26 non-Annex I countries. We compiled a dataset of 95 216 news articles (dated between 2010 and 2020 from 50 sources) in 26 lower-middle and upper-middle income non-Annex I countries. In line with previous research results, we find that most common topics represented are international governance of climate change, the economics of energy transitions, and the impacts of climate change. Advancing current research understanding, we also demonstrate heterogeneity in coverage between non-Annex I countries and discover that a country’s vulnerability to climate change is positively associated with the diversity of topics (based on an article-level entropy index) portrayed by its domestic news media outlets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".