Reporting on scientific studies in newspaper articles on COVID-19 in Italy and the US: missing citations and information bubbles
Bibliographic record
Abstract
Newspapers are a major source of health information, including on COVID-19. Many news articles are based on press releases describing research published in scientific journals. We studied a sample of articles in US and Italian newspapers for their mention of non-pharmacological interventions (face masks and lockdown) and pharmacological ones, both effective and ineffective (convalescent plasma, hydroxychloroquine, ivermectin, vaccines and vitamin D), and whether these were mentioned in a favorable or unfavorable way. We checked for the presence or absence of explicit mentions or links of the primary source of scientific information. Finally, we analyzed whether there was a trend to form “information bubbles” where some opinions around different treatments cluster together.Of 480,819 news in the USA and 767,172 in Italy, vaccines, face masks and lockdown were the most mentioned interventions. Of the pharmacological interventions other than vaccines, ivermectin and hydroxychloroquine were more frequently mentioned in the USA than in Italy (5- and 6-fold, respectively) although, when analyzing a sample of 210 news returned from a search on COVID-19 mentioning a research publication, articles from the USA were less favorable than those in Italy. We also found that the frequency of articles with a negative stance on vaccines was very small, indicating that the main newspapers do not contribute significantly to vaccine hesitancy. There was also evidence of information bubbles, where articles with a favorable or unfavorable view of a non-approved drug had the same stance on other non-approved treatments.Of the 210 news articles analyzed, half specifically mentioned a scientific publication. However, a link (or a complete citation) to the original source was provided in only 16% of Italian newspapers as opposed to 58% of those in the USA. The results highlight the fact that often news stories do not cite the scientific article they are reporting on, thus allowing the reader to verify the original source. This weakness in the citation behavior is particularly evident in Italian newspapers. This study suggests that linking to the original research article, rather than basing the news story on a press release, would improve the trustworthiness of the news and the critical thinking in the reader.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.051 | 0.268 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.032 | 0.035 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".