The private versus public contribution to the biomedical literature during the COVID-19, Ebola, H1N1, and Zika public health emergencies
Bibliographic record
Abstract
BACKGROUND: The private versus public contribution to developing new health knowledge and interventions is deeply contentious. Proponents of commercial innovation highlight its role in late-stage clinical trials, regulatory approval, and widespread distribution. Proponents of public innovation point out the role of public institutions in forming the foundational knowledge undergirding downstream innovation. The rapidly evolving COVID-19 situation has brought with it uniquely proactive public involvement to characterize, treat, and prevent this novel health treat. How has this affected the share of research by industry and public institutions, particularly compared to the experience of previous pandemics, Ebola, H1N1 and Zika? METHODS: Using Embase, we categorized all publications for COVID-19, Ebola, H1N1 and Zika as having any author identified as affiliated with industry or not. We placed all disease areas on a common timeline of the number of days since the WHO had declared a Public Health Emergency of International Concern with a six-month lookback window. We plotted the number and proportion of publications over time using a smoothing function and plotted a rolling 30-day cumulative sum to illustrate the variability in publication outputs over time. RESULTS: Industry-affiliated articles represented 2% (1,773 articles) of publications over the 14 months observed for COVID-19, 7% (278 articles) over 7.1 years observed for Ebola, 5% (350 articles) over 12.4 years observed for H1N1, and 3% (160 articles) over the 5.7 years observed for Zika. The proportion of industry-affiliated publications built steadily over the time observed, eventually plateauing around 7.5% for Ebola, 5.5% for H1H1, and 3.5% for Zika. In contrast, COVID-19's proportion oscillated from 1.4% to above 2.7% and then declined again to 1.7%. At this point in the pandemic (i.e., 14 months since the PHEIC), the proportion of industry-affiliated articles had been higher for the other three disease areas; for example, the proportion for H1N1 was twice as high. CONCLUSIONS: While the industry-affiliated contribution to the biomedical literature for COVID is extraordinary in its absolute number, its proportional share is unprecedentedly low currently. Nevertheless, the world has witnessed one of the most remarkable mobilizations of the biomedical innovation ecosystem in history.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.062 | 0.226 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.055 | 0.081 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.012 | 0.010 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".