First, Do No Harm (Gone Wrong): Total-Scale Analysis of Medical Errors Scientific Literature
Bibliographic record
Abstract
Objective: Medical errors represent a leading cause of patient morbidity and mortality. The aim of this study was to quantitatively analyze the existing scientific literature on medical errors in order to gain new insights in this important medical research area. Study Design: Web of Science (WoS) database was used to identify relevant publications, and bibliometric analysis was performed to quantitatively analyze the identified papers for prevailing research themes, contributing journals, institutions, countries, authors, and citation performance. Results: In total, 12,415 publications concerning medical errors were identified and quantitatively analyzed. The overall ratio of original research articles to reviews was 8.1:1, and temporal subset analysis revealed that the share of original research papers has been increasing over time. The United States contributed to nearly half (46.4%) of the total publications, and 8 of the top 10 most productive institutions were from the United States, with the remaining 2 located in Canada and the United Kingdom. Prevailing (frequently mentioned) and highly impactful (frequently cited) themes were: errors related to drugs/medications, applications related to medicinal information technology, errors related to critical/intensive care units, to children, and to mental conditions associated with medical errors (burnout, depression). Conclusions: The high prevalence of medical errors revealed from the existing literature indicates the high importance of future work invested in preventive approaches. Digital health technology applications are perceived to be of great promise to counteract medical errors, and further effort should be focused to study their optimal implementation in all medical areas, with special emphasis on critical areas such as intensive care and pediatric units.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.007 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".