Facilitating Access to Effective Medical Checklists: A Literature Review and Dataverse Expansion Project
Bibliographic record
Abstract
Medical settings frequently employ checklists to enhance patient care and clinical processes. However, challenges such as poor checklist design and overreliance can limit the effectiveness of checklist interventions. We expanded a collection of checklists hosted on Borealis, a Canadian academic Dataverse repository, to promote widespread access to tested medical checklists. We collected checklists by searching through the Clinicaltrials.gov, PubMed, and CINAHL databases for relevant articles. Articles were imported into Covidence and screened for eligibility by two reviewers. Articles that studied checklists as the primary intervention and had performance metrics to support the conclusions they made about their outcomes were eligible for inclusion. From 725 retrieved articles, 80 were selected for full-text review after title and abstract screening and 25 met the inclusion criteria. Most studies (n=20) reported that checklists improved outcomes by increasing worker task adherence and care quality, while reducing patient morbidity and readmission rates. However, three studies found checklists to be ineffective and two had inconclusive results. Ultimately, 18 articles describing 16 checklists were added to the Borealis repository; two articles had already been added in the past. The articles added included checklists on safe childbirths, anesthesia administration, and cardiopulmonary resuscitation training. The Borealis repository also had its searchability improved with the addition of keyword filters and Medical Subject Heading (MeSH) terms. There are now a total of 34 checklists available in the repository. As medical institutions often create checklists to support their own processes, enabling widespread access to these can prove beneficial to the greater healthcare community.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.161 | 0.288 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.007 | 0.008 |
| Bibliometrics | 0.134 | 0.092 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.009 | 0.011 |
| Open science | 0.006 | 0.013 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.011 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".