Everybody’s talking about equity, but is anyone really listening?: The Case for Better Data-Driven Learning in Health Systems
Bibliographic record
Abstract
Data collection, analysis, and data driven action cycles have been viewed as vital components of healthcare for decades. Throughout the COVID-19 pandemic, case incidence and mortality data have consistently been used by various levels of governments and health institutions to inform pandemic strategies and service distribution. However, these responses are often inequitable, underscoring pre-existing healthcare disparities faced by marginalized populations. This has prompted governments to finally face these disparities and find ways to quickly deliver more equitable pandemic support. These rapid data informed supports proved that learning health systems (LHS) could be quickly mobilized and effectively used to develop healthcare actions that delivered healthcare interventions that matched diverse populations' needs in equitable and affordable ways. Within LHS, data are viewed as a starting point researchers can use to inform practice and subsequent research. Despite this innovative approach, the quality and depth of data collection and robust analyses varies throughout healthcare, with data lacking across the quadruple aims. Often, large data gaps pertaining to community socio-demographics, patient perceptions of healthcare quality and the social determinants of health exist. This prevents a robust understanding of the healthcare landscape, leaving marginalized populations uncounted and at the sidelines of improvement efforts. These gaps are often viewed by researchers as indication that more data is needed rather than an opportunity to critically analyze and iteratively learn from multiple sources of pre-existing data. This continued cycle of data collection and analysis leaves one to wonder if healthcare has a data problem or a learning problem. In this commentary, we discuss ways healthcare data are often used and how LHS disrupts this cycle, turning data into learning opportunities that inform healthcare practice and future research in real time. We conclude by proposing several ways to make learning from data just as important as the data itself.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.388 | 0.417 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.004 |
| Bibliometrics | 0.007 | 0.009 |
| Science and technology studies | 0.013 | 0.094 |
| Scholarly communication | 0.048 | 0.126 |
| Open science | 0.010 | 0.042 |
| Research integrity | 0.027 | 0.057 |
| Insufficient payload (model declined to judge) | 0.015 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".