Opportunity and accessibility: an environmental scan of publicly available data repositories to address disparities in healthcare decision-making
Bibliographic record
Abstract
BACKGROUND: Health disparities, starkly exposed and exacerbated by coronavirus disease 2019, pose a significant challenge to healthcare system access and health outcomes. Integrating health inequalities into health technology assessment calls for robust analytical methodologies utilizing disaggregated data to investigate and quantify the scope of these disparities. However, a comprehensive summary of population datasets that can be used for this purpose is lacking. The objective of this review was to identify publicly accessible health inequalities data repositories that are potential resources for healthcare decision-making and future health technology assessment submissions. METHODS: An environmental scan was conducted in June of 2023 of six international organizations (World Health Organization, Organisation for Economic Co-operation and Development, Eurostat, United Nations Inter-agency Group for Child Mortality Estimation, the United Nations Sustainable Development Goals, and World Bank) and 38 Organisation for Economic Co-operation and Development countries. The official websites of 42 jurisdictions, excluding non-English websites and those lacking English translations, were reviewed. Screening and data extraction were performed by two reviewers for each data repository, including health indicators, determinants of health, and health inequality metrics. The results were narratively synthesized. RESULTS: The search identified only a limited number of country-level health inequalities data repositories. The World Health Organization Health Inequality Data Repository emerged as the most comprehensive source of health inequality data. Some country-level data repositories, such as Canada's Health Inequality Data Tool and England's Health Inequality Dashboard, offered rich local insights into determinants of health and numerous health status indicators, including mortality. Data repositories predominantly focused on determinants of health such as age, sex, social deprivation, and geography. CONCLUSION: Interactive interfaces featuring data exploration and visualization options across diverse patient populations can serve as valuable tools to address health disparities. The data they provide may help inform complex analytical methodologies that integrate health inequality considerations into healthcare decision-making. This may include assessing the feasibility of transporting health inequality data across borders.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.087 | 0.319 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.036 | 0.040 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.009 | 0.014 |
| Open science | 0.003 | 0.012 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.010 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".