Precision, Equity, and Public Health and Epidemiology Informatics – A Scoping Review
Bibliographic record
Abstract
OBJECTIVES: This scoping review synthesizes the recent literature on precision public health and the influence of predictive models on health equity with the intent to highlight central concepts for each topic and identify research opportunities for the biomedical informatics community. METHODS: Searches were conducted using PubMed for publications between 2017-01-01 and 2019-12-31. RESULTS: Precision public health is defined as the use of data and evidence to tailor interventions to the characteristics of a single population. It differs from precision medicine in terms of its focus on populations and the limited role of human genomics. High-resolution spatial analysis in a global health context and application of genomics to infectious organisms are areas of progress. Opportunities for informatics research include (i) the development of frameworks for measuring non-clinical concepts, such as social position, (ii) the development of methods for learning from similar populations, and (iii) the evaluation of precision public health implementations. Just as the effects of interventions can differ across populations, predictive models can perform systematically differently across subpopulations due to information bias, sampling bias, random error, and the choice of the output. Algorithm developers, professional societies, and governments can take steps to prevent and mitigate these biases. However, even if the steps to avoid bias are clear in theory, they can be very challenging to accomplish in practice. CONCLUSIONS: Both precision public health and predictive modelling require careful consideration in how subpopulations are defined and access to data on subpopulations can be challenging. While the theory for both topics has advanced considerably, there is much work to be done in understanding how to implement and evaluate these approaches in practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.052 | 0.212 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.005 |
| Bibliometrics | 0.022 | 0.021 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.010 | 0.009 |
| Open science | 0.003 | 0.006 |
| Research integrity | 0.005 | 0.005 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".