Using Linked Data and Advanced Analytics to Prioritize Health Concerns within Regions
Bibliographic record
Abstract
IntroductionLinked population health data have the potential to inform evidence-based actions targeting serious public health concerns. However, large-scale data integration efforts can produce hundreds of population health indicators, which can overwhelm the ability of decision-makers to synthesize and interpret the information.
 Objectives and ApproachOur research uses an existing semantic web application for population health surveillance, the Population Health Record (PopHR). PopHR automates a computational pipeline for linking data sources, building timely population health indicators, and uses artificial intelligence to organize indicators along a determinants of health framework. To assist users in interpreting the thousands of indicators, we developed computational algorithms combining values of multiple indicators across chronic diseases, to prioritize conditions within each region. This analytic approach can assist regional decision-makers in identifying their region’s priority conditions by facilitating the integration and analysis of multiple types of indicators (e.g. disease burden, temporal patterns).
 ResultsA pilot implementation of the regional prioritization algorithm focused on indicators defined in the Public Health Agency of Canada’s Chronic Disease Indicators Framework. Within this subset of diseases, we developed a computational algorithm to integrate into a priority index regional estimates of incidence, mortality, and prevalence taking into account the relative importance of each indicators’ outlier status and statistical significance of temporal trends. Our results allowed for the development of region-specific data visualizations dashboards, emphasizing the different factors driving the rankings of indicators within and across regions. For example, regions with higher socioeconomic status having generally lower disease burden are presented with visualizations emphasizing temporal trends and other statistically compelling patterns rather than simple indicators of magnitude.
 Conclusion/ImplicationsThis ranking approach represents initial stages ongoing research, expanding our methods to use machine learning strategies and additional expert knowledge. Current and future prioritization analyses within the PopHR platform offer the potential for public health to gain insights from an otherwise challenging complexity and richness of linked data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".