Using Linked Data and Advanced Analytics to Prioritize Health Concerns within Regions
Notice bibliographique
Résumé
IntroductionLinked population health data have the potential to inform evidence-based actions targeting serious public health concerns. However, large-scale data integration efforts can produce hundreds of population health indicators, which can overwhelm the ability of decision-makers to synthesize and interpret the information. Objectives and ApproachOur research uses an existing semantic web application for population health surveillance, the Population Health Record (PopHR). PopHR automates a computational pipeline for linking data sources, building timely population health indicators, and uses artificial intelligence to organize indicators along a determinants of health framework. To assist users in interpreting the thousands of indicators, we developed computational algorithms combining values of multiple indicators across chronic diseases, to prioritize conditions within each region. This analytic approach can assist regional decision-makers in identifying their region’s priority conditions by facilitating the integration and analysis of multiple types of indicators (e.g. disease burden, temporal patterns). ResultsA pilot implementation of the regional prioritization algorithm focused on indicators defined in the Public Health Agency of Canada’s Chronic Disease Indicators Framework. Within this subset of diseases, we developed a computational algorithm to integrate into a priority index regional estimates of incidence, mortality, and prevalence taking into account the relative importance of each indicators’ outlier status and statistical significance of temporal trends. Our results allowed for the development of region-specific data visualizations dashboards, emphasizing the different factors driving the rankings of indicators within and across regions. For example, regions with higher socioeconomic status having generally lower disease burden are presented with visualizations emphasizing temporal trends and other statistically compelling patterns rather than simple indicators of magnitude. Conclusion/ImplicationsThis ranking approach represents initial stages ongoing research, expanding our methods to use machine learning strategies and additional expert knowledge. Current and future prioritization analyses within the PopHR platform offer the potential for public health to gain insights from an otherwise challenging complexity and richness of linked data.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,017 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,006 | 0,005 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,006 | 0,004 |
| Science ouverte | 0,001 | 0,004 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».