Identification of people with low prevalence diseases in administrative healthcare records: A case study of HIV in British Columbia, Canada
Notice bibliographique
Résumé
<h4>Introduction</h4> Case-finding algorithms can be applied to administrative healthcare records to identify people with diseases, including people with HIV (PWH). When supplementing an existing registry of a low prevalence disease, near-perfect specificity helps minimize impacts of adding in algorithm-identified false positive cases. We evaluated the performance of algorithms applied to healthcare records to supplement an HIV registry in British Columbia (BC), Canada. <h4>Methods</h4> We applied algorithms based on HIV-related diagnostic codes to healthcare practitioner and hospitalization records. We evaluated 28 algorithms in a validation sub-sample of 7,124 persons with positive HIV tests (2,817 with a prior negative test) from the STOP HIV/AIDS data linkage–a linkage of healthcare, clinical, and HIV test records for PWH in BC, resembling a disease registry (1996–2020). Algorithms were primarily assessed based on their specificity–derived from this validation sub-sample–and their impact on the estimate of the total number of PWH in BC as of 2020. <h4>Results</h4> In the validation sub-sample, median age at positive HIV test was 37 years (Q1: 30, Q3: 46), 80.1% were men, and 48.9% resided in the Vancouver Coastal Health Authority. For all algorithms, specificity exceeded 97% and sensitivity ranged from 81% to 95%. To supplement the HIV registry, we selected an algorithm with 99.89% (95% CI: 99.76% - 100.00%) specificity and 82.21% (95% CI: 81.26% - 83.16%) sensitivity, requiring five HIV-related healthcare practitioner encounters or two HIV-related hospitalizations within a 12-month window, or one hospitalization with HIV as the most responsible diagnosis. Upon adding PWH identified by this highly-specific algorithm to the registry, 8,774 PWH were present in BC as of March 2020, of whom 333 (3.8%) were algorithm-identified. <h4>Discussion</h4> In the context of an existing low prevalence disease registry, the results of our validation study demonstrate the value of highly-specific case-finding algorithms applied to administrative healthcare records to enhance our ability to estimate the number of PWH living in BC.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».