P22. Prognostic Predictive Model for the Development of Osteoarthritis using Electronic Medical Record Data
Notice bibliographique
Résumé
Background: As the most common joint disorder worldwide (1), osteoarthritis represents a growing concern for older adults. Prognostic predictive models (statistical models used to predict future disease development (2)) may enable the identification of patients at high risk of developing osteoarthritis, allowing for health and lifestyle modifications aimed at reducing the risk of disease development (3,4).\nMethods: For our project, we accessed the DELPHI (Deliver Primary Healthcare Information) database which contains de-identified electronic medical records of more than 60,000 primary care patients in Ontario (5,6). From these data, we constructed a retrospective cohort examining patients’ risk factors and followed them over time to observe incident cases of osteoarthritis. This retrospective cohort was used to develop and test prognostic predictive models, using methods such as logistic regression, to determine the models’ ability to predict development of osteoarthritis. Models were evaluated, examining both discrimination (AUC) and calibration (calibration plots), using a reserved portion of patient data.\nResults: A logistic regression model was built that predicts the incidence of osteoarthritis based on patient age, sex, Body Mass Index (BMI), osteoporosis status, and leg injury status (AUC: 0.73).\nDiscussion & Conclusion: By creating a prognostic predictive model for osteoarthritis, we aim to support primary health care practitioners in estimating an individual patient’s risk of osteoarthritis; thereby allowing practitioners and patients to create unique plans to address the patient’s personal risk factors.\nInterdisciplinary Reflection: This project is highly interdisciplinary as it spans the fields of epidemiology, statistics, health informatics, primary health care, and computer science.\nReferences:\n1. Lopez AD, Mathers CD, Ezzati M, Jamison DT, Murray CJL. Global and regional burden of disease and risk factors, 2001: systematic analysis of population health data. Lancet (London, England) [Internet]. 2006 May 27 [cited 2016 Feb 13];367(9524):1747–57. Available from: http://www.ncbi.nlm.nih.gov/pubmed/16731270\n2. Hendriksen JMT, Geersing GJ, Moons KGM, de Groot JAH. Diagnostic and prognostic prediction models. J Thromb Haemost [Internet]. 2013 Jun [cited 2016 Aug 10];11 Suppl 1:129–41. Available from: http://www.ncbi.nlm.nih.gov/pubmed/23809117\n3. Felson DT, Zhang Y, Anthony JM, Naimark A, Anderson JJ. Weight loss reduces the risk for symptomatic knee osteoarthritis in women. The Framingham Study. Ann Intern Med [Internet]. 1992 Apr 1 [cited 2016 Jun 23];116(7):535–9. Available from: http://www.ncbi.nlm.nih.gov/pubmed/1543306\n4. Felson DT. Weight and osteoarthritis. Am J Clin Nutr [Internet]. 1996 Mar [cited 2016 Jun 23];63(3 Suppl):430S–432S. Available from: http://www.ncbi.nlm.nih.gov/pubmed/8615335\n5. CPCSSN. DELPHI (Deliver Primary Healthcare Information) Project [Internet]. 2013. Available from: http://cpcssn.ca/regional-networks/delphi-deliver-primary-healthcare-information-project/\n6. Birtwhistle R, Keshavjee K, Lambert-Lanning A, Godwin M, Greiver M, Manca D, et al. Building a pan-Canadian primary care sentinel surveillance network: initial development and moving forward. J Am Board Fam Med [Internet]. 2009 Jan [cited 2016 May 19];22(4):412–22. Available from: http://www.ncbi.nlm.nih.gov/pubmed/19587256
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,009 | 0,033 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,003 | 0,002 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,008 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».