OSTEO-AI: A Systematic Review and Meta-Analysis of Artificial Intelligence Models for Osteoarthritis and Osteoporosis Detection and Prognosis
Notice bibliographique
Résumé
Introduction: Osteoarthritis (OA) and osteoporosis are leading degenerative bone diseases that diminish quality of life and impose significant socioeconomic costs. Traditional diagnostic approaches, including imaging and bone density assessments, often fail to detect disease in its early stages, delaying critical interventions. Emerging artificial intelligence (AI) techniques, particularly those employing machine learning (ML) and deep learning (DL), offer promising avenues for early detection and more accurate prognostication. Methods: We conducted a systematic review of AI models developed between 2018 and 2024, assessing their performance in diagnosing and predicting the progression of OA and osteoporosis. Studies utilizing supervised or unsupervised methods applied to imaging modalities (e.g., X-ray, MRI, DXA) or clinical data were included. We evaluated model accuracy, reliability, clinical applicability, and generalizability. Quality and risk of bias were assessed using a modified CLAIM framework, ensuring alignment with transparency, validity, and clinical integration standards. Results: Of 2,300 identified articles, 33 studies met the inclusion criteria. Top-performing models for OA reached up to 97% accuracy, with one study achieving an AUC of 0.93 for MRI-based progression prediction. For osteoporosis, the strongest models attained a C-index of 0.90 using DXA imaging, indicating robust fracture risk prediction. Nevertheless, many studies relied on geographically or demographically homogeneous datasets, limiting broader applicability. Only 15% included external validation, and a substantial proportion lacked interpretability features essential for clinical adoption. Discussion: AI-driven models outperformed conventional diagnostic tools in accuracy and early disease detection. However, the limited dataset diversity, infrequent external validation, and insufficient model interpretability pose barriers to clinical integration. The reliance on male-dominant datasets for osteoporosis and geographically narrow cohorts for OA underscores the need for broader data representation. Standardizing evaluation metrics and improving explainability will enhance cross-study comparisons and support adoption in practice. Conclusion: AI holds transformative potential for improving OA and osteoporosis diagnostics, facilitating earlier interventions, and informing personalized patient management. Future work should prioritize diverse, well-validated datasets; transparent, clinician-friendly interfaces; and standardized performance metrics. Addressing these challenges will enable AI to evolve from a promising innovation into a cornerstone of global musculoskeletal healthcare.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,009 | 0,003 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,004 | 0,000 |
| Bibliométrie | 0,003 | 0,005 |
| Études des sciences et des technologies | 0,000 | 0,002 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».