From machine learning to clinical practice: phenotypic clusters of anti-MDA5 antibody-positive dermatomyositis
Notice bibliographique
Résumé
A recent review written by McLeish et al. elegantly explains the use of machine learning techniques for the diagnosis, prognosis and treatment of idiopathic inflammatory myopathies (IIMs) [1]. The usefulness in clinical practice of such techniques, however, remains uncertain, as the authors themselves admit. A particular problem for the diagnosis, prognosis, and treatment of IIMs is their heterogeneity. Machine learning techniques have, according to McLeish et al., ‘the potential to effectively tackle the heterogeneity of IIMs, offering a promising avenue to enhance the accuracy of predicting disease progression and outcome.’ As some of their examples, the authors summarize studies that used machine learning to generate phenotypic clusters of patients with anti-melanoma differentiation-associated protein 5 (MDA5) antibody-positive dermatomyositis (MDA5-DM). These studies were also summarized in an influential recent review on MDA5-DM [2]. Given the relevance for clinical practice of these studies, they warrant a more critical review. MDA5-DM is a distinct and very heterogeneous type of IIM that can manifest with any combination of myopathy, arthritis, skin lesions and interstitial lung disease (ILD). In about a third of patients, the ILD is rapidly-progressive interstitial lung disease (RP-ILD) and refractory to treatment, being the main explanation of the 6-month mortality being 25%. Other patients experience a milder course of ILD or no ILD at all. Therefore, ‘accurately identifying patient subtypes and predicting their prognoses are crucial for improving patient outcomes’ [2]. Four retrospective studies aimed to do so by classifying patients with MDA5-DM into phenotypic clusters [3–6]. Supplemental Table 1 compares the four studies. One originated in France [3] and the other three in China [4–6]. The studies included patients with dermatomyositis or clinically amyopathic dermatomyositis, depending on different classification criteria. All had anti-MDA5 antibody. None required muscle biopsy. Their definitions of RP-ILD varied—lacking a commonly accepted definition—from worsening respiratory symptoms to respiratory failure. All used unsupervised clustering analysis and decision trees to generate three clusters. Figure 1 compares the clinical characteristics of the three clusters generated in the four studies. All studies identified a cluster with predominant RP-ILD. However, while 67–93% of patients in this cluster had RP-ILD, up to 61% in the other clusters had RP-ILD. The studies variably characterized the other two clusters as ‘typical dermatomyositis’, ‘rheumatologic’ or ‘vasculopathic’, which they based on, respectively, predominance of skin lesions and myopathy, arthritis or Raynaud’s phenomenon. However, as Fig. 1 shows, prevalences of these clinical characteristics differed more between studies within each cluster than between clusters. A comparison of three phenotypic clusters of patients with MDA5 dermatomyositis (MDA5-DM) as reported in four studies [3–6]. The three clusters are grouped horizontally; we grouped and named them according to the predominance of RP-ILD, arthritis or myopathy, respectively. Patient characteristics are grouped vertically with their prevalences scaled from 0 to 100%; we show characteristics that were reported in at least three of the four studies. The four studies are distinguished with shades of grey. Characteristics that were reported as statistically different between the three clusters in each study are indicated by circles in corresponding shades of grey. *Data not available. Studies on phenotypical clusters directly address the quintessential question when diagnosing a patient with MDA5-DM in clinical practice: will this patient develop RP-ILD? The answer to this question guides important initial decisions regarding treatment and monitoring. The clusters presented in these four studies give insufficient answer to this clinical question. The inclusion criteria, clinical characteristics and outcomes lacked clear and common definitions. A multitude of different clinical characteristics were used to generate the clusters. Most characteristics were similarly prevalent among different clusters, as we showed above. RP-ILD was used as one of these characteristics, while the practitioner has it as the unknown outcome that should be prevented. Different combinations of characteristics constituted the decision trees, predicting a patient’s cluster with a risk of misclassification between 17 and 30%. Although mortality was highest in the cluster with predominant RP-ILD, it varied widely among the other clusters. Rather than describing phenotypes, clusters and decision trees provide better guidance for clinical practice if they follow commonly accepted definitions of MDA5-DM and RP-ILD, use a few easily and objectively measured clinical characteristics known at the time of diagnosis, predict a patient’s prognosis reliably and are validated by replication. As a striking illustration of this viewpoint, one of the four studies demonstrated that phenotypic clusters performed equally well if they were based on lymphocyte count only instead of a multitude of characteristics [5]. Another study, not included in the review, showed similar accuracy of phenotypic clusters based on immune cell phenotypes [7]. A simple model to predict RP-ILD based on four simple clinical characteristics—sex, disease duration, C-reactive protein and anti-Ro52—was recently published [8]. Unfortunately, only one study was replicated [9]. The limitations of studies on phenotypic clusters of patients with MDA5-DM, which we have set out here, support the argument of McLeish et al. that sophisticated machine learning techniques hold promise for the integration of complex clinical data into reliable prediction only insofar the results of studies are comparable, reproducible and applicable [1]. Taken together, we call for collaboration and integration of different cohorts and studies, data and techniques, researchers and clinicians. MDA5-DM is a very heterogeneous IIM, with RP-ILD as its most feared and fatal outcome. Machine learning techniques have been used to generate phenotypic clusters that predict RP-ILD. Machine learning techniques have yet limited usefulness in clinical practice, warranting a more collaborative and integrative approach. None declared. May Y. Choi: consulted for Mitogendx, Celltrion, Organon, AstraZeneca, GSK, Werfen, Mallinkrodt Pharmaceutical. The data underlying this article are available in the article and in its online supplementary material. Jacob J.E. Koopman, MD, PhD, is a clinical immunologist and visiting scholar. Katherine A. Buhler, BHSc, is a research assistant and a health scientist experienced in bioinformatics. May Y. Choi, MD, MPH, FRCPC, is a rheumatologist, associate professor in medicine and associate director of MitogenDx Laboratory.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».