A digital twin technique using external observational data to reduce sample sizes in clinical trials on Alzheimer’s disease
Notice bibliographique
Résumé
Abstract Background Randomized placebo‐controlled trials (RCTs) are the gold standard to evaluate efficacy of new drug treatments for Alzheimer’s disease. For example, the United States FDA approved the brain amyloid‐targeting drug lecanemab following CLARITY AD, Biogen and Eisai’s Phase 3 RCT. However, recruiting enough participants for a high‐powered and demographically representative trial is difficult and expensive. Fortunately, historical patient data from existing external observational studies of a disease can help populate RCTs [Thorlund et al. (2020). https://doi.org/10.2147/CLEP.S242097 ]. We propose a new trial framework that uses an external study to source “digital twins” for each trial participant. Using computer‐simulated trials mimicking CLARITY AD’s demographics and 18‐month duration, we show that our digital twin trial (DTT) has increased power compared to a conventional RCT. Method A continuous time linear mixed model tracked CDRSB change‐from‐baseline (CDRSBΔbl) trajectories in 670 ADNI participants satisfying CLARITY AD inclusion criteria [clinicaltrials.gov/study/NCT03887455]. To simulate an RCT, we resampled and added noise to participants’ data, generating a desired sample size of “recruited” participants who we randomized 1:1 to “drug” and “placebo” groups. We calculated participants’ CDRSBΔbl scores at 18 months and simulated the drug effect as a 25% reduction in CDRSBΔbl. For each participant in our DTT, we used Gower’s distance on demographic and clinical baseline variables to identify 20 most‐similar real ADNI participants (the digital twins) from our original 670. Each original ADNI participant’s 18‐month CDRSBΔbl was calculated using the model. A z‐score was then calculated for each DTT participant’s 18‐month CDRSBΔbl relative to their digital twins. T‐tests were used to evaluate DTT drug vs. placebo group difference in mean z‐score and, separately, RCT group difference in mean 18‐month CDRSBΔbl. We simulated each trial 1,000 times. Power is the proportion of simulations with a statistically significant treatment group difference. Result Figure 1 shows that 90% power is reached with approximately 500 fewer recruited participants in simulated DTTs (∼1,600 participants) compared to RCTs (∼2,100 participants). Conclusion DTTs might require substantially fewer recruited participants to achieve the same power as conventional RCTs. This sample size reduction could facilitate recruitment for trials on Alzheimer’s and in rare diseases with low patient numbers.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,317 | 0,561 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,003 | 0,005 |
| Bibliométrie | 0,003 | 0,003 |
| Études des sciences et des technologies | 0,002 | 0,007 |
| Communication savante | 0,003 | 0,004 |
| Science ouverte | 0,005 | 0,007 |
| Intégrité de la recherche | 0,004 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,013 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».