Applying the fragility index to randomized controlled trials evaluating total neoadjuvant therapy for rectal cancer.
Notice bibliographique
Résumé
32 Background: A relatively novel summary measure that is not commonly reported in randomized controlled trials (RCTs) is the fragility index (FI). FI describes the number of additional events required for an outcome to lose statistical significance. It has been applied to a number of medical subspecialities such as critical care, orthopedic surgery and colorectal surgery. Recently, there has been significant interest in, and adoption of, total neoadjuvant therapy (TNT) for locally advanced rectal cancer (LARC). A number of RCTs have assessed TNT, but the robustness of these practice changing RCTs has never been evaluated. As such, we designed the present study to assess the robustness of the RCTs evaluating TNT for LARC using the FI. Methods: Relevant articles were identified through a recently published review article by Johnson et al. in the Canadian Journal of Surgery, that narratively reviewed all of the previously published RCTs evaluating TNT for LARC. We manually searched Google Scholar and PubMed to identify any other relevant RCTs. Outcomes within these RCTs that were either dichotomous outcomes or time to event outcomes were eligible for inclusion if the reported effect size had an associated p-value of less than 0.05. The main outcome was the FI for each statistically significant outcome. Walsh et al.’s method of calculating FI was utilized. A RCTs results were considered fragile if the FI was less than the loss to follow up for a given outcome. Correlations between FI and research characteristics were assessed using the Spearman’s rank correlation coefficients. Results: Ten RCTs were identified with 25 outcomes having statistically significant differences between groups (p-values < 0.05). Eleven outcomes were time-to-event outcomes, while the remainder were dichotomous outcomes. About half (n=13) were oncologic outcomes (i.e., survival, recurrence), while the rest (n=12) were short- and long-term complications. The median FI was 2 (interquartile range [IQR] 1-16). The number of patients lost to follow-up exceeded the FI in 17 outcomes (68.0%) and thus these results were considered “fragile”. Lower FI was associated with high risk of bias (rho=-0.5594) and higher loss to follow-up (i.e., greater than 5% vs. less than 5%) (rho=-0.4394), while higher FI was associated with large sample size (i.e., greater than 500 patients vs. less than 500 patients) (rho=0.5120). Conclusions: The robustness of outcomes from trials assessing TNT for LARC was found to be questionable. Most of these outcomes were fragile, as determined by the FI. In most cases, two or less additional events would have resulted in a loss of statistical significance of the reported results. Those using the results of these studies, including clinicians and health policy experts, should apply caution when interpreting these types of trials.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,343 | 0,702 |
| Méta-épidémiologie (sens strict) | 0,003 | 0,001 |
| Méta-épidémiologie (sens large) | 0,014 | 0,029 |
| Bibliométrie | 0,027 | 0,023 |
| Études des sciences et des technologies | 0,001 | 0,004 |
| Communication savante | 0,008 | 0,007 |
| Science ouverte | 0,004 | 0,005 |
| Intégrité de la recherche | 0,005 | 0,005 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,010 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».