Is a Trial of Perioperative Cognitive Training to Prevent Early Postoperative Cognitive Decline Actually Feasible?
Notice bibliographique
Résumé
To the Editor In the recent randomized-controlled feasibility trial reported by O’Gara et al,1 the authors concluded that based on their study’s findings, “a future trial of cognitive prehabilitation before cardiac surgery in this population is likely to be feasible.” While we share the authors’ overall enthusiasm for pursuing this field of investigation, potential limitations related to their nuanced trial design, conduct, and subsequent analysis make it difficult to echo their optimism regarding feasibility. Several aspects of the trial require careful consideration to inform the interpretation of the study’s results. It appears from the authors’ ClinicalTrials.gov registration (NCT02908464) that this study was initially designed as an efficacy trial, where the primary and secondary outcome measures were related to the incidence of postoperative delirium and postoperative cognitive dysfunction (POCD). Specifically, they described an intent to examine the impact of a neurocognitive training program using a mobile software application designed to train users in the cognitive domains of memory, attention, problem solving, flexibility, and processing speed (Luminosity; Lumos Labs, Inc, San Francisco, CA). This tablet-based platform was to be administered to patients preoperatively as well as at several time points postoperatively. However, it appears from the trial’s registration details (and subsequent protocol publication2) that the primary outcome measure was modified after commencement of enrollment from one related to efficacy (delirium and POCD incidences) to a one of feasibility (focused on enrollment and adherence to the trial’s protocol). While the authors are to be commended for updating the ClinicalTrials.gov registry, it would have been more transparent had they explicitly described this major change in study design in the article’s methods section itself. Indeed, the original registration lists no feasibility endpoints whatsoever and the postenrollment change in design is incomplete in its description of the feasibility endpoints that were subsequently examined. For example, the recruitment endpoint was defined a priori as “enrollment of ≥50% of eligible patients,” whereas the adherence thresholds were not defined and appear to be analyzed post hoc. To illustrate, the authors reported that only 39% of the patients were able to complete the preoperative training, 6% postoperatively, and 19% postdischarge. However, without having a priori defined the anticipated percent completion of the training necessary to meet their feasibility threshold, readers are unable to assess if the adherence objective was truly met. Further confusing the issue of interpreting their enrollment-related feasibility primary endpoint are the differences between the study’s prespecified definition (the percentage of eligible patients enrolled; 45 of 169 [27%]) and reported definition (the percentage of approached patients enrolled; 45 of 69 [65%]). Using the original definition would change the conclusions that a subsequent trial is not feasible, which is contrary to what the authors actually conclude. It is also notable that the reported feasibility endpoint in the published protocol2 (where it is defined as “recruitment”) also differs from that registered (where it is “eligibility” for recruitment). The efficacy (secondary) endpoints also need careful consideration. Delirium was assessed using the Confusion Assessment Method (CAM) and POCD was defined as one standard deviation (SD) decline in the postoperative Montreal Cognitive Assessment (MoCA) score compared to baseline. However, a one SD change from baseline cognition, although stated by the authors as being consistent with other publications, has not generally been reported when using the MoCA as a measure of POCD. Irrespective of this, and recognizing the lack of power for this secondary endpoint of efficacy, it also appears questionable if the MoCA was sufficiently sensitive to detect POCD in the context of this trial. For example, the telephonic MoCA scores did not change among any of the groups between the preoperative to postoperative time points, raising concern regarding a lack of intrinsic validity. Thus, if one extends the feasibility aspects of the study to the assessment of POCD per se, then additional uncertainty as to their conclusions arises. A further limitation in this study relates to the final objective in the study’s introduction, where the authors state that it can be used to “estimate effect sizes of postoperative delirium, and POCD to inform the conduct of future trials.”1 We would argue that small feasibility trials that also capture secondary clinical endpoint data are unreliable to actually inform future trials’ endpoints and effect sizes. That is, small trials such as this tend to overestimate effect sizes. This is partly responsible for why many large efficacy trials that have been preceded by smaller trials have failed to replicate the same results. Furthermore, if one is going to use a feasibility trial such as this to inform future trials as to the incidence of the endpoint being examined (eg, delirium or POCD), then one needs to enroll sufficient numbers to have a sufficiently narrow estimate of precision. Finally, the effect size of any intervention (such as cognitive training) is also usually overestimated by that ascertained from a small-size trial (common with feasibility trials). An alternate more conservative approach might be to use the lower margin of the 95% confidence interval of the effect size’s point estimate in the sample size calculation for a future trial. In summary, as to whether a trial such as this is likely to be feasible, we would suggest that at best, the results are inconclusive, and at worst point to it being not feasible at all. Only by repeating a similar trial that modifies the parameters that prevent the interpretation of the feasibility endpoints in the present pilot study will one know with any confidence. Hilary P. Grocott, MD, FRCPC, FASEDepartment of AnesthesiologyPerioperative and Pain MedicineUniversity of ManitobaWinnipeg, Manitoba, Canada[email protected]Stephan K. W. Schwarz, MD, PhD, FRCPCDepartment of AnesthesiologyPharmacology & TherapeuticsThe University of British ColumbiaVancouver, British Columbia, Canada
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,026 | 0,146 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,003 | 0,002 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,003 | 0,003 |
| Science ouverte | 0,004 | 0,001 |
| Intégrité de la recherche | 0,013 | 0,016 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,008 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».