Is a Trial of Perioperative Cognitive Training to Prevent Early Postoperative Cognitive Decline Actually Feasible?
Bibliographic record
Abstract
To the Editor In the recent randomized-controlled feasibility trial reported by O’Gara et al,1 the authors concluded that based on their study’s findings, “a future trial of cognitive prehabilitation before cardiac surgery in this population is likely to be feasible.” While we share the authors’ overall enthusiasm for pursuing this field of investigation, potential limitations related to their nuanced trial design, conduct, and subsequent analysis make it difficult to echo their optimism regarding feasibility. Several aspects of the trial require careful consideration to inform the interpretation of the study’s results. It appears from the authors’ ClinicalTrials.gov registration (NCT02908464) that this study was initially designed as an efficacy trial, where the primary and secondary outcome measures were related to the incidence of postoperative delirium and postoperative cognitive dysfunction (POCD). Specifically, they described an intent to examine the impact of a neurocognitive training program using a mobile software application designed to train users in the cognitive domains of memory, attention, problem solving, flexibility, and processing speed (Luminosity; Lumos Labs, Inc, San Francisco, CA). This tablet-based platform was to be administered to patients preoperatively as well as at several time points postoperatively. However, it appears from the trial’s registration details (and subsequent protocol publication2) that the primary outcome measure was modified after commencement of enrollment from one related to efficacy (delirium and POCD incidences) to a one of feasibility (focused on enrollment and adherence to the trial’s protocol). While the authors are to be commended for updating the ClinicalTrials.gov registry, it would have been more transparent had they explicitly described this major change in study design in the article’s methods section itself. Indeed, the original registration lists no feasibility endpoints whatsoever and the postenrollment change in design is incomplete in its description of the feasibility endpoints that were subsequently examined. For example, the recruitment endpoint was defined a priori as “enrollment of ≥50% of eligible patients,” whereas the adherence thresholds were not defined and appear to be analyzed post hoc. To illustrate, the authors reported that only 39% of the patients were able to complete the preoperative training, 6% postoperatively, and 19% postdischarge. However, without having a priori defined the anticipated percent completion of the training necessary to meet their feasibility threshold, readers are unable to assess if the adherence objective was truly met. Further confusing the issue of interpreting their enrollment-related feasibility primary endpoint are the differences between the study’s prespecified definition (the percentage of eligible patients enrolled; 45 of 169 [27%]) and reported definition (the percentage of approached patients enrolled; 45 of 69 [65%]). Using the original definition would change the conclusions that a subsequent trial is not feasible, which is contrary to what the authors actually conclude. It is also notable that the reported feasibility endpoint in the published protocol2 (where it is defined as “recruitment”) also differs from that registered (where it is “eligibility” for recruitment). The efficacy (secondary) endpoints also need careful consideration. Delirium was assessed using the Confusion Assessment Method (CAM) and POCD was defined as one standard deviation (SD) decline in the postoperative Montreal Cognitive Assessment (MoCA) score compared to baseline. However, a one SD change from baseline cognition, although stated by the authors as being consistent with other publications, has not generally been reported when using the MoCA as a measure of POCD. Irrespective of this, and recognizing the lack of power for this secondary endpoint of efficacy, it also appears questionable if the MoCA was sufficiently sensitive to detect POCD in the context of this trial. For example, the telephonic MoCA scores did not change among any of the groups between the preoperative to postoperative time points, raising concern regarding a lack of intrinsic validity. Thus, if one extends the feasibility aspects of the study to the assessment of POCD per se, then additional uncertainty as to their conclusions arises. A further limitation in this study relates to the final objective in the study’s introduction, where the authors state that it can be used to “estimate effect sizes of postoperative delirium, and POCD to inform the conduct of future trials.”1 We would argue that small feasibility trials that also capture secondary clinical endpoint data are unreliable to actually inform future trials’ endpoints and effect sizes. That is, small trials such as this tend to overestimate effect sizes. This is partly responsible for why many large efficacy trials that have been preceded by smaller trials have failed to replicate the same results. Furthermore, if one is going to use a feasibility trial such as this to inform future trials as to the incidence of the endpoint being examined (eg, delirium or POCD), then one needs to enroll sufficient numbers to have a sufficiently narrow estimate of precision. Finally, the effect size of any intervention (such as cognitive training) is also usually overestimated by that ascertained from a small-size trial (common with feasibility trials). An alternate more conservative approach might be to use the lower margin of the 95% confidence interval of the effect size’s point estimate in the sample size calculation for a future trial. In summary, as to whether a trial such as this is likely to be feasible, we would suggest that at best, the results are inconclusive, and at worst point to it being not feasible at all. Only by repeating a similar trial that modifies the parameters that prevent the interpretation of the feasibility endpoints in the present pilot study will one know with any confidence. Hilary P. Grocott, MD, FRCPC, FASEDepartment of AnesthesiologyPerioperative and Pain MedicineUniversity of ManitobaWinnipeg, Manitoba, Canada[email protected]Stephan K. W. Schwarz, MD, PhD, FRCPCDepartment of AnesthesiologyPharmacology & TherapeuticsThe University of British ColumbiaVancouver, British Columbia, Canada
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.026 | 0.146 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.004 | 0.001 |
| Research integrity | 0.013 | 0.016 |
| Insufficient payload (model declined to judge) | 0.008 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".