Comparing the predictive accuracy of frailty instruments applied to preoperative electronic health data for adult patients undergoing non-cardiac surgery: a retrospective cohort study.
Notice bibliographique
Résumé
Introduction Twenty to 50% of older surgical patients live with frailty, a multidimensional state characterized by increased vulnerability to stressors due to accumulation of age- and disease-related deficits.1,2 Frailty is a strong perioperative risk factor associated with more than a doubling in the odds of postoperative morbidity, patient-reported disability and mortality, as well as increased healthcare resource use.3–5 Importantly, frailty appears to provide novel prognostic information when assessed in addition to risk factors typically identified before surgery (e.g., age, sex, ASA score, procedural risk), or when added to high-performing multivariable risk models.2,6–8 Preoperative frailty assessment may also represent a novel opportunity to drive optimization of underlying physical, nutritional, and cognitive deficits before surgery.9 Accordingly, routine preoperative frailty assessment is recommended by multiple international, multi-specialty best practice guidelines.6,10,11 Identification and communication of frailty status before surgery is associated with decreased postoperative mortality.12 However, routine frailty assessment in older patients appears to be rarely performed in practice.13,14 Barriers to preoperative frailty assessment are likely multifactorial and include the lack of a single or best-performing instrument as well as added time required to perform an assessment. Automated electronic frailty assessment could help to overcome such barriers by applying a prognostically accurate frailty instrument to preoperative electronic health data. While a recent systematic review identified 22 different electronic frailty instruments studied in the perioperative setting, most did not adhere to consensus definitions of frailty and accepted conceptual frameworks.15–17 Furthermore, among multi-dimensional instruments, there were no head-to-head comparisons performed and little data was available to describe predictive accuracy or added accuracy compared to risk factors that are typically assessed. This means that clinicians and health system planners have limited data to guide the development or implementation of an electronic approach to preoperative frailty assessment. The Frailty Index (FI), Risk Analysis Index (RAI), Adjusted Clinical Groups Frailty Defining Diagnoses indicator (ACG) and Hospital Frailty Risk Score (HFRS) all represent well-studied, multi-dimensional frailty instruments that can be applied to electronic health data. However, they have not been systematically compared when predicting postoperative outcomes relevant to older surgical patients. Therefore, using two population-based, non-cardiac surgery cohorts, our objectives are to: 1) determine the predictive accuracy of each of these frailty instruments in predicting postoperative outcomes (primary: 30-day mortality, secondary: days alive at home (within 30- and 365-days of surgery), length of hospital stay, health systems costs (within 30- and 365-days of surgery), discharge destination and one-year mortality); 2) determine the added predictive accuracy of each instrument beyond that provided by typically assessed risk factors; and 3) compare the predictive accuracy of each instrument head-to-head. Methods and Analysis Design and Data Sources This will be a retrospective, population-based cohort study using linked health administrative data from Ontario, Canada. All relevant data will come from ICES, an independent research institute where data are extracted and coded using validated and standardized processes consistent with national health data standards. As ICES data are anonymized and routinely collected, this study is legally exempt from research ethics review based on provincial health privacy legislation. We will use unique, encrypted patient identifiers to deterministically link several databases to re-construct each patient’s perioperative health system episode, including: the Discharge Abstract Database (DAD), which captures demographic and clinical (diagnoses, comorbidities, procedures, admission characteristics) information about all hospitalizations; the Registered Persons Database (RPDB), which captures all deaths and death dates for Ontarians; the Ontario Drug Database (ODB), which captures all prescription drug claims; the Ontario Health Insurance Claims Database (OHIP) which includes physician claims data for inpatient, outpatient and long-term care settings; the Continuing Care Reporting System (CCRS), which captures details of non-hospital institutional care; the Home Care Data (HCD) which contains home care assessments and service data; the Ontario Cancer Registry (OCR) which captures all tissue diagnoses of malignancy; and the Canadian Census which includes sociodemographic statistics. Reporting will follow relevant guidelines.18–20 Study Population We will derive two distinct cohorts of non-cardiac surgery patients. The first will include patients >65 years of age on the date of their first major, elective, non-cardiac surgery (Apr 2012-Mar 2018). These surgeries (gender neutral major orthopedic, vascular, and oncologic) will be identified using validated Canadian Classification of Intervention (CCI) codes.21 The second will include patients >65 years of age on the date of their first emergency general surgery procedure (Apr 2012-Mar 2018), which will be identified using CCI codes for a core set of EGS procedures that account for >80% of deaths and resource use in the United States.22,23 Sample Size As there is no clearly defined minimally important increase in predictive accuracy measures, our sample size considerations focused on ensuring estimated models were stable and would not be overfit. We estimated the minimum number of individuals that would be required for our models using the methods of Riley and colleagues and the related ‘pmsampsize’ package in R.24 For mortality, assuming a conservatively low R2 value of 0.1, mortality rates of 1% and 16 parameters in our model, we would require 2715 individuals. As a population-based study, we will include all eligible individuals, and previous experience with similar data suggest that we will have more than 100,000 elective participants and 50,000 emergency participants. Exposures Frailty has been defined in electronic perioperative data using at least 22 different instruments.25 However, only a minority align with consensus frailty definitions. Therefore, our study will operationalize frailty exposure in 4 distinct ways: 1) FI,26 2) HFRS,27 3) ACG,28 and 4) RAI.29 As none of the frailty instruments were derived (i.e., weighted) in the data under study, our analyses represent external validation. Covariates Baseline clinical and demographic characteristics will be collected including age, sex, type of surgery, and ASA score (assigned by the intraoperative anesthesiologist). To support sensitivity analyses we will also collect socioeconomic indicators, cancer diagnoses, preoperative resource use, all Elixhauser comorbidities,30 and preoperative receipt of home- or institution-based support services. Outcomes Outcomes have been selected based on their valid availability in electronic data and relevance to patients and healthcare systems. The primary outcome will be 30-day mortality. Secondary outcomes will include days alive at home (calculated as the number of days alive within 30 or 365 days of surgery minus time in acute care hospitals (index or readmission) or institutionalized),31,32 length of hospital stay (date of discharge minus date of surgery), health system costs (using validated costing algorithms incorporating direct and indirect costs),33 and non-home discharge (hospital discharge to a non-home location or death in hospital (which is a competing risk)). Analysis All data manipulation and analyses will be performed using SAS version 9.4 for Windows (SAS Institute, Cary NC). Descriptive statistics will be computed separately in each cohort to compare characteristics between people who did, or did not, die within 30 days of surgery. Differences will be quantified using absolute standardized differences, where a value >0.1 is considered to represent a substantive difference. Agreement between dichotomized representations of each instrument will be quantified using kappa statistics. To compare predictive accuracy of different frailty instruments, we require a modelling framework that addresses two key considerations. First, we need to identify whether each frailty instrument adds predictive accuracy above that provided by risk factors typically assessed before surgery. While ‘typical’ preoperative variables used for risk assessment will vary, our methods draw on those of the METS study, as well as previous comparisons of clinical frailty instruments.2,34 Specifically, our baseline risk model will include age (as a restricted cubic spline), sex (binary), ASA score (categorical) and procedural risk (categorical using each CCI code). These variables will be used to estimate the ‘typical’ accuracy with which outcomes can be predicted without frailty assessment. Second, we need to compare whether a given frailty instrument adds greater accuracy than comparator instruments. Therefore, we will add each frailty instrument (separately) to the baseline model to estimate the predictive accuracy of the baseline model plus each frailty instrument. For binary outcomes (death, non-home discharge), logistic regression will be used. The predictive accuracy measures will be: discrimination (c-statistic: whether a model assigns a greater predicted probability of outcome to people who did experience the outcome than those who did not); calibration (calibration plots and integrated calibration index (ICI): extent to which predicted risks match observed outcomes); explained variance (Nagelkerke R2: extent that the model accounts for observed outcome variation); event reclassification (continuous net reclassification index (NRI): the propo
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,016 | 0,049 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,003 |
| Bibliométrie | 0,003 | 0,005 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,002 | 0,002 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».