Reliability of a Method to Evaluate Frailty Using Medical Records of Hospitalized Octogenarians
Bibliographic record
Abstract
To the Editor: Many community-based frailty scales exist, but valid instruments to classify hospitalized individuals according to baseline frailty status using data from medical records are lacking.1-7 In the acute care setting, the prevalence of frailty can be as high as 70%.8 Evaluating the baseline frailty status of hospitalized individuals from retrospective datasets is difficult, because physical performance tests recorded in the hospital chart may not reflect preadmission capabilities.1, 2, 4 This letter describes a method of reliably ascertaining baseline frailty status from the medical records of hospitalized octogenarians. Twenty files were randomly selected from a list of 356 individuals aged 80 and older with nonvalvular atrial fibrillation or flutter admitted to a medical or surgical ward of an academic hospital in Montreal, Canada, between January 1, 2013, and December 31, 2013. Information was recorded on age, sex, weight, diagnosis upon admission, living arrangement (home or long-term care), past medical history, medications, ability to perform instrumental activities of daily living (IADLs) and activities of daily living (ADLs), need for a caregiver, mobility (including falls), and cognitive status and used to construct a clinical vignette about each individual according to a theoretical framework of frailty2 (Table 1). The 20 clinical vignettes were distributed to five academic geriatricians using an electronic survey. Each geriatrician was asked to independently apply the nine-level Canadian Study of Health and Aging Clinical Frailty Scale (CFS) to rate the frailty status of each individual.2 The CFS relies on the clinician's assessment to attribute a frailty score ranging from 1 (near-perfect health) to 9 (terminal).2 The CFS has been validated against the 70-item Frailty Index (correlation coefficient = 0.80), with each one-category increment of the scale significantly increasing the medium-term risks of death and entry into an institution.1 Geriatricians were asked to indicate in a text box the variables that most influenced their frailty rating for each case. Cases without perfect concordance among the five reviewers were redistributed to the geriatricians with an invitation to change their scores based on the ratings and rationales of their colleagues. For each round of the Delphi process, interrater reliability on the frailty scores for the 20 clinical vignettes was calculated using a Cronbach alpha intraclass correlation coefficient (ICC) using all nine levels. The CFS was also dichotomized, with scores of 6 or less classified as fit, mild or moderately frail and scores of 7 or greater as severely frail, based on clinical consensus. Interrater reliability for the dichotomous ratings was calculated using the kappa multirater coefficient. All statistical analyses were conducted using SPSS version 21 (SPSS, Inc., Chicago, IL). Concordance on the frailty ratings was only moderate during the first round of the Delphi process using the nine-level scale (ICC = 0.593, 95% confidence interval (CI) = 0.399–0.778). Discrepancies occurred in 12 of 20 cases. The multirater kappa for the dichotomous rating was 0.546 (P < .001). After the second round, the cases were reevaluated, and the ICC increased to 0.859 (95% CI = 0.731–0.936). The reliability for dichotomously classifying individuals as frail versus nonfrail increased to 1 (P < .001). One reason for initial disagreement was missing data on baseline prehospital functional status. Chart data were often ambiguous about whether individuals were necessarily dependent in ADLs or IADLs (CFS score = 7) or simply did not perform these tasks for cultural or traditional reasons (CFS score ≥6). A history of falls also tended to yield overestimation of frailty scores by geriatricians on the first assessment. If individuals were noted to be completely independent, with only one or two accidental falls, the frailty score was readjusted upward to indicate more-robust status. These findings suggest that the CFS is reliable to use with medical chart data if the information is entered into a template, the prehospital functional status of the individuals is recorded appropriately, and potential sources of bias related to overestimating the importance of falling are acknowledged. Although a previous study using a screening software program was adapted from the 70-item Frailty Index to encode the electronic medical records of a retrospective cohort of community-dwelling individuals aged 60 and older, interrater reliability of the method was not assessed.9 Experience with the clinical vignette template and the CFS significantly improved interrater reliability ratings in the current study, so it is recommended that research teams interested in applying the CFS to retrospective chart data pilot the scale with 20 clinical vignettes and engage in discussions with seasoned geriatricians to determine potential sources of interrater variability before embarking on frailty research using medical records. The authors thank Dr. Fadi Massoud, Dr. Judith Latour, Dr. Isabelle Payot, Dr. Thien Tuong Minh Vu, and Dr. Marie-Jeanne Kergoat for conducting the geriatrician ratings for the interrater reliability testing. We also express gratitude to Martin Ladouceur, biostatistician, for his statistical assistance. Marie-Claude D. Lefebvre, Maude St-Onge, Maude Glazer-Cavanagh and Laurence Bell contributed equally as co-first authors as part of their Pharmacy Residency Master's Project. The research reported in this letter was supported by the Faculty of Pharmacy, Université de Montréal. Conflict of Interest: None. Author Contributions: All authors contributed to the study concept and design. Lefebvre, St-Onge, Glazer-Cavanagh, Bell: acquisition of subjects and data. Lefebvre, St-Onge, Glazer-Cavanagh, Bell, Tannenbaum: data analysis and interpretation. Nguyen, LeFebrevre, Tannenbaum: preparation of manuscript. Sponsor's Role: This study received a small research stipend from the Faculty of Pharmacy at the Université de Montréal to pay for statistical consulting services. The sponsor had no role in the design, methods, subject recruitment, data collection, analysis, and preparation of the paper. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.025 | 0.207 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".