Inter‐rater Agreement Between Self‐rated and Staff‐rated Clinical Frailty Scale Scores in Older Emergency Department Patients: A Prospective Observational Study
Bibliographic record
Abstract
Frailty is a state of vulnerability arising from multiple medical and psychosocial problems primarily affecting older people.1 It is associated with adverse outcomes including mortality, prolonged hospitalization, and functional dependence after discharge.2 Identifying frailty early may help concentrate resources on patients at high risk of iatrogenesis, functional decline, and death.3 Despite the growing number of emergency department (ED) visits by older people, frailty is relatively underexamined in the ED setting. The Canadian Study of Health and Aging Clinical Frailty Scale (CFS) is a commonly used frailty assessment tool. Derived from a 5-year cohort study involving over 10,000 older Canadians, the CFS employs a 9-point scale based on clinical judgment.4 A CFS score of 1 to 3 represents nonfrailty; 4 represents vulnerability to frailty; 5 and 6 represent mild and moderate frailty, respectively; and a score of 7 or higher represents severe frailty. It is quick to administer and predicts patient-important outcomes including mortality, adverse discharge (e.g., long-term care), and functional decline.5, 6 Dresden et al.7 found that MD-assigned CFS scores in the ED were associated with adverse discharge destination. Lewis et al.6 found that ED CFS scores were as accurate in predicting poor outcomes as more time-intensive instruments based on objective measures of physical frailty, but more practical and less disruptive. As Dresden et al. note, the level of agreement between self-rated and provider-rated frailty sheds light on whether providers and patients interpret frailty in the same way, helping to inform research and choice of screening instrument.8 Acceptable agreement between provider CFS scores would justify the use of a single instrument by various categories of ED providers, where registered nurses (RNs) are responsible for initial evaluation and management before assessment by a physician or advanced practice provider (e.g., nurse practitioner or physician assistant [APP]). Dresden et al.7 reported moderate agreement between patient and MD CFS scores. However, that study did not involve RNs or APPs, the latter of whom play a growing role in Canadian EDs. The objective of this study was to determine the level of agreement between self-rated and staff-rated (RN and MD/APP) CFS scores in older ED patients. This was a prospective, observational study of patients aged 75 or older presenting to an urban Canadian academic ED (annual census 65,000) between December 2018 and April 2019. At this institution, approximately 12.5% of the annual ED census are 75 years and older. Research personnel obtained written informed consent to participate. Patients requiring resuscitation, those with a Glasgow Coma Scale score of <14, and patients who could not understand English were excluded. The study received institutional research ethics board approval (18-0271-E). Patients received a copy of the CFS and were asked to circle the CFS score that best described themselves. Each patient's RN and MD/APP were given a copy of the same instrument and asked to circle the frailty rating that best described the patient. All raters were blinded to each other's scores. To estimate our desired sample size, we posited the three categories of CFS scores used in the literature (nonfrail [CFS ≤ 4], mildly to moderately frail [CFS 5 or 6], and severely frail [CFS ≥ 7]) 4. We assumed that each of these categories would have a frequency of 0.33. To detect a kappa of 0.80 with 80% power, we estimated that 144 patient encounters would be required.8 We added 10% to account for withdrawals and missing data, resulting in a final sample size of 160 patient encounters. For the primary analysis, we estimated inter-rater agreement between ordinal (i.e., 1 to 9) CFS scores using quadratic-weighted kappa statistics with 95% confidence intervals (CIs) for these dyads: patient–RN, patient–MD/APP, and RN–MD/APP. In a secondary analysis, CFS scores were dichotomized as frail (CFS ≥ 5) or nonfrail (CFS ≤ 4; a cut-point prevalent in the literature). We then estimated interrater agreement using Cohen's kappa statistics with 95% CI.4, 6 Over the 4-month study period, 159 of 160 patient encounters were included. One encounter was excluded due to a missing MD CFS score. Mean (±SD) age of the included patients was 82.9 (±6.0) and 86 (±54.1%) were female. Figure 1 displays distributions of CFS scores by rater category. A high proportion of patients (n = 31, 19.4%) rated themselves as “very fit” (CFS = 1), whereas fewer patients were rated as “very fit” by RNs (n = 17, 10.6%) and MD/APPs (n = 15, 9.4%). Overall, very few patients (n = 5, 3.2%), RNs (n = 7, 4.4%), or MD/APPs (n = 8; 5.0%) assigned ratings of “severely frail” (CFS ≥ 7). These patterns resemble those reported by Dresden et al.7 Inter-rater agreement between patient and RN scores was moderate (0.59, 95% CI = 0.46 to 0.71), as was the agreement between patient-rated frailty scores and those reported by MD/APPs (0.53, 95% CI = 0.42 to 0.64). Inter-rater agreement between RN and MD/APP scores was good (0.74, 95% CI = 0.67 to 0.81). When CFS scores were dichotomized as “nonfrail” (CFS < 5) or “frail” (CFS ≥ 5), patient–RN and patient–MD/APP inter-rater agreement was 0.51 (95% CI = 0.35 to 0.67) and 0.42 (95% CI = 0.30 to 0.59), respectively, and RN–MD/APP inter-rater agreement was 0.72 (95% CI = 0.60 to 0.83). This was a small study performed in a single academic ED and may not be generalizable to other centers. Patients in this study were not consecutively enrolled, which may have introduced risk of selection bias. For instance, research and clinical staff may have tended not to approach patients who seemed unwell, agitated, or somnolent (e.g., because of dementia and/or delirium). This may have yielded a sample that was less frail than the population of interest (all older ED patients). It is possible that some RNs and MD/APPs were biased by their impression of the patient's current health status (i.e., acute condition), rather than their baseline condition. We did not distinguish between different APP categories, nor assess staff's years of experience. Finally, we did not evaluate patient characteristics such as level of education, cognitive impairment, or functional status, which may represent potential confounders. Our permissive inclusion criteria (all stable ED patients over the age of 75 without serious neurologic impairment) and the fact that we did not specially train or select participating ED staff enhance our study's generalizability to other ED settings. However, by not including seriously ill or cognitively impaired older people, we may have excluded the frailest ED patients (e.g., with terminal illness or advanced dementia). Moreover, excluding non–English-speaking participants may affect generalizability of results to the actual ED population, in that culture may affect individuals' understanding and interpretation of frailty. The objective of this study was to determine the level of agreement between self-rated and ED staff-rated CFS scores in older ED patients. Patient-rated CFS scores showed moderate agreement with those of both RNs and MD/APPs. RNs and MD/APP CFS scores showed good inter-rater agreement. The difference between provider–provider and provider–patient agreement is a novel finding. It may be explained partially by the higher propensity of some patients, also reported by Dresden et al., to rate themselves as “very fit” (CFS = 1) or “fit” (CFS = 2).7 This hypothesis is supported by a subgroup analysis of our 63 participants with self-rated CFS scores of ≤2. In this group, there was no statistically significant agreement between patient–RN (0.06, 95% CI = −0.05 to 0.17) or patient–MD/APP (0.07, 95% CI = −0.20 to 0.16) dyads, whereas RN–MD/APP agreement was moderate (0.59, 95% CI = 0.39 to 0.80). Absent an objective criterion standard for ED frailty, we cannot conclude whether the phenomenon is one of underrating by patients, overrating by providers, or both. Contributing factors may include negative perceptions of the term “frailty” and the fact that it carries different meanings in the lay and medical contexts. Older adults' inaccurate understanding of their objective health status, particularly among frail individuals, may play a role.10 Conversely, ED providers may tend focus on patients' burden of comorbidities, while minimizing patient-important considerations such as functional status and mobility. Our findings support further study of the CFS in the ED as an alternative to time-intensive comprehensive geriatric assessments. Subsequent research should continue to explore both staff- and self-rated frailty and examine the relationship between CFS scores assigned at triage and a range of patient- and system-important outcomes, including resource utilization, length of hospital stay, ED revisits, functional changes, and survival. We recognize the study participants for their contributions as well as the staff of the Schwartz/Reisman Emergency Department at Mount Sinai Hospital, Toronto, Ontario, Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.007 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".