Identification of Frailty using EMR and Admin data: A complex issue
Bibliographic record
Abstract
IntroductionFrailty is a state of vulnerability to diverse stressors emphasizing the importance of identifying the frail to support them. The burden of frailty in Canada is steadily growing. Today, approximately 25% of people over age 65 and 50% past age 85 – over one million Canadians – are medically frail. Objectives and ApproachTo develop an administrative data definition of frailty to facilitate clinical and health system planning. We will validate the definition by linking the administrative data to electronic medical records (EMR) data. The EMR definition is based on a Machine Learning binarized frailty flag for patients with a Rockwood Clinical Frailty Score > 5 on physician chart audit. The sensitivity of the Machine Learning was disappointing: 28% (95% CI: 21% to 36%).specificity was: 94% (95% CI: 93% to 96%), positive predictive value: 53% (95% CI: 42% to 64%), negative predictive value: 86% (95% CI: 83% to 88%). ResultsThere was little overlap between the EMR and administrative data definitions using the same population. Of the 29,382 eligible administrative data community dwelling patients over 65 years old, with a linkable EMR record, 2398 (8.15%) were identified as frail using the administrative data definition, but only 16.1% of these were frail according to the EMR definition. Of the 2396 who were identified as frail in EMR data, only 375 (15.7%) were identified as frail using the administrative data definition. Conclusion/ImplicationsWe are not yet able to develop a reliable administrative data definition of frailty to identify community living individuals to support health service planning. The lack of agreement between the results obtained from EMR and administrative data definitions suggests that further refinement is necessary. Identification of frailty remains complex.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.078 | 0.194 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.008 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.006 | 0.004 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".