Agreement between physician and patient performance status ratings in an outpatient setting.
Bibliographic record
Abstract
66 Background: Performance status (PS) scores are generally completed by physicians, although patient PS measures are increasingly available. The objective of this study was to assess agreement between physician-recorded Eastern Cooperative Oncology Group (ECOG) PS and a patient-rated measure of performance status based on the ECOG measure in an outpatient oncology palliative care clinic. Methods: Physicians recorded ECOG PS for every patient as part of routine clinical practice. Patients recorded their PS using a measure based on the ECOG, where 0 = normal with no limitations; 1 = not my normal self, but able to be up and about with fairly normal activities; 2 = not feeling up to most things, but in bed or chair less than half the day; 3 = able to do little activity and spend most of the day in bed or chair; and 4 = pretty much bedridden, rarely out of bed. Patients also completed the ESAS-CS measure. Medical and demographic data were abstracted from the patients’ charts. We examined correlation between physician and patient PS measures as well as factors associated with a difference in patient versus physician measures for 949 patients with a visit between August 1, 2013 and December 31, 2014. Results: Weighted Kappa statistics indicated moderate inter-rater agreement at 0.32 (95%CI: 0.28-0.36). On average, patients rated their ECOG higher by 0.31 (95%CI: 0.25-0.37, p < 0.0001); 40.7% of patients rated their PS worse, 41.9% the same, and 17.3% better than did physicians. Worst agreement was for ECOG 0 (5 of 25 patients rated 0 by physicians also rated 0 themselves, 20% agreement); the highest agreement was in those with physician-rated ECOG 3 (87/136 patients, 64% agreement). Older patients reported better PS than physicians (p = 0.002) while those with worse ESAS distress scores reported worse ECOG than physicians (p < .0001). Conclusions: Patients tend to rate their PS as worse than rated by physicians, particularly those with greater symptom burden. However, older patients tend to rate their PS as better. This could be due to age bias of physicians or due to older patients assessing their PS optimistically. Future research will assess the survival data of participants to assess correlation of PS assessments of each group with prognosis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.043 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".