Assessing Physician and Patient Agreement on Whether Patient Outcomes Captured in Clinical Progress Notes Reflect Treatment Success: Cross-Sectional Study
Bibliographic record
Abstract
BACKGROUND: It remains unclear if there is agreement between physicians and patients on the definition of treatment success following orthopedic treatment. Clinical progress notes are generated during each health care encounter and include information on current disease symptoms, rehabilitation progress, and treatment outcomes. OBJECTIVE: This study aims to assess if physicians and patients agree on whether patient outcomes captured in clinical progress notes reflect a successful treatment outcome following orthopedic care. METHODS: We performed a cross-sectional analysis of a subset of clinical notes for patients presenting to a Level-1 Trauma Center and Regional Health System for follow-up for an acute proximal humerus fracture (PHF). This study was part of a larger study of 1000 patients with PHF receiving initial treatment between 2019 and 2021. From the full dataset of 1000 physician-labeled notes, a stratified random sample of 25 notes from each outcome label group was identified for this study. A group of 2 patients then reviewed the sample of 100 clinical notes and labeled each note as reflecting treatment success or failure. Cohen κ statistics were used to assess the degree of agreement between physicians and patients on clinical note content. RESULTS: The average age of the patients in the sample was 67 (SD 13) years and 82% of the notes came from female patients. Patients were primarily White (91%) and had Medicare insurance coverage (65%). The note sample came from fracture-related encounters ranging from the second to the tenth encounter after the index PHF visit. There were no significant differences in patient or visit characteristics across concordant and discordant notes labeled by physicians and patients. Among agreement levels ranging from poor to perfect agreement, physician and patient evaluators exhibited only a fair level of agreement in what they deemed as treatment success based on a Cohen κ of 0.32 (95% CI 0.10-0.55; P=.01). Furthermore, interpatient and interphysician agreement also demonstrated relatively low levels of agreement. CONCLUSIONS: The findings suggest that physicians and patients demonstrated low levels of agreement when assessing whether a patient's clinical note reflected a successful outcome following treatment for a PHF. As low levels of agreement were also observed within physician and patient groups, it is clear the definition of success varied highly across both physicians and patients. Further research is needed to elucidate physician and patient perceptions of treatment success. As outcome measurement and demonstrating the value of orthopedic treatment remain important priorities, it is important to better define and reach a consensus on what treatment success means in orthopedic medicine.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".