Lack of Congruence in the Ratings of Patients' Health Status by Patients and Their Physicians
Bibliographic record
Abstract
PURPOSE: The purpose of this study was to examine if physician assessments of their patients' health status after the medical encounter are comparable to their patients' self-assessment of their own health. METHODS: Consecutive patients with musculoskeletal diseases were recruited when they attended 1 of the rheumatology outpatient clinics selected for the study. Five physicians participated in the study, 4 based at an academic center and 1 in the community. Patients were interviewed after seeing the physician; they completed health status questionnaires (mHAQ and SF-12) and rated their pain, worry about disease, and overall health status on visual analog scales. Standard gamble techniques were used to obtain patient utilities in relation to their health status, "gambling" on the probability of obtaining perfect health from an intervention with varying risks of death. After the medical encounter, physicians were asked to rate their patients' health status with similar instruments and with standard gamble elicitation techniques, blinded to the patients' responses. RESULTS: A total of 105 patients participated in the study; 70% were female; mean age was 54+/-16 years; 64% had a connective tissue disease, most commonly rheumatoid arthritis; and the other diseases in this group included soft tissue rheumatism, osteoarthritis, or low back pain. Statistically significant differences were observed between patient and physician ratings for pain, overall health, and standard gamble. On average, physicians rated their patients' health status higher than the patients themselves and were less willing to gamble on the risk of death versus perfect health. Intraclass correlation coefficients (ICC) were low: 0.42 for pain, 0.11 for worry, 0.11 for overall health, and 0.04 for standard gamble utilities. Similar findings were observed when subgroup analysis was performed for individual physicians and for patients with connective tissue diseases. No specific patient characteristic consistently related to increased divergence in the ratings. CONCLUSIONS: These findings suggest that the communication between physicians and their patients at the time of the medical encounter needs to be enhanced. An understanding of their patients' health perceptions may assist physicians in suggesting appropriate interventions, taking into account their patients' benefit-risk preferences.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.073 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".