Observer Ratings of Shared Decision Making Do Not Match Patient Reports: An Observational Study in 5 Family Medicine Practices
Bibliographic record
Abstract
Background Measuring shared decision making (SDM) in clinical practice is important to improve the quality of health care. Measurement can be done by trained observers and by people participating in the clinical encounter, namely, patients. This study aimed to describe the correlations between patients’ and observers’ ratings of SDM using 2 validated and 2 nonvalidated SDM measures in clinical consultations. Methods In this cross-sectional study, we recruited 238 complete dyads of health professionals and patients in 5 university-affiliated family medicine clinics in Canada. Participants completed self-administered questionnaires before and after audio-recorded medical consultations. Observers rated the occurrence of SDM during medical consultations using both the validated OPTION-5 (the 5-item “observing patient involvement” score) and binary questions on risk communication and values clarification (RCVC-observer). Patients rated SDM using both the 9-item Shared Decision-Making Questionnaire (SDM-Q9) and binary questions on risk communication and values clarification (RCVC-patient). Results Agreement was low between observers’ and patients’ ratings of SDM using validated OPTION-5 and SDM-Q9, respectively (ρ = 0.07; P = 0.38). Observers’ ratings using RCVC-observer were correlated to patients’ ratings using either SDM-Q9 ( r pb = −0.16; P = 0.01) or RCVC-patients ( r pb = 0.24; P = 0.03). Observers’ OPTION-5 scores and patients’ ratings using RCVC-questions were moderately correlated ( r φ = 0.33; P = 0.04). Conclusion There was moderate to no alignment between observers’ and patients’ ratings of SDM using both validated and nonvalidated measures. This lack of strong correlation emphasizes that observer and patient perspectives are not interchangeable. When assessing the presence, absence, or extent of SDM, it is important to clearly state whose perspectives are reflected.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.087 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".