Multisource Feedback and Narrative Comments: Polarity, Specificity, Actionability, and CanMEDS Roles
Bibliographic record
Abstract
INTRODUCTION: Multisource feedback is a questionnaire-based assessment tool that provides physicians with data about workplace behaviors and may combine numeric and narrative (free-text) comments. Little attention has been paid to wording of requests for comments, potentially limiting its utility to support physician performance. This study tested the phrasing of two different sets of questions. METHODS: Two sets of questions were tested with family physicians, medical and surgical specialists, and their medical colleague and coworker respondents. One set asked respondents to identify one thing the participant physician does well and one thing the physician could target for action. Set 2 questions asked what does the physician do well and what might the physician do to enhance practice. Resulting free-text comments provided by respondents were coded for polarity (positive, neutral, or negative), specificity (precision and detail), actionability (ability to use the feedback to direct future activity), and CanMEDS roles (competencies) and analyzed descriptively. RESULTS: Data for 222 physicians (111 physicians per set) were analyzed. A total of 1824 comments (8.2/physician) were submitted, with more comments from coworkers than medical colleagues. Set 1 yielded more comments and were more likely to be positive, semi specific, and very actionable than set 2. However, set 2 generated more very specific comments. Comments covered all CanMEDS roles with more comments for collaborator and leader roles. DISCUSSION: The wording of questions inviting free-text responses influences the volume and nature of the comments provided. Individuals designing multisource feedback tools should carefully consider wording of items soliciting narrative responses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.027 | 0.179 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".