Development and Validation of a Novel Measure for the Direct Assessment of Empathy in Veterinary Students
Bibliographic record
Abstract
Empathy is a requisite clinical skill for health professionals and empathy scores have been positively associated with professionalism, clinical competency, confidence, well-being, and emotional intelligence. In order to improve empathy in the veterinary field, it is critical to measure the construct of empathy accurately. Most research has relied on self-reporting measures to assess empathy, while some studies have recently implemented the use of simulated client encounters in veterinary education. Building on this research, the aim of the current study was to develop and validate a novel quantitative assessment tool-the Empathy Clinical Evaluation Exercise (ECEX)-designed to measure empathy based on directly observable behaviors, using simulated clients. To evaluate empathy, evaluators used the ECEX to assess the performance of student clinicians in a simulated client encounter, which contained a pre-determined number of opportunities designed to elicit empathic responses from student clinicians. Statistical analysis suggests the test has a high degree of inter-rater reliability. In addition, there was moderate correlation between average empathy scores using ECEX and previously validated measures of empathy, compassion satisfaction, and burnout. Using these methods, we found the majority of students we studied had increased empathy scores at the completion of their primary care rotations. These results provide preliminary support for the use of the ECEX as a direct and quantitative tool for the assessment of empathy. Health professionals could use this novel empathy assessment tool to teach students, evaluate teaching strategies, and improve communication competencies in a wide variety of clinical settings. Our broad aim was to examine the utility of a direct and quantitative assessment tool for measuring empathy-the ECEX-in order to answer the following questions: (1) Does the tool have good inter-rater reliability? (2) Does the tool correlate with previously validated empathy measures? and (3) Does the tool correlate with similar constructs of compassion fatigue and burnout? Our secondary aim was to evaluate the change in empathy scores over the course of a 4-month (16-week) primary care rotation (pre- to -post).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.033 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".