An Assessment of the Ability of Diplomates, Practitioners, and Students to Describe and Interpret Recordings of Heart Murmurs and Arrhythmia
Bibliographic record
Abstract
The ability of clinicians, ie, 10 veterinary students, 10 general practitioners, and 10 board certified internists, to describe and interpret common normal and abnormal heart sounds was assessed. Recordings of heart sounds from 7 horses with a variety of normal and abnormal rhythms, heart sounds, and murmurs were analyzed by digital sonography. The perception of the presence or absence of the heart sounds S1, S2, and S4 was similar for clinicians irrespective of their level of training and was in agreement with the sonographic interpretation on 89, 82, and 78% of occasions, respectively. However, practitioners were less likely to correctly describe the presence of S3. The heart rhythm was correctly described as being regular or irregular on 89% of occasions, and this outcome was not affected by level of training. Differentiation of the type of irregularity was less reliable. The perception of the intensity of a heart murmur was accurate and correlated with the grade assigned in the living horses, R2 = .68, and with sonographic measurements of the murmur's intensity, R2 = .69. Clinicians overestimated the duration of cardiac murmurs, particularly that of the loud systolic murmur. Only diplomates could reliably differentiate systolic from diastolic murmurs. The ability to diagnose the underlying cardiac problem was significantly affected by training; diplomates, practitioners, and undergraduates made the correct diagnosis on 53, 33, and 29% of occasions, respectively. The poor diagnostic ability of practitioners and the lack of improvement in diagnostic skill after the 2nd year of veterinary school emphasizes the need for better teaching of these skills. Digital sonograms that combine sound files with synchronous visual interpretations may be useful in this regard.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".