DIAGNOSTIC SENSITIVITY AND INTEROBSERVER AGREEMENT OF RADIOGRAPHY AND ULTRASONOGRAPHY FOR DETECTING TROCHLEAR RIDGE OSTEOCHONDROSIS LESIONS IN THE EQUINE STIFLE
Bibliographic record
Abstract
Osteochondrosis lesions commonly occur on the femoral trochlear ridges in horses and radiography and ultrasonography are routinely used to diagnose these lesions. However, poor correlation has been found between radiographic and arthroscopic findings of affected trochlear ridges. Interobserver agreement for ultrasonographic diagnoses and correlation between ultrasonographic and arthroscopic findings have not been previously described. Objectives of this study were to describe diagnostic sensitivity and interobserver agreement of radiography and ultrasonography for detecting and grading osteochondrosis lesions of the equine trochlear ridges, using arthroscopy as the reference standard. Twenty-two horses were sampled. Two observers independently recorded radiographic and ultrasonographic findings without knowledge of arthroscopic findings. Imaging findings were compared between observers and with arthroscopic findings. Agreement between observers was moderate to excellent (κ 0.48-0.86) for detecting lesions using radiography and good to excellent (κ 0.74-0.87) for grading lesions using radiography. Agreement between observers was good to excellent (κ 0.78-0.94) for detecting lesions using ultrasonography and very good to excellent (κ 0.86-0.93) for grading lesions using ultrasonography. Diagnostic sensitivity was 84-88% for radiography and 100% for ultrasonography. Diagnostic specificity was 89-100% for radiography and 60-82% for ultrasonography. Agreement between radiography and arthroscopy was good (κ 0.64-0.78). Agreement between ultrasonography and arthroscopy was very good to excellent (κ 0.81-0.87). Findings from this study support ultrasound as a preferred method for predicting presence and severity of osteochondrosis lesions involving the femoral trochlear ridges in horses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".