Accuracy, intermethod agreement, and inter-reviewer agreement for use of magnetic resonance imaging and myelography in small-breed dogs with naturally occurring first-time intervertebral disk extrusion
Bibliographic record
Abstract
OBJECTIVE: To determine accuracy, intermethod agreement, and inter-reviewer agreement for multisequence magnetic resonance imaging (MRI) and 2-view orthogonal myelography in small-breed dogs with first-time intervertebral disk (IVD) extrusion. DESIGN: Prospective evaluation study. ANIMALS: 24 dogs with thoracolumbar IVD extrusion. PROCEDURES: Each dog underwent MRI and myelography. Images obtained with each modality were independently evaluated and assigned standardized scores in a blinded manner by 3 reviewers. Results were compared with surgical findings. Inter-reviewer and intermethod agreements were assessed via κ statistics. Accuracy was assessed as the percentage of dogs for which ≥ 2 of 3 reviewers recorded findings identical to those determined surgically. RESULTS: Inter-reviewer agreement was substantial for site (κ = 0.70) and side of IVD extrusion (κ = 0.62) in T2-weighted magnetic resonance images and was substantial for site (κ = 0.72) and fair for side of extrusion (κ = 0.37) in myelographic images. Agreement for site between each modality and surgical findings was near perfect (κ = 0.94 and 0.88 for MRI and myelography, respectively). Intermethod agreement was substantial for site (κ = 0.71) and moderate for side of extrusion (κ = 0.40). Accuracy of MRI for site and side was 100% when results for T1-weighted, T2-weighted, and contrast-enhanced T1-weighted sequences were combined. Accuracy of myelography was 90.9% and 54.5% for site and side, respectively. CONCLUSIONS AND CLINICAL RELEVANCE: Agreement between imaging results and surgical findings for identification of IVD extrusion sites in small-breed dogs was similar for MRI and myelography. However, MRI appeared to be more accurate than myelography and allowed evaluation of extradural compressive mass composition.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".