Automatic Detection of Meniscus Tears Using Backbone Convolutional Neural Networks on Knee <scp>MRI</scp>
Bibliographic record
Abstract
BACKGROUND: Timely diagnosis of meniscus injuries is key for preventing knee joint dysfunction and improving patient outcomes because it decreases morbidity and facilitates treatment planning. PURPOSE: To train and evaluate a deep learning model for automated detection of meniscus tears on knee magnetic resonance imaging (MRI). STUDY TYPE: Bicentric retrospective study. SUBJECTS: In total, 584 knee MRI studies, divided among training (n = 234), testing (n = 200), and external validation (n = 150) data sets, were used in this study. The public data set MRNet was used as a second external validation data set to evaluate the performance of the model. SEQUENCE: A 3 T, coronal, and sagittal images from T1-weighted proton density (PD) fast spin-echo (FSE) with fat saturation and T2-weighted FSE with fat saturation sequences. ASSESSMENT: The detection system for meniscus tear was based on the improved YOLOv4 model with Darknet-53 as the backbone. The performance of the model was also compared with that of three radiologists of varying levels of experience. The determination of the presence of a meniscus tear from surgery reports was used as the ground truth for the images. STATISTICAL TESTS: Sensitivity, specificity, prevalence, positive predictive value, negative predictive value, accuracy, and receiver operating characteristic curve were used to evaluate the performance of the detection model. Two-way analysis of variance, Wilcoxon signed-rank test, and Tukey's multiple tests were used to evaluate differences in performance between the model and radiologists. RESULTS: The overall accuracies for detecting meniscus tears using our model on the internal testing, internal validation, and external validation data sets were 95.4%, 95.8%, and 78.8%, respectively. One radiologist had significantly lower performance than our model in detecting meniscal tears (accuracy: 0.9025 ± 0.093 vs. 0.9580 ± 0.025). DATA CONCLUSION: The proposed model had high sensitivity, specificity, and accuracy for detecting meniscus tears on knee MRIs. EVIDENCE LEVEL: 3 TECHNICAL EFFICACY: Stage 2.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".