Detection of Differences in Longitudinal Cartilage Thickness Loss Using a Deep‐Learning Automated Segmentation Algorithm: Data From the Foundation for the National Institutes of Health Biomarkers Study of the Osteoarthritis Initiative
Bibliographic record
Abstract
OBJECTIVE: To study the longitudinal performance of fully automated cartilage segmentation in knees with radiographic osteoarthritis (OA), we evaluated the sensitivity to change in progressor knees from the Foundation for the National Institutes of Health OA Biomarkers Consortium between the automated and previously reported manual expert segmentation, and we determined whether differences in progression rates between predefined cohorts can be detected by the fully automated approach. METHODS: The OA Initiative Biomarker Consortium was a nested case-control study. Progressor knees had both medial tibiofemoral radiographic joint space width loss (≥0.7 mm) and a persistent increase in Western Ontario and McMaster Universities Osteoarthritis Index pain scores (≥9 on a 0-100 scale) after 2 years from baseline (n = 194), whereas non-progressor knees did not have either of both (n = 200). Deep-learning automated algorithms trained on radiographic OA knees or knees of a healthy reference cohort (HRC) were used to automatically segment medial femorotibial compartment (MFTC) and lateral femorotibial cartilage on baseline and 2-year follow-up magnetic resonance imaging. Findings were compared with previously published manual expert segmentation. RESULTS: The mean ± SD MFTC cartilage loss in the progressor cohort was -181 ± 245 μm by manual segmentation (standardized response mean [SRM] -0.74), -144 ± 200 μm by the radiographic OA-based model (SRM -0.72), and -69 ± 231 μm by HRC-based model segmentation (SRM -0.30). Cohen's d for rates of progression between progressor versus the non-progressor cohort was -0.84 (P < 0.001) for manual, -0.68 (P < 0.001) for the automated radiographic OA model, and -0.14 (P = 0.18) for automated HRC model segmentation. CONCLUSION: A fully automated deep-learning segmentation approach not only displays similar sensitivity to change of longitudinal cartilage thickness loss in knee OA as did manual expert segmentation but also effectively differentiates longitudinal rates of loss of cartilage thickness between cohorts with different progression profiles.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".