Novel Arthroscopic Classification of Osteochondritis Dissecans of the Knee
Bibliographic record
Abstract
BACKGROUND: Several systems have been proposed for classifying osteochondritis dissecans (OCD) of the knee during surgical evaluation. No single classification includes mutually exclusive categories that capture all of the salient features of stability, chondral fissuring, and fragment detachment. Furthermore, no study has assessed the reliability of these classification systems. PURPOSE: To determine the intra- and interobserver reliability of a novel, comprehensive arthroscopic classification system with mutually exclusive OCD lesion types. STUDY DESIGN: Cohort study (diagnosis); Level of evidence, 3. METHODS: The Research in OsteoChondritis of the Knee (ROCK) study group developed a classification system for arthroscopic evaluation of OCD of the knee that includes 6 arthroscopic categories-3 immobile types and 3 mobile types. To optimize comprehensibility and applicability, each was developed with a memorable name, a brief description, a line diagram corresponding to the archetypal arthroscopic appearance, and an arthroscopic photograph depicting this archetype. Thirty representative arthroscopic videos were evaluated by 10 orthopaedic surgeon raters, who classified each lesion. After 4 weeks, the raters again classified the OCD lesions depicted in the 30 videos in a new, randomly selected order. Reliability was assessed via the intraclass correlation coefficient (ICC). RESULTS: The interobserver reliability of this novel arthroscopy classification was estimated by an ICC of 0.94 (95% CI, 0.91-0.97) for the first round and 0.95 (95% CI, 0.93-0.98) for the second round. According to the standards for the magnitude of the reliability coefficient of Altman, these ICCs indicate that interobserver reliability was very good. The intraobserver reliability was estimated by an ICC of 0.96 (95% CI, 0.95-0.97), which indicates that the intraobserver reliability was similarly very good. CONCLUSION: The ROCK OCD knee arthroscopy classification system demonstrated excellent intra- and interobserver reliability. In light of this reliability, this classification system may be used clinically and to facilitate future research, including multicenter studies for OCD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.013 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".