Reliability of the ROCK Osteochondritis Dissecans Knee Arthroscopy Classification System - Multi-center Validation Study
Bibliographic record
Abstract
Objectives: Very few predictors of healing have been identified to help guide OCD treatment decisions for patients, families and physicians. As this condition is not common, multi-center study groups will be necessary to determine optimal diagnostic and treatment strategies. For staging systems to be useful, there must be agreement among observers of each stage, and from different centers. Although arthroscopic staging systems exist for OCD, none have been tested for intra-observer and inter-observer reliability. Using an expert consensus method, the ROCK OCD study group developed an arthroscopy classification system for OCD of the knee. The purpose of this study was to determine the reliability of an OCD classification system in a multicenter study group. Methods: We developed a classification system for arthroscopic evaluation of OCD of the knee based on the experience of a 13 centers experienced in the care of OCD. The classification system produced 6 arthroscopic categories (Cue Ball, Shadow, Wrinkle in the Rug, Locked Door, Trap Door, and Crater). A training module including arthroscopic photos, iconic sketches, and representative videos was developed to describe each stage. A total of 30 representative arthroscopic videos were viewed by 10 orthopedic surgeons who had not participated in the video case selection and preparation. After 60 days, the 30 videos were reviewed a second time in a new, randomly selected order and classified. An inter-rater reliability assessment was performed using the intra-class correlation method. Results: The intra-class correlation coefficient was 0.92, indicating a very good to excellent reliability of this classification system amongst orthopedic surgeons within the ROCK group. Conclusion: The ROCK OCD Knee arthroscopy classification system demonstrated high reliability. Relatively rare conditions will require multi-center study groups to perform high quality outcome studies. This classification system will facilitate multi-center studies for OCD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.026 | 0.044 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".