Development of the machine learning model that is highly validated and easily applicable to predict radiographic knee osteoarthritis progression
Bibliographic record
Abstract
Many models using the aid of artificial intelligence have been recently proposed to predict the progression of knee osteoarthritis. However, previous models have not been properly validated with an external data set or have reported poor predictive performances. Therefore, the purpose of this study was to design a machine learning model for knee osteoarthritis progression, focusing on high validation quality and clinical applicability. A retrospective analysis was conducted on prospectively collected data, using the Osteoarthritis Initiative data set (5966 knees) for model development and the Multicenter Osteoarthritis Study data set (3392 knees) for validation. The analysis aimed to predict Kellgren-Lawrence grade (KLG) progression over 4-5 years in knees with initial KLG of 0, 1, or 2. Possible predictors included demographics, comorbidities, history of meniscectomy, gait speed, Western Ontario and McMaster Universities Osteoarthritis Index (WOMAC) scores, and radiological findings. The Random Forest algorithm was employed for the predictive model development. Baseline KLG, contralateral knee osteoarthritis, lateral joint space narrowing (JSN) grade, BMI, medial JSN grade, and total WOMAC score were six features selected for the model in descending order of importance. Odds ratios of baseline KLG, contralateral knee osteoarthritis, and lateral JSN grade were 1.76, 2.59, and 4.74, respectively (all p < 0.001). The area-under-the-curve of the ROC curve in the validation set was 0.76 with an accuracy of 0.68 and an F1-score of 0.56. The progression of knee osteoarthritis in 4 ~ 5 years could be well-predicted using easily available variables. This simple and validated model may aid surgeons in knee osteoarthritis patient management.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.013 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".