Identification of the most important features of knee osteoarthritis structural progressors using machine learning methods
Bibliographic record
Abstract
OBJECTIVES: The aim was to identify the most important features of structural knee osteoarthritis (OA) progressors and classification using machine learning methods. METHODS: Participants, features and outcomes were from the Osteoarthritis Initiative. Features were from baseline (1107), including articular knee tissues (135) assessed by quantitative magnetic resonance imaging (MRI). OA progressors were ascertained by four outcomes: cartilage volume loss in medial plateau at 48 and 96 months (Prop_CV_48M, 96M), Kellgren-Lawrence (KL) grade ⩾ 2 and medial joint space narrowing (JSN) ⩾ 1 at 48 months. Six feature selection models were used to identify the common features in each outcome. Six classification methods were applied to measure the accuracy of the selected features in classifying the subjects into progressors and non-progressors. Classification of the best features was done using an automatic machine learning interface and the area under the curve (AUC). To prioritize the top five features, sparse partial least square (sPLS) method was used. RESULTS: For the classification of the best common features in each outcome, Multi-Layer Perceptron (MLP) achieved the highest AUC in Prop_CV_96M, KL and JSN (0.80, 0.88, 0.95), and Gradient Boosting Machine for Prop_CV_48M (0.70). sPLS showed the baseline top five features to predict knee OA progressors are the joint space width, mean cartilage thickness of the medial tibial plateau and sub-regions and JSN. CONCLUSION: = 1107) and MRI outcomes in addition to radiological outcomes, we identified the best features and classification methods for knee OA structural progressors. Data revealed baseline X-ray and MRI-based features could predict early OA knee progressors and that MLP is the best classification method.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".