Validation of prediction models of severe disease course and non-achievement of remission in juvenile idiopathic arthritis: part 1—results of the Canadian model in the Nordic cohort
Bibliographic record
Abstract
BACKGROUND: Models to predict disease course and long-term outcome based on clinical characteristics at disease onset may guide early treatment strategies in juvenile idiopathic arthritis (JIA). Before a prediction model can be recommended for use in clinical practice, it needs to be validated in a different cohort than the one used for building the model. The aim of the current study was to validate the predictive performance of the Canadian prediction model developed by Guzman et al. and the Nordic model derived from Rypdal et al. to predict severe disease course and non-achievement of remission in Nordic patients with JIA. METHODS: The Canadian and Nordic multivariable logistic regression models were evaluated in the Nordic JIA cohort for prediction of non-achievement of remission, and the data-driven outcome denoted severe disease course. A total of 440 patients in the Nordic cohort with a baseline visit and an 8-year visit were included. The Canadian prediction model was first externally validated exactly as published. Both the Nordic and Canadian models were subsequently evaluated with repeated fine-tuning of model coefficients in training sets and testing in disjoint validation sets. The predictive performances of the models were assessed with receiver operating characteristic curves and C-indices. A model with a C-index above 0.7 was considered useful for clinical prediction. RESULTS: The Canadian prediction model had excellent predictive ability and was comparable in performance to the Nordic model in predicting severe disease course in the Nordic JIA cohort. The Canadian model yielded a C-index of 0.85 (IQR 0.83-0.87) for prediction of severe disease course and a C-index of 0.66 (0.63-0.68) for prediction of non-achievement of remission when applied directly. The median C-indices after fine-tuning were 0.85 (0.80-0.89) and 0.69 (0.65-0.73), respectively. Internal validation of the Nordic model for prediction of severe disease course resulted in a median C-index of 0.90 (0.86-0.92). CONCLUSIONS: External validation of the Canadian model and internal validation of the Nordic model with severe disease course as outcome confirm their predictive abilities. Our findings suggest that predicting long-term remission is more challenging than predicting severe disease course.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.028 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".