Predicting Treatment Outcomes in Patients With Psoriatic Arthritis or Axial Spondyloarthritis: An Artificial Intelligence–Driven Approach
Bibliographic record
Abstract
Objective To develop machine learning (ML) models to predict the probability at baseline of achieving low disease activity (LDA) and high health-related quality of life (HRQOL) in patients with psoriatic arthritis (PsA) or axial spondyloarthritis (axSpA) treated with secukinumab (SEC). Methods AQUILA is an ongoing multicenter, prospective, noninterventional study assessing the effectiveness and safety of SEC in patients with active PsA or axSpA in Germany. Data from 1961 participants were used to develop ML models for predicting treatment outcomes. We investigated baseline prediction of achieving LDA and high HRQOL at week 16 using binary ML algorithms, identifying main predictors for LDA and high HRQOL and their direction of influence. In addition, explainable artificial intelligence (XAI) estimated the importance and impact of each predictor based on how it affected the change in individual patient predictions. Results In PsA, the main LDA predictors were patient global assessment, physician global assessment, pretreatment with biologic disease-modifying antirheumatic drugs (bDMARDs), tender joint count (TJC), and age; high HRQOL predictors were PsA Impact of Disease, Beck Depression Inventory (BDI), height, TJC, and BMI (kg/m 2 ). In axSpA, the main LDA predictors were Bath Ankylosing Spondylitis Disease Activity Index (BASDAI), pretreatment with bDMARDs, C-reactive protein, Assessment of SpondyloArthritis international Society Health Index (ASAS HI), and height; high HRQOL predictors were ASAS HI, BDI, BMI, height, and age. Conclusion XAI provides significant value by enabling explanations of individual patient predictions and their visualizations. This modeling approach may help in the development of a clinical decision support system for patient management.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".