Interpretable Artificial Neural Network Models for Predicting Anti-Adalimumab Immune Complex and Serum Drug Level in Crohn’s Disease: A Proof-of-Concept Study
Bibliographic record
Abstract
Background: The development of anti-drug antibodies (ADAs) and resulting immune complexes are key mechanisms behind the secondary loss of response to adalimumab in Crohn’s disease (CD). Despite their clinical importance, routine immunogenicity assays are limited, underscoring the need for alternative predictive approaches. Objective: This study aimed to develop interpretable artificial neural network (ANN) models to predict immune complex formation and estimate serum adalimumab levels using routinely available clinical and laboratory data from CD patients. Methods: A prospective analysis was performed on 58 CD patients on maintenance adalimumab. Immune complexes and serum adalimumab were measured via ELISA and lateral flow assays. ANN and ensemble regression models were trained on demographic, clinical, and inflammatory data, with performance evaluated by five-fold cross-validation. Interpretability was enhanced using Garson’s algorithm and permutation importance. Results: The ANN-based classification model accurately predicted ADA immune complex formation, achieving an accuracy of 77.47% and an area under the curve (AUC) of 82.63%. The main predictive variables included extraintestinal manifestations, perianal disease, disease behavior, and age at diagnosis. For estimating serum adalimumab levels measured by ELISA, the model performed modestly (accuracy 59.89%, AUC 79.72%), incorporating factors such as Montreal classification, perianal disease, C-reactive protein, immunosuppressant use, and disease duration. Conclusions: Interpretable ANN models robustly predict anti-adalimumab immune complexes and, to a lesser extent, serum adalimumab, using clinically available data, including perianal disease. This proof-of-concept study is limited by the relatively small, single-center dataset (n = 58), which may affect model generalizability and increase the risk of overfitting. External validation in larger and multicenter cohorts is required before clinical implementation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".