Predicting Clinical Remission of Chronic Urticaria Using Random Survival Forests: Machine Learning Applied to Real-World Data
Bibliographic record
Abstract
INTRODUCTION: The time required to reach clinical remission varies in patients with chronic urticaria (CU). The objective of this study is to develop a predictive model using a machine learning methodology to predict time to clinical remission for patients with CU. METHODS: Adults with ≥ 2 ICD-9/10 relevant CU diagnosis codes/CU-related treatment > 6 weeks apart were identified in the Optum deidentified electronic health record dataset (January 2007 to June 2019). Clinical remission was defined as ≥ 12 months without CU diagnosis/CU-related treatment. A random survival forest was used to predict time from diagnosis to clinical remission for each patient based on clinical and demographic features available at diagnosis. Model performance was assessed using concordance, which indicates the degree of agreement between observed and predicted time to remission. To characterize clinically relevant groups, features were summarized among cohorts that were defined based on quartiles of predicted time to remission. RESULTS: Among 112,443 patients, 73.5% reached clinical remission, with a median of 336 days from diagnosis. From 1876 initial features, 176 were retained in the final model, which predicted a median of 318 days to remission. The model showed good performance with a concordance of 0.62. Patients with predicted longer time to remission tended to be older with delayed CU diagnosis, and have more comorbidities, more laboratory tests, higher body mass index, and polypharmacy during the 12-month period before the first CU diagnosis. CONCLUSIONS: Applying machine learning to real-world data enabled accurate prediction of time to clinical remission and identified multiple relevant demographic and clinical variables with predictive value. Ongoing work aims to further validate and integrate these findings into clinical applications for CU management.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".