In silico validation of the Autoinflammatory Disease Damage Index
Bibliographic record
Abstract
INTRODUCTION: Autoinflammatory diseases can cause irreversible tissue damage due to systemic inflammation. Recently, the Autoinflammatory Disease Damage Index (ADDI) was developed. The ADDI is the first instrument to quantify damage in familial Mediterranean fever, cryopyrin-associated periodic syndromes, mevalonate kinase deficiency and tumour necrosis factor receptor-associated periodic syndrome. The aim of this study was to validate this tool for its intended use in a clinical/research setting. METHODS: The ADDI was scored on paper clinical cases by at least three physicians per case, independently of each other. Face and content validity were assessed by requesting comments on the ADDI. Reliability was tested by calculating the intraclass correlation coefficient (ICC) using an 'observer-nested-within-subject' design. Construct validity was determined by correlating the ADDI score to the Physician Global Assessment (PGA) of damage and disease activity. Redundancy of individual items was determined with Cronbach's alpha. RESULTS: The ADDI was validated on a total of 110 paper clinical cases by 37 experts in autoinflammatory diseases. This yielded an ICC of 0.84 (95% CI 0.78 to 0.89). The ADDI score correlated strongly with PGA-damage (r=0.92, 95% CI 0.88 to 0.95) and was not strongly influenced by disease activity (r=0.395, 95% CI 0.21 to 0.55). After comments from disease experts, some item definitions were refined. The interitem correlation in all different categories was lower than 0.7, indicating that there was no redundancy between individual damage items. CONCLUSION: The ADDI is a reliable and valid instrument to quantify damage in individual patients and can be used to compare disease outcomes in clinical studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".