Outcome measures in systemic lupus erythematosus: constructing a meaningful response index from existing clinical trial data
Bibliographic record
Abstract
The purpose of this project is to develop a systemic lupus erythematosus (SLE) response index as a standard outcome measure in future therapeutic trials. Currently, there is no widely validated method for defining response to therapy. Most SLE trials to date have failed to meet predesigned endpoints, leading to controversy over whether it is drug treatments or outcome measures that are unsuccessful in SLE. A similar controversy in rheumatoid arthritis (RA) years ago was resolved by examining data from placebo-controlled trials with drugs that were only modestly effective. Important clinical variables were selected, criteria for patient improvement determined, and an index was developed that distinguished treated patients from those getting placebo. This index (ACR 20/50/70) used in RA trials has led to approval of more than 20 drug therapies. Now that large-scale SLE clinical trial data exist, we propose to use the approach that was successful in RA. We will perform a post-hoc analysis of the raw data from the BLISS-52 and BLISS-76 trials investigating belimumab for SLE. The disease activity indices (SELENA SLEDAI and BILAG) will be deconstructed and individual clinical and laboratory parameters will be identified (for example, rash, complement). The variables that are present in the majority of patients, improve over time, and have face validity will be selected for this index. Both the physician global assessment and a patient-related measure of quality of life will be included. Study data will be split 50/50 into a training set and a validation set. Baseline values of variables will be compared with values at the end of the study to determine the degree of improvement or deterioration occurring in individual patients during the study. We will examine various threshold percent-improvement cutoff points across sets of variables, selecting those that produce the largest difference between placebo-treated and drug-treated patients while retaining an acceptably low proportion of improved placebo-treated patients. We will follow the methodology outlined by Harold Paulus in previous work for RA. The index will be tested by applying it to the remaining set of study subjects (validation set) used to derive the criterion. Performance measures will include discriminative ability, calibration and overall accuracy. This new composite index will be simple to use, based on real individual patient clinical trial data, and will include patient-reported outcome measures. The index should serve to prevent useful drugs from being discarded due to inadequate trial designs. Preliminary data will be presented.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.056 | 0.051 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".