The Disease Activity Score is not suitable as the sole criterion for initiation and evaluation of anti–tumor necrosis factor therapy in the clinic: Discordance between assessment measures and limitations in questionnaire use for regulatory purposes
Bibliographic record
Abstract
OBJECTIVE: The Disease Activity Score (DAS) is widely used in clinical trials. A DAS of 5.1 defines the level of severe rheumatoid arthritis (RA) and is the criterion for the initiation of anti-tumor necrosis factor therapy in the UK and The Netherlands. In North America, similar rules are sometimes imposed. However, it is not known how accurately the DAS characterizes RA activity. The present study was undertaken to determine the concordance between DAS scores and physicians' assessments of RA activity, to investigate factors relating to discrepancies, and to assess the suitability of using the DAS in individual patients. METHODS: Six hundred sixty-nine RA patients were assessed using the DAS and other clinical measures. A physician's global estimate of RA activity was performed using an 11-point predefined scale and a standard definition of disease activity. RESULTS: The DAS and physician global assessment had substantially different distributions of values. The level of agreement (Kendall's tau-a) between DAS scores and physician global assessments was 49% (95% confidence interval 45-53%), Lin's coefficient of concordance was 0.62, and the Bland-Altman 95% limits of agreement were -3.17 and 3.99. These results suggest poor-to-moderate concordance between the 2 measures of disease activity. CONCLUSION: The DAS and the physician's assessment of RA activity do not approach, value, and weight RA variables to the same extent, suggesting that RA activity is not evaluated similarly by North American physicians and with the DAS. The scales do not have acceptable levels of concordance. There is too much inherent variability in the DAS and other RA scales (e.g., the Health Assessment Questionnaire) to recommend them as sole determinants of RA activity for clinical or regulatory purposes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".