Should quantitative assessment of rheumatoid arthritis include measures of joint damage and patient distress, in addition to measures of apparent inflammatory activity?
Bibliographic record
Abstract
Recent controversies concerning patient global assessment (PATGL) in rheumatoid arthritis (RA) remission criteria (1, 2) largely ignore issues that emerge when, as noted by Felson et al, “remission criteria designed and validated in clinical trials are applied to…help assess treatment ‘success’ in clinical practice…and could serve as a ‘treat-to-target’ goal” (1). All six RA core data measures and indices—including swollen joint count and tender joint count, as well as PATGL, Disease Activity Score in 28 (DAS 28) Joints, and other indices—are elevated significantly not only by inflammatory activity, but also by secondary osteoarthritis (3-5), depression (6), and/or fibromyalgia (FM) (7, 8), regardless of levels of inflammatory activity (see Table 1 of representative reports). This phenomenon has limited impact in clinical trials, in which PATGL is as efficient as any measure (including swollen and tender joint counts) to distinguish active from control treatments (9) in patient groups toward meeting regulatory requirements to market a new agent. However, only 5% to 30% of all patients with RA meet eligibility criteria for clinical trials (10, 11), whereas more than 50% of patients seen in routine care have clinically important osteoarthritis, FM, and/or depression (3, 5-8), which elevate all clinical RA measures, affecting a goal “low activity or remission” according to RA indices in individual patients (12). Furthermore, self-report measures of patients with RA are elevated in 35% to 60% of normal elderly people who do not report any arthritis (13). Quantitative, pragmatic measures are available to assess joint damage and/or patient distress in order to better interpret whether elevated RA measures and indices result from these problems, rather than from—or in addition to—inflammatory activity. Joint damage may be quantitated as deformed and/or limited motion on a 28-joint count, as described in the initial report (14). Patient distress may be quantitated by disease-specific questionnaires for FM (15) and depression (16), and/or by indices for FM and depression on a single multidimensional health assessment questionnaire (MDHAQ) which agree more than 80% with reference questionnaires (17, 18). A RheuMetric checklist includes pragmatic 0-10 physician estimates for global status, inflammation, damage, and distress (19, 20). Quantitative assessment of comorbid joint damage and patient distress may be informative even in clinical trials, eg, to explain in part why 30% to 40% of patients treated with powerful biological therapies do not meet American College of Rheumatology 20 (ACR 20) response criteria (21), a relatively low target. Quantitative measurement of joint damage and patient distress—in addition to inflammatory activity—in routine care, long-term databases, and even clinical trials may clarify RA management, outcomes, and possible new remission criteria. All authors were involved in drafting the article or revising it critically for important intellectual content, and all authors approved the final version to be published.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".