Establishing a Core Domain Set to Measure Rheumatoid Arthritis Flares: Report of the OMERACT 11 RA Flare Workshop
Bibliographic record
Abstract
OBJECTIVE: The OMERACT Rheumatoid Arthritis (RA) Flare Group (FG) is developing a data-driven, patient-inclusive, consensus-based RA flare definition for use in clinical trials, longterm observational studies, and clinical practice. At OMERACT 11, we sought endorsement of a proposed core domain set to measure RA flare. METHODS: Patient and healthcare professional (HCP) qualitative studies, focus groups, and literature review, followed by patient and HCP Delphi exercises including combined Delphi consensus at Outcome Measures in Rheumatology 10 (OMERACT 10), identified potential domains to measure flare. At OMERACT 11, breakout groups discussed key domains and instruments to measure them, and proposed a research agenda. Patients were active research partners in all focus groups and domain identification activities. Processes for domain selection and patient partner involvement were case studies for OMERACT Filter 2.0 methodology. RESULTS: A pre-meeting combined Delphi exercise for defining flare identified 9 domains as important (>70% consensus from patients or HCP). Four new patient-reported domains beyond those included in the RA disease activity core set were proposed for inclusion (fatigue, participation, stiffness, and self-management). The RA FG developed preliminary flare questions (PFQ) to measure domains. In combined plenary voting sessions, OMERACT 11 attendees endorsed the proposed RA core set to measure flare with ≥78% consensus and the addition of 3 additional domains to the research agenda for OMERACT 12. CONCLUSION: At OMERACT 11, a core domain set to measure RA flare was ratified and endorsed by attendees. Domain validation aligning with Filter 2.0 is ongoing in new randomized controlled clinical trials and longitudinal observational studies using existing and new instruments including a set of PFQ.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.172 | 0.081 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.003 | 0.012 |
| Research integrity | 0.004 | 0.008 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".