Outcome Measures in Acute Gout: A Systematic Literature Review
Bibliographic record
Abstract
OBJECTIVE: Five core domains have been endorsed by Outcome Measures in Rheumatology (OMERACT) for acute gout: pain, joint swelling, joint tenderness, patient global assessment, and activity limitation. We evaluated instruments for these domains according to the OMERACT filter: truth, feasibility, and discrimination. METHODS: A systematic search strategy for instruments used to measure the acute gout core domains was formulated. For each method, articles were assessed by 2 reviewers to summarize information according to the specific components of the OMERACT filter. RESULTS: Seventy-seven articles and abstracts met the inclusion criteria. Pain was most frequently reported (76 studies, 20 instruments). The pain instruments used most often were 100 mm visual analog scale (VAS) and 5-point Likert scale. Both methods have high feasibility, face and content validity, and within- and between-group discrimination. Four-point Likert scales assessing index joint swelling and tenderness have been used in numerous acute gout studies; these instruments are feasible, with high face and content validity, and show within- and between-group discrimination. Five-point Patient Global Assessment of Response to Treatment (PGART) scales are feasible and valid, and show within- and between-group discrimination. Measures of activity limitations were infrequently reported, and insufficient data were available to make definite assessments of the instruments for this domain. CONCLUSION: Many different instruments have been used to assess the acute gout core domains. Pain VAS and 5-point Likert scales, 4-point Likert scales of index joint swelling and tenderness and 5-point PGART instruments meet the criteria for the OMERACT filter.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.010 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".