Systematic Review of Discriminating Power of Outcome Measures Used in Clinical Trials of Fibromyalgia
Bibliographic record
Abstract
OBJECTIVE: Fibromyalgia (FM) comprises many symptoms and features. Consequently, studies on the condition have used a wide variety of outcome measures and assessment instruments. We investigated those outcome measures and instruments in association with the OMERACT (Outcome measures in Rheumatoid Arthritis Clinical Trials) FM Workshop initiative to define core outcome measures that should be used to assess FM. METHODS: A systematic literature review up to December 2007 was carried out using the keywords "fibromyalgia," "treatment" or "management," and "trial." Data were extracted on outcome measures and assessment instruments used and the pre and post mean and standard deviation to calculate effect sizes (ES). Further sensitivity analysis was carried out according to treatment type, blinding status, and study outcome. RESULTS: The outcome domains identified fell largely within those defined by OMERACT. Morning stiffness was frequently assessed and therefore has been included here. The number of assessment instruments used was wide-ranging, so sensitivity analysis was only carried out on the top 5 within each domain. ES ranged from 0.54 to 3.77 for the key OMERACT domains. Health-related quality of life (HRQOL) was the only exception that had no instrument with moderate sensitivity. Of the secondary domains, dyscognition was lacking any sensitive instrument, as were fatigue and anxiety in pharmacological trials. CONCLUSION: Each of the key OMERACT domains has an instrument that appears to be sensitive to change, with the exception of HRQOL, which requires further research. Dyscognition, fatigue, and anxiety would all benefit from more research into their assessment instruments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.032 | 0.045 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.021 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".