Systematic Review of Discriminating Power of Outcome Measures Used in Clinical Trials of Fibromyalgia
Bibliographic record
Abstract
OBJECTIVE: Fibromyalgia (FM) comprises many symptoms and features. Consequently, studies on the condition have used a wide variety of outcome measures and assessment instruments. We investigated those outcome measures and instruments in association with the OMERACT (Outcome measures in Rheumatoid Arthritis Clinical Trials) FM Workshop initiative to define core outcome measures that should be used to assess FM. METHODS: A systematic literature review up to December 2007 was carried out using the keywords "fibromyalgia," "treatment" or "management," and "trial." Data were extracted on outcome measures and assessment instruments used and the pre and post mean and standard deviation to calculate effect sizes (ES). Further sensitivity analysis was carried out according to treatment type, blinding status, and study outcome. RESULTS: The outcome domains identified fell largely within those defined by OMERACT. Morning stiffness was frequently assessed and therefore has been included here. The number of assessment instruments used was wide-ranging, so sensitivity analysis was only carried out on the top 5 within each domain. ES ranged from 0.54 to 3.77 for the key OMERACT domains. Health-related quality of life (HRQOL) was the only exception that had no instrument with moderate sensitivity. Of the secondary domains, dyscognition was lacking any sensitive instrument, as were fatigue and anxiety in pharmacological trials. CONCLUSION: Each of the key OMERACT domains has an instrument that appears to be sensitive to change, with the exception of HRQOL, which requires further research. Dyscognition, fatigue, and anxiety would all benefit from more research into their assessment instruments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.111 | 0.399 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.023 | 0.018 |
| Bibliometrics | 0.014 | 0.012 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".