Patient-Reported Outcome Measures Used for Neck Disorders: An Overview of Systematic Reviews
Bibliographic record
Abstract
BACKGROUND: The evaluation of patient-reported outcome measures for the neck from multiple systematic reviews will provide a broader view of, and may identify potential conflicting or consistent results for, their psychometric properties. OBJECTIVES: The purpose of this study was to conduct an overview of systematic reviews and synthesize evidence to establish the current state of knowledge on psychometric properties of patient-reported outcome measures for patients with neck disorders. METHODS: In this overview of systematic reviews, an electronic search of 6 databases (MEDLINE, Embase, CINAHL, ILC, the Cochrane Central Register of Controlled Trials, and LILACS) was conducted to identify reviews that addressed at least one measurement property of outcome measures for people with neck pain. Only systematic reviews with patient-reported outcome measures were included in the analysis. Risk of bias was assessed with A MeaSurement Tool to Assess systematic Reviews (AMSTAR). Data on measurement properties were extracted from each systematic review. RESULTS: From 13 systematic reviews, 8 patient-reported outcome measures were evaluated in 2 or more reviews. Risk-of-bias scores ranged from moderate (5-7) to high (4 and lower). Findings on internal consistency, test-retest reliability, construct validity, responsiveness to change, and content and structural validity were synthesized for the Neck Disability Index (NDI) in 11 systematic reviews; the Northwick Park Neck Pain Questionnaire and Neck Pain and Disability scale (NPDS) in 6 systematic reviews; the Copenhagen Neck Functional Disability Scale in 5 systematic reviews; the Neck Bournemouth Questionnaire in 4 systematic reviews; the Core Neck Pain Questionnaire and Patient-Specific Functional Scale in 3 systematic reviews, and the Whiplash Disability Questionnaire in 2 systematic reviews. CONCLUSION: High-quality evidence was found of good to excellent internal consistency and moderate to excellent test-retest reliability for the NDI. Moderate-quality evidence was found of good to excellent internal consistency and good test-retest reliability for the Northwick Park Neck Pain Questionnaire. High-quality evidence was found of excellent test-retest reliability and good to strong construct validity with pain scales for the Copenhagen Neck Functional Disability Scale. Moderate-quality evidence was found of unclear to excellent internal consistency and moderate to strong concurrent associations with the NDI and global assessment of change for the Neck Pain and Disability scale. Moderate-quality evidence was found of excellent internal consistency for the Whiplash Disability Questionnaire and of high test-retest reliability for the Patient-Specific Functional Scale. J Orthop Sports Phys Ther 2018;48(10):775-788. Epub 22 Jun 2018. doi:10.2519/jospt.2018.8131.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.006 | 0.003 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".