Validation of Self-Report Pain Scales in Children
Bibliographic record
Abstract
BACKGROUND AND OBJECTIVES: The Faces Pain Scale-Revised (FPS-R) and Color Analog Scale (CAS) are self-report pain scales commonly used in children but insufficiently validated in the emergency department setting. Our objectives were to determine the psychometric properties (convergent validity, discriminative validity, responsivity, and reliability) of the FPS-R and CAS, and to determine whether degree of validity varied based on age, sex, and ethnicity. METHODS: We conducted a prospective, observational study of English- and Spanish-speaking children ages 4 to 17 years. Children with painful conditions indicated their pain severity on the FPS-R and CAS before and 30 minutes after analgesia. We assessed convergent validity (Pearson correlations, Bland-Altman method), discriminative validity (comparing pain scores in children with pain against those without pain), responsivity (comparing pain scores pre- and postanalgesia), and reliability (Pearson correlations, repeatability coefficient). RESULTS: Of 620 patients analyzed, mean age was 9.2 ± 3.8 years, 291(46.8%) children were girls, 341(55%) were Hispanic, and 313(50.5%) were in the younger age group (<8 years). Pearson correlation was 0.85, with higher correlation in older children and girls. Lower convergent validity was noted in children <7 years of age. All subgroups based on age, sex, and ethnicity demonstrated discriminative validity and responsivity for both scales. Reliability was acceptable for both the FPS-R and CAS. CONCLUSIONS: The FPS-R and CAS overall demonstrate strong psychometric properties in children ages 4 to 17 years, and between subgroups based on age, sex, and ethnicity. Convergent validity was questionable in children <7 years old.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".