A Systematic Review of Faces Scales for the Self-report of Pain Intensity in Children
Bibliographic record
Abstract
CONTEXT: Numerous faces scales have been developed for the measurement of pain intensity in children. It remains unclear whether any one of the faces scales is better for a particular purpose with regard to validity, reliability, feasibility, and preference. OBJECTIVES: To summarize and systematically review faces pain scales most commonly used to obtain self-report of pain intensity in children for evaluation of reliability and validity and to compare the scales for preference and utility. METHODS: Five major electronic databases were systematically searched for studies that used a faces scale for the self-report measurement of pain intensity in children. Fourteen faces pain scales were identified, of which 4 have undergone extensive psychometric testing: Faces Pain Scale (FPS) (scored 0-6); Faces Pain Scale-Revised (FPS-R) (0-10); Oucher pain scale (0-10); and Wong-Baker Faces Pain Rating Scale (WBFPRS) (0-10). These 4 scales were included in the review. Studies were classified by using psychometric criteria, including construct validity, reliability, and responsiveness, that were established a priori. RESULTS: From a total of 276 articles retrieved, 182 were screened for psychometric evaluation, and 127 were included. All 4 faces pain scales were found to be adequately supported by psychometric data. When given a choice between faces scales, children preferred the WBFPRS. Confounding of pain intensity with affect caused by use of smiling and crying anchor faces is a disadvantage of the WBFPRS. CONCLUSIONS: For clinical use, we found no grounds to switch from 1 faces scale to another when 1 of the scales is in use. For research use, the FPS-R has been recommended on the basis of utility and psychometric features. Data are sparse for children below the age of 5 years, and future research should focus on simplified measures, instructions, and anchors for these younger children.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.056 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.008 | 0.007 |
| Bibliometrics | 0.019 | 0.018 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".