Measurement properties of pain scoring instruments in farm animals: A systematic review using the COSMIN checklist
Bibliographic record
Abstract
This systematic review aimed to investigate the measurement properties of pain scoring instruments in farm animals. According to the PRISMA guidelines, a registered report protocol was previously published in this journal. Studies reporting the development and validation of acute and chronic pain scoring instruments based on behavioral and/or facial expressions of farm animals were searched. Data extraction and assessment were performed individually by two investigators using the Consensus-based Standards for the Selection of Health Measurement Instruments (COSMIN) guidelines. Nine categories were assessed: two for scale development (general design requirements and development, and content validity and comprehensibility) and seven for measurement properties (internal consistency, reliability, measurement error, criterion and construct validity, responsiveness and cross-cultural validity). The overall strength of evidence (high, moderate, low, or very low) of each instrument was scored based on methodological quality, number of studies and studies' findings. Twenty instruments for three species (bovine, ovine and swine) were included. There was considerable variability concerning their development and measurement properties. Three behavior-based instruments scored high for strength of evidence: UCAPS (Unesp-Botucatu Unidimensional Composite Pain Scale for assessing postoperative pain in cattle), USAPS (Unesp-Botucatu Sheep Acute Composite Pain Scale) and UPAPS (Unesp-Botucatu Pig Composite Acute Pain Scale). Four instruments scored moderate for strength of evidence: MPSS (Multidimensional Pain Scoring System for bovine), SPFES (Sheep Pain Facial Expression Scale), LGS (Lamb Grimace Scale) and PGS-B (Piglet Grimace Scale-B). Most instruments (n = 13) scored low or very low for final overall evidence. Construct validity was the most reported measurement property followed by criterion validity and reliability. Instruments with reported validation are urgently required for pain assessment of buffalos, goats, camelids and avian species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".