Effects of Human Observer Presence on Pain Assessment Using Facial Expressions in Rabbits
Bibliographic record
Abstract
The goal of this study was to evaluate the effect of a human observer on Rabbit Grimace Scale (RbtGS) scores. The study scored video footage taken of 28 rabbits before and after orthopedic surgery, as follows: 24 h before surgery ( baseline ), 1 h after surgery ( pain ), 3 h after analgesia administration ( analgesia ), and 24 h after surgery ( 24h ) in the presence and absence of an observer. Videos were assessed twice in random order by 3 evaluators who were blind to the collection time and the presence or absence of an observer. Responses to pain and analgesia were evaluated by comparing the 4 time points using the Friedman test, followed by the Dunn test. The influence of the presence or absence of the observer at each time point was evaluated using the Wilcoxon test. Intra- and interrater reliabilities were estimated using the intraclass correlation coefficient. The scale was responsive to pain, as the scores increased after surgery and had decreased by 24 h after surgery. The presence of the observer reduced significantly the RbtGS scores (median and range) at pain (present, 0.75, 0 to 1.75; absent, 1, 0 to 2) and increased the scores at baseline (present, 0.2, 0 to 2; absent, 0, 0 to 2) and 24h after surgery (present, 0.33, 0 to 1.75; absent, 0.2, 0 to 1.5). The intrarater reliability was good (0.69) to very good (0.82) and interrater reliability was moderate (0.49) to good (0.67). Thus, the RbtGS appeared to detect pain when scored from video footage of rabbits before and after orthopedic surgery. In the presence of the observer, the pain scores were underestimated at the time considered to be associated with the greatest pain and overestimated at the times of little or no pain.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".