Utility and reliability of various on-farm welfare and health indicators for preweaning dairy calves
Bibliographic record
Abstract
The first objective of this study was to quantify the utility of welfare and health indicators for preweaning dairy calves based on the opinion of bovine veterinarians.The second objective was to assess these indicators' inter-rater and intra-rater reliability.A total of 37 veterinarians interested in the health and welfare of preweaning calves were initially identified in a previous study.An email invitation to participate in the utility assessment through an online questionnaire was sent, and 24 of them agreed to participate.Thirty-two dairy calf welfare indicators were evaluated, with each indicator assigned a utility value on a visual analog scale from 0 (no utility) to 10 (high utility).Indicators were categorized into low (≤3.4), average (3.5-6.9), or high (≥7.0)utility based on their median values.Each indicator utility's interquartile range (IQR) was stratified into 3 categories (low, average, and high) based on percentiles.In the second phase, 4 trained observers (3 veterinarians, including 2 PhD students in clinical science and 1 postdoctoral veterinarian working on calf welfare, and 1 veterinary student) were selected to assess reliability.The student was replaced by a professor (veterinarian) with expertise in calf health for the reliability assessment using pictures and videos, due to the student's involvement in selecting the material.Twenty-six indicators were included in the inter-rater reliability assessment and 21 in the intra-rater reliability assessment.Reliability was evaluated using both onfarm and online approaches.In the on-farm approach, 40 calves were assessed by the trained observers 3 times on the same day, whereas the online approach involved rating indicators based on pictures and videos of calves and their living environment.The intraclass correlation coef-ficient was used to assess reliability for quantitative indicators, with benchmarks of <0.5 (poor), 0.5-0.75(moderate), 0.75-0.9(good), and >0.9 (excellent).For qualitative indicators, Gwet's agreement coefficients AC 1 /AC 2 were employed, with benchmarks of <0.2 (poor), 0.21 to 0.40 (fair), 0.41 to 0.60 (moderate), 0.61 to 0.80 (good), and 0.81 to 1.00 (very good).The majority (30/32) of the indicators had a high median utility (≥7.0).The highest median utilities were observed for rectal temperature (10/10), dehydration (10/10), body condition (lean calf), and lesions (9.5/10).Indicators with high median utility and low IQR included rectal temperature (10/10, IQR = 1), dehydration (10/10, IQR = 1), navel discharge (9/10, IQR = 1.25), bedding wetness (calf area, 8/10, IQR = 1.25), bedding cleanliness in the calf area (8/10, IQR = 1.25), calf hygiene score (rear, 7.5/10, IQR = 1.25), wall cleanliness (calf area, 7/10, IQR = 1), and calf hygiene score (belly, 7/10, IQR = 1.25).The indicators classified with average utility and high IQR were those with the lowest median utility (i.e., umbilical hernia [6.5/10, IQR = 3] and avoidance [5/10, IQR = 2.5]).Most of the selected welfare and health indicators had good inter-and intra-rater agreements.Five indicators had moderate inter-rater reliability: hip height, length from the withers to the lumbosacral junction, swollen navel, avoidance, and dehydration (skin tent test), and only hip height had a moderate intra-rater agreement.No indicator had poor or fair agreement for the reliability assessment.This study highlights indicators with high utility but also emphasizes the importance of considering utility variability when assessing welfare at the herd level.Indicators with high reliability were identified, and for those with moderate reliability, better rater training, adjustments to the categories, or using other indicators are encouraged to improve reliability.It also represents an essential step for implementing these indicators in assessing calf welfare across multiple farms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.034 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".