Transferability across countries of equations developed using milk mid-infrared spectroscopy to estimate daily body condition score change in dairy cows
Bibliographic record
Abstract
Routine milk samples are commonly subjected to spectroscopic analysis within the mid-infrared (MIR) region of the electromagnetic spectrum to estimate macro-constituents of milk such as fat, protein, lactose, and urea content. These spectra, however, can also be used to predict other traits, such as daily BCS change (ΔBCS). The objective of the present study was to assess the transferability across countries of equations for predicting daily ΔBCS that were developed using milk MIR data collected in Ireland and in Canada. Body condition was scored on a scale from 1 (emaciated) to 5 (obese) in both countries. A total of 347,254 BCS records from 80,400 Canadian cows were available along with 73,193 BCS records from 6,572 Irish cows. Partial least squares regression and neural networks were separately used to predict daily ΔBCS. Two scenarios were studied: (1) using Canadian and Irish data combined as the calibration dataset to predict daily ΔBCS in Canada and in Ireland separately, and (2) Canadian and Irish data used separately to predict daily ΔBCS in each country separately. These prediction methods were applied to data with and without pretreatment (i.e., first derivative of the spectrum) as well as with and without standardizing daily ΔBCS across countries. For all the scenarios investigated, the correlation between actual and predicted daily ΔBCS when calibrated and validated (using cross-validation) in the same country ranged from 0.92 to 0.94, and from 0.85 to 0.87 for the Canadian and Irish datasets, respectively. When the data from Canada and Ireland were combined in the calibration process to predict daily ΔBCS, the correlations between actual and predicted ΔBCS were ≥0.90 and ≥0.80 for Canadian and Irish daily ΔBCS, respectively indicating no improvement in predictive ability. Predictive performance when calibrated using only Canadian data and validated using only Irish data was poor, and vice versa. Nonetheless, when developing equations for a country for which a limited database (i.e., 100 records) of gold standard and MIR data was available, predictive performance improved when the limited database was supplemented with the large dataset from the other country. In general, for some of the investigated scenarios, standardizing the daily ΔBCS data within country before undertaking the calibration improved prediction accuracy. Overall, the benefits of merging data from different countries, at least based on the trait (i.e., daily ΔBCS) and countries (i.e., Ireland and Canada) considered in the present study, were limited and, in some cases, counterproductive.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.031 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".