Prediction of bulk tank somatic cell count violations based on monthly individual cow somatic cell count data
Bibliographic record
Abstract
The regulatory limit in Canada for bulk tank somatic cell count (BTSCC) was recently lowered from 500,000 to 400,000 cells/mL. Herd indices based on changes in cow somatic cell count over 2 consecutive months (e.g., proportion of healthy or chronically infected cows, cows cured, and new intramammary infection rate) could be used as predictors for BTSCC violations. The objective of this study was to develop a predictive model for exceeding the limit of 400,000 cells/mL in the next month using these herd indices. Dairy Herd Improvement (DHI) data were used from 924 dairy herds in Québec, Canada. Test-day BTSCC was estimated by dividing the sum of all cows' DHI test-day somatic cell count times DHI test-day milk production by the total volume of milk produced by the herd on that test-day. In total, 986 of 8,681 (11.4%) estimated BTSCC exceeded 400,000 cells/mL. The final predictive model included 6 variables: mean herd somatic cell score at the current test-month, proportion of cows >500,000 cells/mL at the current test-month, proportion of healthy cows during lactation at the current test-month, proportion of chronically infected cows at the current test-month, average days in milk at the current test-month, and annual mean daily milk production. The optimized sensitivity and specificity of the model were 76 and 74%, respectively. The positive predictive value and negative predictive value were 25 and 95%, respectively. This low positive predictive value and high negative predictive value demonstrated that the model was less accurate at predicting herds that would violate the estimated BTSCC threshold but very accurate at identifying herds that would not. In addition, the area under the curve for the receiver operating characteristic curve was 0.82, suggesting that the model had excellent discrimination between test-months that did and did not exceed 400,000 cells/mL. An internal validation was completed using a bootstrapped resampling-based estimation method and confirmed that the final model provided a validated estimate of predictive accuracy. This model could be used to monitor and advise clients on impending risks of exceeding the BTSCC limit.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".