Impact of Censored or Penalized Data in the Genetic Evaluation of Two Longevity Indicator Traits Using Random Regression Models in North American Angus Cattle
Bibliographic record
Abstract
This study aimed to evaluate the impact of different proportions (i.e., 20%, 40%, 60% and 80%) of censored (CEN) or penalized (PEN) data in the prediction of breeding values (EBVs), genetic parameters, and computational efficiency for two longevity indicators (i.e., traditional and functional longevity; TL and FL, respectively). In addition, three different criteria were proposed for PEN: (1) assuming that all cows with censored records were culled one year after their last reported calving; (2) assuming that cows with censored records older than nine years were culled one year after their last reported calving, while censored (missing) records were kept for cows younger than nine years; and (3) assuming that cows with censored records older than nine years were culled one year after their last reported calving, while cows younger than nine years were culled two years after their last reported calving. All analyses were performed using random regression models based on fourth order Legendre orthogonal polynomials. The proportion of commonly selected animals and EBV correlations were calculated between the complete dataset (i.e., without censored or penalized data; COM) and all simulated proportions of CEN or PEN. The computational efficiency was evaluated based on the total computing time taken by each scenario to complete 150,000 Bayesian iterations. In summary, increasing the CEN proportion significantly (p-value < 0.05 by paired t-tests) decreased the heritability estimates for both TL and FL. When compared to CEN, PEN tended to yield heritabilities closer to COM, especially for FL. Moreover, similar heritability patterns were observed for all three penalization criteria. High proportions of commonly selected animals and EBV correlations were found between COM and CEN with 20% censored data (for both TL and FL), and COM and all levels of PEN (for FL). The proportions of commonly selected animals and EBV correlations were lower for PEN than CEN for TL, which suggests that the criteria used for PEN are not adequate for TL. Analyses using COM and CEN took longer to finish than PEN analyses. In addition, increasing the amount of censored records also tended to increase the computational time. A high proportion (>20%) of censored data has a negative impact in the genetic evaluation of longevity. The penalization criteria proposed in this study are useful for genetic evaluations of FL, but they are not recommended when analyzing TL.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".