MétaCan
Menu
Back to cohort
Record W4396581169 · doi:10.1186/s12711-024-00892-9

Estimating genomic relationships of metafounders across and within breeds using maximum likelihood, pseudo-expectation–maximization maximum likelihood and increase of relationships

2024· article· en· W4396581169 on OpenAlexaff
Andrés Legarra, Matias Bermann, Quanshun Mei, Ole Fredslund Christensen

Bibliographic record

VenueGenetics Selection Evolution · 2024
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetic and phenotypic traits in livestock
Canadian institutionsCanadian Wood Council
Fundersnot available
KeywordsMaximum likelihoodBiologyRestricted maximum likelihoodMaximizationExpectation–maximization algorithmGenomic selectionSelection (genetic algorithm)StatisticsEvolutionary biologyComputational biologyGeneticsMathematicsComputer scienceMachine learningGeneMathematical optimizationGenotype

Abstract

fetched live from OpenAlex

Abstract Background The theory of “metafounders” proposes a unified framework for relationships across base populations within breeds (e.g. unknown parent groups), and base populations across breeds (crosses) together with a sensible compatibility with genomic relationships. Considering metafounders might be advantageous in pedigree best linear unbiased prediction (BLUP) or single-step genomic BLUP. Existing methods to estimate relationships across metafounders $${\varvec{\Gamma}}$$ Γ are not well adapted to highly unbalanced data, genotyped individuals far from base populations, or many unknown parent groups (within breed per year of birth). Methods We derive likelihood methods to estimate $${\varvec{\Gamma}}$$ Γ . For a single metafounder, summary statistics of pedigree and genomic relationships allow deriving a cubic equation with the real root being the maximum likelihood (ML) estimate of $${\varvec{\Gamma}}$$ Γ . This equation is tested with Lacaune sheep data. For several metafounders, we split the first derivative of the complete likelihood in a term related to $${\varvec{\Gamma}}$$ Γ , and a second term related to Mendelian sampling variances. Approximating the first derivative by its first term results in a pseudo-EM algorithm that iteratively updates the estimate of $${\varvec{\Gamma}}$$ Γ by the corresponding block of the H-matrix. The method extends to complex situations with groups defined by year of birth, modelling the increase of $${\varvec{\Gamma}}$$ Γ using estimates of the rate of increase of inbreeding ( $$\Delta F$$ Δ F ), resulting in an expanded $${\varvec{\Gamma}}$$ Γ and in a pseudo-EM+ $$\Delta F$$ Δ F algorithm. We compare these methods with the generalized least squares (GLS) method using simulated data: complex crosses of two breeds in equal or unsymmetrical proportions; and in two breeds, with 10 groups per year of birth within breed. We simulate genotyping in all generations or in the last ones. Results For a single metafounder, the ML estimates of the Lacaune data corresponded to the maximum. For simulated data, when genotypes were spread across all generations, both GLS and pseudo-EM(+ $$\Delta F$$ Δ F ) methods were accurate. With genotypes only available in the most recent generations, the GLS method was biased, whereas the pseudo-EM(+ $$\Delta F$$ Δ F ) approach yielded more accurate and unbiased estimates. Conclusions We derived ML, pseudo-EM and pseudo-EM+ $$\Delta F$$ Δ F methods to estimate $${\varvec{\Gamma}}$$ Γ in many realistic settings. Estimates are accurate in real and simulated data and have a low computational cost.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.011
metaresearch head score (Gemma)0.047
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.011
Threshold uncertainty score0.056

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0110.047
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0030.003
Science and technology studies0.0010.001
Scholarly communication0.0020.002
Open science0.0020.002
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0030.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.022
GPT teacher head0.261
Teacher spread0.239 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations7
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueGenetics Selection EvolutionSame topicGenetic and phenotypic traits in livestockFrench-language works237,207