MétaCan
Menu
Back to cohort
Record W4283739205 · doi:10.1101/2022.06.26.495014

Genotyping, sequencing and analysis of 140,000 adults from the Mexico City Prospective Study

2022· preprint· en· W4283739205 on OpenAlexaff
Andrey Ziyatdinov, Jason Torres, Jesús Alegre-Díaz, Joshua Backman, Joelle Mbatchou, Michael Turner, Sheila M. Gaynor, Tyler Joseph, Yuxin Zou, Daren Liu, Rachel Wade, Jeffrey Staples, Razvan Panea, Alex Popov, Xiaodong Bai, Suganthi Balasubramanian, Lukas Habegger, Rouel Lanche, Alex Lopez, Evan K. Maxwell, Marcus B. Jones, Humberto Garcia‐Ortíz, Raúl Ramírez-Reyes, Rogelio Santacruz-Benítez, Abhishek Nag, Katherine R. Smith, Mark Reppell, Sebastian Zöllner, Eric Jorgenson, William Salerno, Slavé Petrovski, John D. Overton, Jeffrey G. Reid, Timothy A. Thornton, Gonçalo R. Abecasis, Jaime Berúmen, Lorena Orozco, Rory Collins, Aris Baras, Michael Hill, Jonathan Emberson, Jonathan Marchini, Pablo Kuri‐Morales, Roberto Tapia‐Conyer

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2022
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetic Associations and Epidemiology
Canadian institutionsDiscovery Centre
FundersMedical Research CouncilUniversity of OxfordCancer Research UKBritish Heart FoundationWellcome TrustNational Institute for Health and Care ResearchRegeneron PharmaceuticalsUniversidad Nacional Autónoma de MéxicoAstraZeneca
Keywords1000 Genomes ProjectGenotypingExome sequencingImputation (statistics)Allele frequencyBiologyPopulationAncestry-informative markerGeneticsIndigenousExomeGenotypeGeographyEvolutionary biologyDemographySingle-nucleotide polymorphismMissing dataGeneMutationEcology

Abstract

fetched live from OpenAlex

Abstract The Mexico City Prospective Study (MCPS) is a prospective cohort of over 150,000 adults recruited two decades ago from the urban districts of Coyoacán and Iztapalapa in Mexico City. We generated genotype and exome sequencing data for all individuals, and whole genome sequencing for 10,000 selected individuals. We uncovered high levels of relatedness and substantial heterogeneity in ancestry composition across individuals. Most sequenced individuals had admixed Native American, European and African ancestry, with extensive admixture from indigenous groups in Central, Southern and South Eastern Mexico. Native Mexican segments of the genome had lower levels of coding variation, but an excess of homozygous loss of function variants compared with segments of African and European origin. We estimated population specific allele frequencies at 142 million genomic variants, with an effective sample size of 91,856 for Native Mexico at exome variants, all available via a public browser. Using whole genome sequencing, we developed an imputation reference panel which outperforms existing panels at common variants in individuals with high proportions of Central, South and South Eastern Native Mexican ancestry. Our work illustrates the value of genetic studies in populations with diverse ancestry and provides foundational imputation and allele frequency resources for future genetic studies in Mexico and in the United States where the Hispanic/Latino population is predominantly of Mexican descent.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.034
Threshold uncertainty score0.068

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0010.000
Scholarly communication0.0010.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0030.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.014
GPT teacher head0.241
Teacher spread0.227 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations15
Published2022
Admission routes1
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicGenetic Associations and EpidemiologyFrench-language works237,207