MétaCan
Menu
Back to cohort
Record W2524054538 · doi:10.1101/078600

Quality control analysis of the 1000 Genomes Project Omni2.5 genotypes

2016· preprint· en· W2524054538 on OpenAlexaffabout
Nicole M. Roslin, Weili Li, Andrew D. Paterson, Lisa J. Strug

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2016
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetic Mapping and Diversity in Plants and Animals
Canadian institutionsPublic Health OntarioUniversity of TorontoSickKids FoundationHospital for Sick Children
Fundersnot available
Keywords1000 Genomes ProjectGenomePopulationCitationInbreedingGenotypeGeneticsSingle-nucleotide polymorphismBiologyQuality (philosophy)DemographyLibrary scienceComputer scienceGeneSociology

Abstract

fetched live from OpenAlex

Citation For any use of the 1000 Genomes Project data, please use the citation as noted here: http://www.1000genomes.org/faq/how-do-i-cite-1000-genomes-project . To cite this report or the lists described here, please use the following: Roslin NM, Li W, Paterson AD, Strug LJ. Quality control analysis of the 1000 Genomes Project Omni2.5 genotypes (Abstract/Program #576/F). Presented at the 66 th Annual Meeting of The American Society of Human Genetics, October 18-22, 2016, Vancouver, Canada. Data Summary Chips : IlluminaHumanOmni2.5-4v1_B and Illumina HumanOmni25M-8v1-1_B Initial number of SNPs : 2 458 861 Initial number of samples : 2318 Number of SNPs passing QC : 1 989 184 (80.9%) Number of samples passing QC : 2318 (100%) Number of quasi-unrelated samples with consistent ethnicity and well inferred sex : 1736 Abstract The 1000 Genomes Project genotype 2318 individuals (48.1% male) from 19 populations in 5 continental groups on the Illumina Omni2.5 platform. The data are publicly available, and will prove a valuable resource to obtain ethnic-specific allele frequencies, as well as exploring population histories through principal components analysis (PCA), estimation of inbreeding coefficients, and admixture analysis. As in any study, the data should be cleaned prior to analysis, to remove individuals or markers of questionable quality. Furthermore, a thorough understanding of the relationships between individuals must be established. Here we report our findings after comprehensive examination of the data for quality control. The basic quality of the genotypes was assessed using standard procedures. KING version 1.4 was used to confirm the relationships in the provided pedigrees, and also to detect undeclared relationships. PCA was used to examine the similarities and differences between individuals among and between population groups. In general, the data was found to be of high quality. No samples were removed due to low call rate (<97%) or excess heterozygosity. Sex chromosome genotypes showed two individuals with discrepancies between reported and inferred sex, and were unable to determine sex in an additional 20 individuals; the sex for these was changed to unknown. Relationship checking found discrepancies between first-degree relationships in the provided pedigrees and the genotypes in 9 families, including one instance where a reported parent/child pair was unrelated, two instances where full sibs were unrelated, and one set of three individuals who formed a newly defined trio. A set of 1756 individuals who were inferred to be more distant than 3 rd degree relatives was extracted and used in PCA. These individuals clustered in a pattern that is consistent with other published reports of global populations. We identified 4 individuals whose genotypes clustered more closely with a different geographic region than the one in the provided data. Although the genotype data is of high quality, errors exist in the publicly available dataset that require attention prior to using the genotypes. PLINK-format files including SNPs with good quality metrics and revised pedigree structures is available at http://tcag.ca . Files with distantly related or unrelated individuals, with sex inference consistent with provided gender, and with PCA consistent with continental group are also available.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.558
Threshold uncertainty score0.975

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.015
GPT teacher head0.236
Teacher spread0.221 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations25
Published2016
Admission routes2
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicGenetic Mapping and Diversity in Plants and AnimalsFrench-language works237,207