From Genetic Association to Forensic Prediction: Computational Methods and Tools for Identifying Phenotypically Informative Single Nucleotide Polymorphisms
Bibliographic record
Abstract
Pigmentation genetics has become an important pillar in the field of forensic genomics for its application in DNA-based prediction of externally visible characteristics (EVCs). EVCs such as hair color, eye color, and skin color are complex traits that are influenced by several loci. When traditional short tandem repeat DNA profiling does not reveal any matches, pigmentation-associated loci can be informative of an individual's EVCs through a process known as forensic DNA phenotyping (FDP). Current FDP panels contain a combined set of over 40 polymorphisms that have been identified as being significantly associated with skin, hair, and eye color. A comprehensive understanding of the genetics underlying pigmentation traits is required to improve the precision and accuracy of FDP estimations. Presented herein is a summary of methods and tools for conducting a genome-wide association study (GWAS) to identify forensically relevant, phenotypically informative single nucleotide polymorphisms. The pipeline described focuses on post-genotyping (i.e., in silico) analyses, with emphasis on association analyses, post-association analyses, and first-pass functional annotation. Using eye color as an example, we demonstrate how the pipeline uses GWAS data to draw preliminary conclusions regarding the location and function of pigmentation-associated variants. The experiment specifically investigates eye color associated variants in individuals with a blue eye color background (rs12913832:GG genotype) in a Canadian dataset. While methodologies and tools available for GWAS and post-GWAS processing continue to evolve and advance, the presented approaches have been applied successfully in numerous association analyses among hundreds of thousands of individuals in a wide range of disciplines. As such, they may offer a road map for future genomics investigations of pigmentation traits as well as other EVCs, ultimately serving to improve statistical predictions of phenotypes in forensic settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".