MétaCan
Menu
Back to cohort
Record W4364352504 · doi:10.1111/1556-4029.15251

Commentary on: Elkins <scp>KM</scp>, Garloff <scp>AT</scp>, Zeller <scp>CB</scp>. Additional predictions for forensic <scp>DNA</scp> phenotyping of externally visible characteristics using the <scp>ForenSeq</scp> and Imagen kits. J Forensic Sci. 2023;68(2):608–13. https://doi.org/10.1111/1556‐4029.15215

2023· letter· en· W4364352504 on OpenAlexaff
Frank R. Wendt

Bibliographic record

VenueJournal of Forensic Sciences · 2023
Typeletter
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicForensic and Genetic Research
Canadian institutionsPublic Health OntarioUniversity of Toronto
Fundersnot available
KeywordsPleiotropyComputational biologyGenotypingGeneticsBiologyComputer sciencePhenotypeGeneGenotype

Abstract

fetched live from OpenAlex

See Original Article here See Author’s Response here Editor, I read with interest the article recently published titled “Additional predictions for forensic DNA phenotyping of externally visible characteristics using the ForenSeq and Imagen kits” by Elkins et al. [1] describing the potential for off-target prediction of externally visible characteristics when genotyping samples using the ForenSeq and Imagen kits from Verogen (recently acquired by QIAGEN) [2, 3]. I agree with the authors’ motivation to investigate possible off-target effects of commercially available genotyping kits for forensic purposes [4]. This type of investigation is motivated by two factors about common complex traits: polygenicity and pleiotropy. Polygenicity describes the presence of tens to thousands of genetic variants conferring small, yet often independent, effects on a trait of interest [5, 6]. Pleiotropy describes a feature of polygenic architectures such that each genetic variant may associate with many similar and/or disparate traits [5, 6]. However, the literature review performed by the authors (i) grossly underestimates the prevalence of off-target pleiotropic events found in these commercial DNA typing kits, largely due to their definition of an externally visible characteristics and (ii) overstates the influence and utility of these events for making a prediction of one's phenotype. Below I describe the breadth of these oversights (Figure 1). Elkins et al. [1] investigated their target loci by “reviewing published Genome-Wide Association Study (GWAS) NGS data reports and a systematic and exhaustive screening of related journal articles.” However, their limited results from table 1 support a gross omission of observations from the literature. A quick search of the 15 single-nucleotide polymorphisms (SNPs) from table 1 in the phenome-wide association study (PheWAS) feature of the GWAS Atlas [7] shows 3282 SNP-trait associations at nominal p-value threshold (p < 0.05) and 204 associations surviving the conventional GWAS p-value threshold (p < 5 × 10−8; Table S1). The 204 genome-wide significant associations involving these 15 loci are unsurprisingly enriched for dermatological traits (37.2-fold enrichment, p = 1.25 × 10−60; estimated by hypergeometric tests; Table S2); however, they also show enrichment of other potentially externally visible characteristics: activities like use of sun protection and heavy do-it-yourself physical activity (domain enrichment = 2.46-fold enrichment, p = 0.001) and social interactions like being a part of a religious group (domain enrichment = 4.81, p = 0.008). Furthermore, these 15 loci are enriched for neoplasms (domain enrichment = 7.36-fold, p = 2.64 x 10−9) and metabolic traits like total cholesterol (domain enrichment = 1.24-fold, p = 0.007). I also investigated the per-locus enrichment across 27 domains in the GWAS Atlas (Table S3) [7]. It is unsurprising once again that nearly all of the tested loci are enriched for association with dermatological traits. For example, the SNP rs3827760 indeed influences dermatological traits (enrichment = 134.3-fold, p = 2.82 × 10−13); however, the cited trait associations by Elkins et al. [1] do not support a genome-wide significant finding and certainly is not reflective of the large number of associations for this SNP. The GWAS Atlas [7] shows seven genome-wide significant associations (p < 5 × 10−8) with this SNP involving curly versus strait hair, excessive hairiness, and male balding patterns. Notably, there also is an association between rs3827760 and alcohol consumption (p = 5.84 × 10−13). Elkins et al. [1] also report a short-listed selection of traits deemed outwardly visible. For rs12203592, they report three traits, yet the GWAS Atlas [7] shows 43 associations at the level of genome-wide significance—26 of which could be considered outwardly visible: cheese intake (p = 3.58 x 10−10), belonging to a religious group (p = 1.46 × 10−15), number of vehicles owned by members of the household (p = 3.55 x 10−9), etc. The second major shortcoming of Elkins et al. [1] is the language orienting their findings as “predictive.” While the authors are transparent about the fact that SNP associations are not causal, they proceed to demonstrate how phenotypic prediction may be performed using genetic information in MetaHuman. It is critical that authors investigating the genetic influence on outwardly/externally visible characteristics take clear and unbiased stances on the racial and ethnic biases present in this type of work. Large studies of diverse ancestries are overwhelmingly lacking in the GWAS literature such that most findings identify Eurocentric genetic factors that fail to generalize across populations. There are growing concerns about the promotion of equitable application of machine learning prediction in biomedical and applied genetics fields that cannot be understated [8]. As it stands, the findings from Elkins et al. [1] are mere associations and are by no means “predictive” of an outcome. In summary, I agree with Elkins et al. [1] motivation for performing their investigation into off-target effects of commercially available kits used in forensic DNA typing. Their approach applies a very narrow definition of what it means for a trait to be externally visible and, as such, their list of possible off-target details understates the true pleiotropic nature of each variant. Furthermore, I implore readers to take these results at face value as the “predictive” message of Elkins et al. [1] fails to clearly state the many nuances of translating genotype–phenotype associations to trait prediction. Table S1-S3 Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.008
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Science and technology studies, Research integrity
Consensus categoriesMeta-epidemiology (narrow), Science and technology studies, Research integrity
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.748
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0040.008
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0020.002
Science and technology studies0.0030.007
Scholarly communication0.0010.000
Open science0.0030.002
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.035
GPT teacher head0.293
Teacher spread0.258 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Forensic SciencesSame topicForensic and Genetic ResearchFrench-language works237,207