Commentary on: Elkins <scp>KM</scp>, Garloff <scp>AT</scp>, Zeller <scp>CB</scp>. Additional predictions for forensic <scp>DNA</scp> phenotyping of externally visible characteristics using the <scp>ForenSeq</scp> and Imagen kits. J Forensic Sci. 2023;68(2):608–13. https://doi.org/10.1111/1556‐4029.15215
Notice bibliographique
Résumé
See Original Article here See Author’s Response here Editor, I read with interest the article recently published titled “Additional predictions for forensic DNA phenotyping of externally visible characteristics using the ForenSeq and Imagen kits” by Elkins et al. [1] describing the potential for off-target prediction of externally visible characteristics when genotyping samples using the ForenSeq and Imagen kits from Verogen (recently acquired by QIAGEN) [2, 3]. I agree with the authors’ motivation to investigate possible off-target effects of commercially available genotyping kits for forensic purposes [4]. This type of investigation is motivated by two factors about common complex traits: polygenicity and pleiotropy. Polygenicity describes the presence of tens to thousands of genetic variants conferring small, yet often independent, effects on a trait of interest [5, 6]. Pleiotropy describes a feature of polygenic architectures such that each genetic variant may associate with many similar and/or disparate traits [5, 6]. However, the literature review performed by the authors (i) grossly underestimates the prevalence of off-target pleiotropic events found in these commercial DNA typing kits, largely due to their definition of an externally visible characteristics and (ii) overstates the influence and utility of these events for making a prediction of one's phenotype. Below I describe the breadth of these oversights (Figure 1). Elkins et al. [1] investigated their target loci by “reviewing published Genome-Wide Association Study (GWAS) NGS data reports and a systematic and exhaustive screening of related journal articles.” However, their limited results from table 1 support a gross omission of observations from the literature. A quick search of the 15 single-nucleotide polymorphisms (SNPs) from table 1 in the phenome-wide association study (PheWAS) feature of the GWAS Atlas [7] shows 3282 SNP-trait associations at nominal p-value threshold (p < 0.05) and 204 associations surviving the conventional GWAS p-value threshold (p < 5 × 10−8; Table S1). The 204 genome-wide significant associations involving these 15 loci are unsurprisingly enriched for dermatological traits (37.2-fold enrichment, p = 1.25 × 10−60; estimated by hypergeometric tests; Table S2); however, they also show enrichment of other potentially externally visible characteristics: activities like use of sun protection and heavy do-it-yourself physical activity (domain enrichment = 2.46-fold enrichment, p = 0.001) and social interactions like being a part of a religious group (domain enrichment = 4.81, p = 0.008). Furthermore, these 15 loci are enriched for neoplasms (domain enrichment = 7.36-fold, p = 2.64 x 10−9) and metabolic traits like total cholesterol (domain enrichment = 1.24-fold, p = 0.007). I also investigated the per-locus enrichment across 27 domains in the GWAS Atlas (Table S3) [7]. It is unsurprising once again that nearly all of the tested loci are enriched for association with dermatological traits. For example, the SNP rs3827760 indeed influences dermatological traits (enrichment = 134.3-fold, p = 2.82 × 10−13); however, the cited trait associations by Elkins et al. [1] do not support a genome-wide significant finding and certainly is not reflective of the large number of associations for this SNP. The GWAS Atlas [7] shows seven genome-wide significant associations (p < 5 × 10−8) with this SNP involving curly versus strait hair, excessive hairiness, and male balding patterns. Notably, there also is an association between rs3827760 and alcohol consumption (p = 5.84 × 10−13). Elkins et al. [1] also report a short-listed selection of traits deemed outwardly visible. For rs12203592, they report three traits, yet the GWAS Atlas [7] shows 43 associations at the level of genome-wide significance—26 of which could be considered outwardly visible: cheese intake (p = 3.58 x 10−10), belonging to a religious group (p = 1.46 × 10−15), number of vehicles owned by members of the household (p = 3.55 x 10−9), etc. The second major shortcoming of Elkins et al. [1] is the language orienting their findings as “predictive.” While the authors are transparent about the fact that SNP associations are not causal, they proceed to demonstrate how phenotypic prediction may be performed using genetic information in MetaHuman. It is critical that authors investigating the genetic influence on outwardly/externally visible characteristics take clear and unbiased stances on the racial and ethnic biases present in this type of work. Large studies of diverse ancestries are overwhelmingly lacking in the GWAS literature such that most findings identify Eurocentric genetic factors that fail to generalize across populations. There are growing concerns about the promotion of equitable application of machine learning prediction in biomedical and applied genetics fields that cannot be understated [8]. As it stands, the findings from Elkins et al. [1] are mere associations and are by no means “predictive” of an outcome. In summary, I agree with Elkins et al. [1] motivation for performing their investigation into off-target effects of commercially available kits used in forensic DNA typing. Their approach applies a very narrow definition of what it means for a trait to be externally visible and, as such, their list of possible off-target details understates the true pleiotropic nature of each variant. Furthermore, I implore readers to take these results at face value as the “predictive” message of Elkins et al. [1] fails to clearly state the many nuances of translating genotype–phenotype associations to trait prediction. Table S1-S3 Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,008 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,003 | 0,007 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,003 | 0,002 |
| Intégrité de la recherche | 0,002 | 0,003 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».