Cross-validation of technologies for genotyping <i>CYP2D6</i> and <i>CYP2C19</i>
Bibliographic record
Abstract
Abstract Background CYP2D6 and CYP2C19 are cytochrome P450 enzymes involved in the metabolism of many medications from multiple therapeutic classes. Associations between patterns of variants (known as haplotypes) in the genes encoding them ( CYP2D6 and CYP2C19 ) and enzyme activities are well described. The genes in fact comprise 21% of biomarkers in drug labels. Despite this, genotyping is not common, partly attributable to its challenging nature ( CYP2D6 having >100 haplotypes, including those with sequence from an adjacent pseudogene, and gene duplications). We cross-validated different methodologies for identifying haplotypes in these genes against each other. Methods Ninety-two samples with a variety of CYP2D6 and CYP2C19 genotypes according to prior AmpliChip CYP450 and TaqMan CYP2C19*17 data were selected from the Genome-based therapeutic drugs for depression (GENDEP) study. Genotyping was performed with TaqMan copy number variant (CNV) and single nucleotide variant (SNV) analysis, the next generation sequencing-based Ion S5 AmpliSeq Pharmacogenomics Panel, PharmacoScan, long-range polymerase chain reaction (L-PCR) followed by amplicon analysis, and Agena for CYP2C19 . Variant pattern to haplotype translation was automated. Results The inter-platform concordance for CYP2C19 was high (up to 100% for available data). For CYP2D6 , the IonS5-PharmacoScan concordance was 94% for a range of variants tested apart from those with at least one extra copy of a CYP2D6 gene (occurring at a frequency of 3.8%, 33/853), or those with substantial sequence derived from pseudogene, known as hybrids (3%, 26/853). Conclusions Inter-platform concordance for CYP2C19 was high, and, moreover, the Ion S5 and PharmacoScan data were 100% concordant with that from a TaqMan CYP2C19*2 assay. We have also demonstrated feasibility of using an NGS platform for genotyping CYP2D6 and CYP2C19 , with automated data interpretation methodology. This points the way to a method of making CYP2D6 and CYP2C19 genotyping more readily accessible.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.026 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".