Single molecule long-read real-time amplicon-based sequencing of <i>CYP2D6</i> : a proof-of-concept with hybrid haplotypes
Notice bibliographique
Résumé
Abstract CYP2D6 is a widely expressed human xenobiotic metabolizing enzyme, best known for its role in the hepatic phase I cytochrome P450 enzyme system, where it metabolizes ∼20% of medications. It is also expressed in other organs including the brain, where its potential role in physiology and mental health traits and disorders is under further investigation. Owing to the presence of homologous pseudogenes in the CYP2D locus and transposable repeat elements in the intergenic regions, the gene encoding the CYP2D6 enzyme, CYP2D6 , is one of the most hypervariable known human genes - with more than 165 core haplotypes. Haplotypes include structural variants, with a subtype of these known as hybrid haplotypes or fusion genes comprising part of CYP2D6 and part of its adjacent pseudogene, CYP2D7 . The fusion genes are particularly challenging to identify. High fidelity (HiFi) single molecule real-time (SMRT) long-read sequencing can cover whole CYP2D6 haplotypes in a single continuous sequence, and is therefore ideal for structural variant detection. In addition, it is highly accurate and suitable for novel haplotype identification, which is necessary as new CYP2D6 haplotypes are continuously being discovered, and many more likely remain to be identified in relatively understudied populations such as Indigenous Peoples. The aim of the present work was to develop an efficient and accurate HiFi SMRT amplicon-based method capable of detecting the full range of CYP2D6 haplotypes including fusion genes. We report proof-of-concept for 24 amplicons including three positive controls, aligned to fusion gene haplotypes, with prior cross-validation data. Amplicons with CYP2D7-D6 fusion genes, including positive controls, aligned to the *13 subhaplotypes predicted ( *13F , *13A2 ) with 100% accuracy, with the exception of one that aligned at 99.9%. Alignment of the *68 was 100% and above 99.9% to the CYP2D6*68 partial sequences EU5300606 and JF307779, respectively. The best alignments for the remaining CYP2D6-2D7 fusion genes were ≥99.7% (to 3 significant figures). Lower percentage alignment for CYP2D6-2D7 fusion genes may reflect imperfect PCR optimization and/or the possibility that we may have haplotypes not yet in public databases. Further work on these is in progress. Moreover, we have adapted this method for non-hybrid haplotypes. This technique could therefore suffice for the characterization of the full range of CYP2D6 haplotypes. The method that we have developed could be extended to other complex loci and to other species in a multiplexed high throughput assay.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».