CYP2D6 Genotype as a Marker for Benefit of Adjuvant Tamoxifen in Postmenopausal Women: Lessons Learned
Bibliographic record
Abstract
He who does not know history is doomed to repeat it . Since 2003, there has been considerable interest and a good degree of credence given to the role of cytochrome P450 2D6 (CYP2D6) genotyping to predict tamoxifen benefit in adjuvant therapy ( 1–5 ). This construct was based on the fact that a variety of CYP2D6 polymorphisms lead to reduced CYP2D6 enzyme activity and hence result in lower plasma concentrations of endoxifen, a clinically active metabolite of tamoxifen. Tamoxifen itself is known to have a relatively weak affinity for the estrogen receptor (ER) and to undergo extensive primary and secondary metabolism mainly by the CYP2D6 enzyme, forming clinically active metabolites, which include 4-hydroxytamoxifen and endoxifen (4-hydroxytamoxifen and N-desmethyltamoxifen). These latter two metabolites have 30- to 100-fold greater affinity for ER compared with tamoxifen, and endoxifen has been assumed to be the most clinically important metabolite based on its high level in the plasma and high affinity for ER. Patients homozygous for functional CYP2D6 alleles, or heterozygous or homozygous for loss of functional CYP2D6 alleles, have been phenotypically classified as extensive, intermediate, or poor metabolizers, respectively, reflecting plasma concentrations of endoxifen ( 1 , 2 , 6 ). It was hypothesized that tamoxifen would be less effective in breast cancer patients with poor and intermediate metabolizer phenotypes and that these patients would experience fewer or less severe tamoxifen-induced hot flushes ( 4 , 7 , 8 ). This theory that certain CYP2D6 genotypes and phenotypes were associated with lower endoxifen concentrations and worse breast cancer outcome was supported by laboratory ( 1 , 6 , 9 ), cohort ( 9–11 ), and case–control ( 12 ) studies, as well as the prototype retrospective analysis of prospectively collected samples from a North Central Cancer Treatment Group adjuvant breast cancer trial by Goetz et al. ( 4 , 13 ). A variety of subsequent articles both supported and refuted these observations ( 14 ). In this issue of the Journal, two articles by Regan et al. ( 15 ) and Rae et al. ( 16 ) appear to refute the utility of CYP2D6 genotyping for predicting either tamoxifen benefit in the adjuvant setting or producing hot flushes. In these two large randomized studies comparing tamoxifen to an aromatase inhibitor, relatively large subsets were analyzed. In the study by Rae et al. ( 16 ), UK-only patients from the Arimidex, Tamoxifen, Alone or in Combination (ATAC) trial were included, and in the study by Regan et al. ( 15 ), more than half of the patients in the Breast International Group (BIG) 1-98 Trial, for whom tumor tissues could be obtained were analyzed. Extracted DNA was used for genotyping and compared in multivariable analysis with endpoints of breast cancer recurrence ( 15 , 16 ) or occurrence of hot flushes ( 15 ). In neither study, was there a statistically significant association between the presence of poor or intermediate metabolizer phenotype and breast cancer outcome. In the study by Rae et al. ( 16 ), measurements of UDP-glucuronosyltransferase-2B7 (UGT2B7), whose gene product inactivates endoxifen, were also examined in relation to breast cancer outcome, and no association was observed. Although each of these studies has their weaknesses, the randomized design of the original trials, the examination of 4861 of 8010 patients randomized to receive tamoxifen or letrozole from BIG 1-98 ( 15 ), and 1203 of the 2145 patients randomly assigned to receive tamoxifen or anastrozole in the United Kingdom as part of the ATAC study (20% of the total trial) ( 16 ) provided large sample sizes, which were basically selected in an unbiased fashion. The blinded genotyping and the statistical analysis were also strong methodological features. The statistical power may still be insufficient to show a positive association between CYP2D6 activity and outcome in patients taking tamoxifen, but possibly further data from other large randomized controlled trials will become available. However, the fact that these two studies confirm each other suggests that this matter has likely been laid to rest. Why has such a good hypothesis gone wrong? First, it is clear from the discussion in the study by Regan et al. ( 15 ) that endoxifen may not be the important metabolite in this setting. It seems that one of the reasons for considerable efficacy of tamoxifen is that several of its active metabolites actually totally saturate ER ( 17 ). Furthermore, endoxifen has a different mechanism of action than 4-hydroxytamoxifen by targeting ER-alpha for degradation as opposed to stabilizing it and inhibiting estradiol-mediated overexpression of amphiregulin, a ligand of epidermal growth factor ( 18 ). Thus, 4-hydroxytamoxifen may, in fact, be the important metabolite. Furthermore, plasma concentrations of tamoxifen and a primary metabolite N-desmethyltamoxifen are higher than concentrations of either endoxifen or 4-hydroxytamoxifen. Tamoxifen metabolites N-desmethyltamoxifen, di-desmethyltamoxifen, and 4-hydroxytamoxifen have been estimated to nearly saturate ER with 99.94% occupancy ( 17 ). Interestingly, the Women's Healthy Eating and Living Study (WHEL), although findings showed no relationship between endoxifen levels and outcome in 1370 women, did observe a poorer outcome in the 20% of women with a plasma endoxifen concentration in the lowest quintile compared with higher quintiles, an observation that may illustrate a minimum critical concentration threshold ( 19 ). This may suggest that endoxifen plays a role at low concentrations, but in general, tamoxifen is being dosed at a level that is more than sufficient for even poor metabolizers to derive full benefit from this drug. During the last 9 years, a great industry of CYP2D6 measurement has arisen. The Food and Drug Administration (FDA) Clinical Pharmacology Subcommittee suggested that the tamoxifen label should be updated to include information about increased risk of breast cancer recurrence in CYP2D6 poor metabolizers taking tamoxifen. Although no consensus was reached regarding routine CYP2D6 genotype testing, many laboratories began testing for CYP2D6 allelic variants. Additional tests and charges were administered to patients, and therapies were presumably adjusted accordingly, all prematurely. What lessons can we learn? 1. We must avoid bias. The use of randomized trials prospectively designed and retrospectively analyzed is critical to assess the role of biomarkers such as CYP2D6. 2. Large confirmatory studies are required for decisions regarding the use of therapeutic agents. Such studies should be required for the adoption of biomarkers as well. 3. Validation of predictive biomarkers is a complicated process. The clinical importance of valid laboratory observations can be obscured by multiple factors. For example, when the baseline risk of recurrence is low, sometimes surgery alone is adequate without tamoxifen. Second, CYP2D6 phenotype would have little impact on tamoxifen benefit when the breast cancer is already de novo resistant to endocrine therapy. Finally, it is likely that there are many undefined drug interactions between tamoxifen and other medications that affect production of endoxifen and its metabolites, and ultimately treatment response. In the end, it is crucial to obtain data from randomized trials for clinical demonstration of associations between biomarkers and disease outcomes. To advance breast cancer therapy, laboratory observations that raise hypotheses must be at the very core of what we do; however, it is only after independent validation that they can begin guide clinical practice. None.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.014 | 0.013 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".