Bibliographic record
Abstract
Sir, Determination of circulating anti-Müllerian hormone (AMH) has become a cornerstone of the infertility workup (Seifer et al., 2011). AMH has been reported to be strongly associated with age, antral follicle count (AFC) and FSH; elevated levels have been associated with polycystic ovarian syndrome (PCOS). It has also been reported to be an indicator of response to ovarian stimulation (Broer et al., 2013), a predictor of pregnancy outcome IVF/ICSI (Khader et al., 2013), and a marker for the onset of menopause (Dolleman et al., 2013). It has been suggested that AMH should be considered a replacement for the ‘gold standard biomarker’ AFC (Usta and Oral, 2012; Nelson, 2013) based upon the contention that AFC is not sufficiently reliable due to variable skills of operators performing the baseline scans. Implicit in all the proposed attributes and uses for AMH, therefore, is the necessity of reliability and reproducibility of AMH detection. Li et al. (2012) recently concluded that age-appropriate ranges should be derived by each laboratory and that AMH levels from different assays are not directly comparable. There have been several commercially available assays for AMH: the Beckman Coulter's Immunotech (IOT) assay and the Diagnostic Systems Lab assay. These have been gradually replaced by Beckman Coulter's Gen II assay which has been widely used internationally for almost 3 years (Nelson and La Marca, 2011; Wallace et al., 2011). Several ‘urgent field safety notices’ for this kit have been issued over the past 10 months. They included both a product recall and several contradictory revisions to correct original manufacturer's instructions. Different notices issued between November 2012 (http://www.imb.ie/images/uploaded/documents/fsn/FSNDec2012/V16335_FSN.pdf) and July 2013 (http://www.mhra.gov.uk/home/groups/fsn/documents/fieldsafetynotice/con297532.pdf) indicated that the original methodology resulted in potentially (but not definitely) either falsely high or low AMH values, apparently due to complement interference in the assay (Han et al., 2013). This situation naturally raises considerable concern. Should patients be recalled and retested if their AMH levels (measured in the 3 years prior to the product correction notices) fell outside expected age-related ranges? Were patients counseled or treated accordingly? Should those AMH values that, perhaps spuriously, fell within expected age-related ranges be accepted as accurate? Should the Gen II AMH body of work prior to July 2013 be re-analyzed entirely, knowing in hindsight that values generated by the assay and any conclusions drawn may have been unreliable? There is no current standard for AMH measurement although the UK NEQAS has been conducting an international pilot scheme for the past 2 years (Syme et al., 2013). Standardization is a necessary step for the determination of assay accuracy. However, it is concerning that the target levels in samples distributed to participating labs are established not against a known standard, but by using an all laboratory trimmed mean (ALTM). Labs using the Gen II AMH kit over the last 3 years, therefore, were unknowingly reporting potentially falsely high or low levels, with the possibility of erroneous values actually being incorporated into the calculation of expected values. Use of the ALTM to derive expected values renders standardization of AMH measurement both vulnerable to and unable to detect technical error, especially when laboratories are increasingly relying upon only one assay worldwide. AMH has rather quickly acquired a very significant clinical role in the assessment and treatment of infertility, perhaps before its reliability and utility have been sufficiently and accurately defined. The decline in AMH throughout women's reproductive years, until it is undetectable by menopause, has been consistently observed, regardless of the assay employed or the demographic studied. However, within each age group, regardless of obstetric history, there is a wide variation in AMH (La Marca et al., 2010, 2012; Seifer et al., 2011). Recently, there have been suggestions that measurement of AMH is less convincing as a superior biomarker for ovarian reserve and its reproducibility is under question (Loh and Maheshwari, 2011; Rustamov et al., 2012; Fitzgerald et al., 2013). Many groups have reported no correlation between AMH and live birth following IVF/ICSI, although they did find low age-appropriate AMH was associated with poor ovarian response (Tremellen and Kolo, 2010; Broer et al., 2013; Li et al., 2013). La Marca et al. (2012) reported that in women with a normal reproductive history, there was no association between AMH level and either miscarriage or pregnancy. In addition, there are reports of intra- and inter-cycle variation among individuals, and even circadian variations among patients with PCOS (Bungum et al., 2013; Hadlow et al., 2013). Conflicting findings as well as the recent technical difficulties indicate a need for caution when evaluating the evolving clinical significance of AMH. It is important for physicians to educate their patients regarding (i) the relative nature of AMH determination; (ii) the absence of a generalized cut-off that either precludes or predicts successful IVF/ICSI outcome; and (iii) the inadvisability of attributing diagnostic or prognostic significance to an absolute AMH value. While AMH has a role to play in the infertility workup, it would be premature to recommend it as a stand-alone measure of ovarian reserve. We believe it should continue to be interpreted in context with all other laboratory, radiologic and clinical findings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.005 | 0.004 |
| Scholarly communication | 0.004 | 0.005 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.070 | 0.052 |
| Insufficient payload (model declined to judge) | 0.006 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".