Liver iron overload assessment by <i>T</i> magnetic resonance imaging in pediatric patients: An accuracy and reproducibility study
Bibliographic record
Abstract
T magnetic resonance imaging (MRI) provides rapid quantification of liver iron content (LIC). The reciprocal of T is directly proportional to iron and has been calibrated against LIC. There has, however, been few independent validation of the T method in a clinical setting. In 100 MRI studies on 75 pediatric patients being investigated for liver iron overload, we assess the accuracy and reproducibility of T-measured LIC, using regulatory approved T2-based FerriScan® for reference measurements. Results from independent analyses by two observers demonstrated robust inter- and intra-observer agreement (intraclass correlation coefficient (ICC) = 0.99 and 1.0, respectively). T-measured and reference LIC were strongly correlated (r = 0.94, P < 0.0001), with a regression slope of 0.97 over the range 0–25 mg Fe/g. The T technique is shown to be accurate and reproducible for rapid, non-invasive LIC quantification. Iron is used in the production of hemoglobin and is usually stored in the liver and spleen. In several pathologies, iron accumulates as a result of frequent blood transfusions to treat anemia (e.g., thalassemia, sickle cell disease) or as a result of excess iron absorption (e.g., hereditary hemochromatosis). Without treatment, excess iron can result in damage to various body organs, particularly endocrine, liver, and heart. The goal of treatment is to prevent liver damage and heart failure. Monitoring body iron content is, therefore, critical in managing patients with iron overload. Assessment of iron levels is conventionally performed with a liver biopsy, as this is the organ where iron first accumulates. The measured LIC is thus used as a surrogate for total body iron stores. However, biopsy is invasive and limited by sampling variation, thus prompting the need for a non-invasive alternative. MRI has been studied for over two decades for non-invasive measurement of LIC [1, 2]. MRI relaxation times T2 and T shorten in the presence of iron, and both have been calibrated against LIC [3-5]. The accuracy of the T2 approach has been validated in numerous clinical studies [6-8], and it is currently offered as a regulatory approved service (FerriScan®), replacing biopsy procedures in many centers worldwide. There are fewer studies of the T approach for LIC quantification [9], and validation in a clinical setting has been scarce. T imaging offers a significantly shorter scan time compared to the T2 approach, and is ideal for all patients and particularly for children to eliminate the need for sedation. In this study, we evaluate the accuracy and reproducibility of T-measured LIC against reference T2-based measurements in children with iron overload. To our knowledge, this is the first large-scale clinical validation and comparison of the T approach against the current non-invasive gold-standard, Ferriscan®, for liver iron quantification. One-hundred T2 and T-MRI examinations were performed on 75 patients. One examination had an extremely high LIC that exceeded the detection range of Ferriscan®; this examination was excluded from the analysis. Children generally displayed a uniform distribution of iron throughout the liver, as demonstrated on the T maps of a patient with thalassemia major (Fig. 1). There is a marked reduction in LIC as a result of iron-chelationtherapy. Representative axial T maps of a liver iron-overloaded patient. A: A patient with thalassemia major presented with a heavily iron-overloaded liver, with T and reference LIC measurements of 26 and 30 mg Fe/g, respectively. B: After iron-chelation therapy for 15 months, the LIC returned to normal range, with T and reference LIC measurements of 4.1 and 2.7 mg Fe/g, respectively. Regions-of-interest drawn to outline the liver is shown in red. [Color figure can be viewed in the online issue, which is available at wileyonlinelibrary.com.] Reference T2-measured LIC averaged in all examinations was 9.1 ± 7.3 mg Fe/g (range, 0.8–40.9 mg Fe/g). The T method yielded a mean LIC of 8.6–8.8 mg Fe/g amongst the two observers (range, 1.3–30.9 mg Fe/g). Figure 2 illustrates the accuracy and reproducibility of T-measured LIC. There was a strong correlation with reference measurements (r = 0.94, P < 0.0001) and a linear regression slope of 0.97 over the range 0–25 mg Fe/g. Above this threshold, where iron loading is considered extremely high, a consistent underestimation exists. Both inter-observer (ICC = 0.99) and intra-observer (ICC = 1.0) agreements were robust. Accuracy and reproducibility of T MRI measurement of LIC in iron-overloaded pediatric patients. A: Plot of absolute LIC measured by the T method versus reference T2-based measurements shows a high correlation (r = 0.94, P < 0.0001) and a linear regression slope of 0.97 over the LIC range 0–25 mg Fe/g. B: A difference plot between T-measured and reference LIC shows accurate absolute measurements up to 25 mg Fe/g, beyond which a consistent underestimation occurs in heavily iron-overloaded patients. C: A difference plot between T-measured LIC made by independent observers shows small variations (ICC = 0.99). D: A difference plot between T-based LIC measurements made by the same observer on different slices in the same patient shows negligible variations (ICC = 1.0). The therapeutic target for iron-chelation therapy in iron-overloaded patients is a LIC below 7 mg Fe/g. The intensity of iron-chelator therapy is guided by how the measured LIC compares with this threshold. Generally, the higher the LIC, the more intense the therapy, but above a certain limit of around 20 mg Fe/g, the therapy regimen remains the same at maximal tolerable doses. Therefore, it is most important that LIC measurement is accurate in the low- to mid-range (<25 mg Fe/g) and able to indicate high LIC in the high-range (>25 mg Fe/g) even if not on a one-to-one scale. The accuracy of our results agree with other recent studies that have reported on the T method [9, 10]. These studies considered only LIC levels below 25 mg Fe/g, which are lower than ours but nonetheless demonstrate the value of the technique for iron quantification. In our study, where patients with extremely high iron overload were included, our T method correctly classified them in the heavy iron overload category despite consistent underestimation compared to reference measurements. The underestimation can be alleviated with the use of shorter echo times around the 1.0 msec range. Future work will address improved accuracy in the high LIC regime using very short echo time acquisitions. Our study provides the largest reported pediatric patient population to date with paired T- and T2-based measurements of LIC. The short scan time requirement of a T-MRI examination offers a distinct advantage to young children, who would otherwise require sedation in order to remain still for the duration of a long T2-MRI examination. Our results demonstrate that T-measured LIC is as reliable as T2-derived measurements that are the current non-invasive gold-standard. Furthermore, our T method is highly robust, showing excellent reproducibility between operators and between repeated assessment by the same operator. Our results corroborate previous findings [4,9,10] and provide further evidence for the use of T-MRI for rapid, accurate, and reproducible non-invasive LIC quantification in iron-overloaded patients. Seventy-five patients, aged 12.4 ± 5.0 (range, ages 5–23), with iron overload were enrolled in this Institutional Review Board-approved prospective study. The diverse population included patients with thalassemia major (38), thalassemia intermedia (3), sickle cell disease (23), Diamond–Blackfan anemia (3), Fanconi anemia (2), hereditary spherocytosis (1), autoimmue hemolytic anemia (1), congenital sideroblastic anemia (1), pyruvate kinase deficiency (1), Hodgkin lymphoma (1), and hereditary hemochromatosis (1). Of these patients, 70 received frequent red cell transfusions. T2 and T data were acquired on all patients on a clinical 1.5 Tesla Siemens scanner (Avanto). The T2 protocol employed a multi-slice spin-echo sequence (scan time = 30 min 20 sec): 11 axial slices, slice thickness (TH) = 5 mm, matrix = 192 × 256, repetition time (TR) = 2500 msec, number of excitations (NEX) = 1; repeated for five echo times (TE) = 6, 9, 12, 15, 18 ms. The T protocol employed a non-breath-held multi-slice multi-echo gradient-echo sequence (scan time = 3 min 22 sec): 11 axial slices, TH = 6 mm, matrix = 101 × 192, variable field-of-view (FOV) (default = 350 mm), flip angle = 60°, TR = 500 msec, NEX = 4; 11 equally spaced TE = (2.3–30) msec. Reference LIC calculations based on T2 data were provided by FerriScan®. T data were analyzed using in-house software (Matlab v.7.0) on a single axial slice chosen by the observer. Analysis was based on the optimal method described in Ref. 11 for achieving accuracy. R (1/T) was computed at every pixel location using a constant offset model (S = Soe−TE × R2* + C) (eq. 1). A region-of-interest (ROI) was manually drawn on the R map to encompass the entire liver, excluding obvious blood vessels and ducts. The LIC (mg/g) was computed from the ROI-median R value using the calibration curve [Fe]= 0.0254 R + 0.202 (eq. 2) (Ref. 4). Two independent observers performed analysis and prescribed ROIs with no knowledge of FerriScan's results. One observer repeated the analysis in each patient on a different axial slice to determine slice-dependent and intra-observer variation. Descriptive statistics were used to describe the sample. Inter-observer and intra-observer agreement of measurements was assessed in terms of the intraclass correlation coefficient (ICC). Pearson correlation was used to measure the agreement between T-measured and reference LIC. Linear regression was used to estimate the slope of the regression line fitted to the absolute LIC measured by the T method versus reference T2-based measurements. H.L.M.C. developed the T-MRI analysis for liver iron quantification, designed and conducted the study, analyzed and interpreted the data, and wrote the manuscript; S.H. analyzed the data; R.M. performed statistical analysis and wrote the manuscript; I.O. designed and conducted the study and wrote the manuscript. Hai-Ling Margaret Cheng* , Stephanie Holowka , Rahim Moineddin§, Isaac Odame¶ **, * Department of Medical Biophysics, Faculty of Medicine, University of Toronto, Toronto, Canada, Physiology & Experimental Medicine, The Research Institute, The Hospital for Sick Children, Toronto, Canada, Department of Diagnostic Imaging, The Hospital for Sick Children; Toronto, Canada, § Department of Family and Community Medicine, Faculty of Medicine, University of Toronto; Toronto, Canada, ¶ Division of Haematology/Oncology, The Hospital for Sick Children; Toronto, Canada, ** Department of Pediatrics, Faculty of Medicine, University of Toronto, Toronto, Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".