Breast MRI in the evaluation of locally recurrent or new breast cancer in the postoperative patient: correlation of morphology and enhancement features with the BI-RADS category
Bibliographic record
Abstract
BACKGROUND: While breast magnetic resonance imaging (MRI) is a highly sensitive test for detecting breast carcinoma, its specificity is lower, and several methods have been described on how to optimize specificity. PURPOSE: To compare the specificity and sensitivity of the BI-RADS category with the Fischer score in breast MRI for diagnosing cancer in women previously treated for breast cancer. MATERIAL AND METHODS: Women referred for evaluation of possible local recurrence or new breast cancer underwent breast MRI examination. Morphologic and kinetic enhancement characteristics were evaluated. BI-RADS category and Fischer score were assigned for each enhancing lesion and compared using a chi-square test. Sensitivity, specificity,and positive predictive values for 27 morphologic and enhancement characteristics were calculated. Pathologic diagnosis was obtained in all patients with enhancing lesions who had ultrasound or mammographic correlation. In those without correlate, 6-, 12-, and 24-month follow-up breast MRIs were obtained. Interobserver kappa correlation was determined for each variable studied. RESULTS: 34 benign and 32 malignant lesions were identified in 26 of 30 patients. BIRADS category yielded a specificity of 77.1% and a sensitivity of 81.8%. Fischer score had a lower specificity and sensitivity (62.9% and 72.7%, respectively) (P<0.0001). Of the 27 variables studied, >100% enhancement was more sensitive than BI-RADS for malignant lesions. Specificity was highest for rim enhancement (97.1%), but sensitivity was low (24.2%). Interobserver kappa correlation was good for all 27 characteristics(k=0.84), and highest for BI-RADS assessment (k=0.91). CONCLUSION: BI-RADS category in breast MRI had the highest combination of specificity and sensitivity, and the highest interobserver correlation. Fischer score and other morphologic and enhancement features lack sensitivity or specificity and do not have high positive predictive values when analyzed as single independent variables.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".