Deep learning-assisted distinguishing breast phyllodes tumours from fibroadenomas based on ultrasound images: a diagnostic study
Bibliographic record
Abstract
OBJECTIVES: To evaluate the performance of ultrasound-based deep learning (DL) models in distinguishing breast phyllodes tumours (PTs) from fibroadenomas (FAs) and their clinical utility in assisting radiologists with varying diagnostic experiences. METHODS: We retrospectively collected 1180 ultrasound images from 539 patients (247 PTs and 292 FAs). Five DL network models with different structures were trained and validated using nodule regions annotated by radiologists on breast ultrasound images. DL models were trained using the methods of transfer learning and 3-fold cross-validation. The model demonstrated the best evaluation index in the 3-fold cross-validation was selected for comparison with radiologists' diagnostic decisions. Two-round reader studies were conducted to investigate the value of DL model in assisting 6 radiologists with different levels of experience. RESULTS: Upon testing, Xception model demonstrated the best diagnostic performance (area under the receiver-operating characteristic curve: 0.87; 95% CI, 0.81-0.92), outperforming all radiologists (all P < .05). Additionally, the DL model enhanced the diagnostic performance of radiologists. Accuracy demonstrated improvements of 4%, 4%, and 3% for senior, intermediate, and junior radiologists, respectively. CONCLUSIONS: The DL models showed superior predictive abilities compared to experienced radiologists in distinguishing breast PTs from FAs. Utilizing the model led to improved efficiency and diagnostic performance for radiologists with different levels of experience (6-25 years of work). ADVANCES IN KNOWLEDGE: We developed and validated a DL model based on the largest available dataset to assist in diagnosing PTs. This model has the potential to allow radiologists to discriminate 2 types of breast tumours which are challenging to identify with precision and accuracy, and subsequently to make more informed decisions about surgical plans.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".