Intra-System Repeatability of S-Detect for Breast Ultrasound Classification on Identical Static Images: A Single-Center Retrospective Repeatability Study (Preprint)
Bibliographic record
Abstract
Background: Computer-aided diagnostic systems such as S-Detect (Samsung Medison) are increasingly integrated into breast ultrasound workflows. Notwithstanding extensive past evaluation of S-Detect's diagnostic accuracy, its intrasystem repeatability at the software level with identical static images, a fundamental prerequisite for clinical reliability, has not been systematically investigated. Objective: This study aimed to evaluate the intrasystem repeatability of the S-Detect computer-aided diagnostic system in classifying breast nodules in identical static ultrasound images. Methods: This retrospective, registered, blinded repeatability study analyzed 398 breast nodules from 261 women (mean age 43.10, SD 12.57 years) who underwent surgery between February 2019 and March 2020 at a single institution. Identical stored static ultrasound images, acquired by a single experienced sonographer on 1 Samsung RS80A ultrasound system, were each analyzed twice using the same S-Detect workstation: immediately after acquisition (S-Detect 1) and again at least 4 weeks later under blinded conditions with manual cursor repositioning (S-Detect 2). Repeatability was assessed using concordance rate and Cohen κ. The diagnostic performance of each run was compared against surgical histopathology. Results: Of 398 nodules, 156 (39.2%) were initially classified as possibly benign, and 242 (60.8%) were initially classified as possibly malignant. On repeat analysis, 37.4% (149/398) and 62.6% (249/398) of the nodules were classified as possibly benign and malignant, respectively. A total of 4.5% (7/156) of the nodules initially classified as benign were reclassified as malignant, whereas no malignant-to-benign changes occurred. The overall concordance rate was 98.2% (391/398), with a Cohen κ of 0.95 (95% CI 0.94-0.99; P<.001). Diagnostic performance remained stable across runs (area under the curve=0.913 vs 0.902; P=0.702). Conclusions: Under controlled conditions with identical static images, S-Detect showed high intrasystem repeatability, underscoring strong software-level consistency, although its translation to real-world clinical reproducibility requires further validation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | low |
| gpt | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | low |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.021 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".