Improvement in thyroid ultrasound report quality with radiologists’ adherence to 2015 ATA or 2017 TIRADS: a population study
Bibliographic record
Abstract
Objectives: There has been slow adoption of thyroid ultrasound guidelines with adherence rates as low as 30% and no population-based studies investigating adherence to guideline-based malignancy risk assessment. We therefore evaluated the impact of adherence to the 2015 ATA guidelines or 2017 ACR-TIRADS guidelines on the quality of thyroid ultrasound reports in our healthcare region. Methods: We reviewed 899 thyroid ultrasound reports of patients who received fine-needle aspiration biopsy and were diagnosed with Bethesda III or IV nodules or thyroid cancer. Ultrasounds were reported by radiology group 1, group 2, or other groups, and were divided into pre-2018 (before guideline adherence) or 2018 onwards. Reports were given a utility score (0-6) based on how many relevant nodule characteristics were included. Results: Group 1 had a pre-2018 utility score of 3.62 and 39.4% classification reporting rate, improving to 5.77 and 97.0% among 2018-onwards reports. Group 2 had a pre-2018 score of 2.8 and reporting rate of 11.5%, improving to 5.58 and 93.3%. Other radiology groups had a pre-2018 score of 2.49 and reporting rate of 32.2%, improving to 3.28 and 61.8%. Groups 1 and 2 had significantly higher utility scores and reporting rates in their 2018-onward reports when compared to other groups' 2018-onward reports, pre-2018 group 1 reports, and pre-2018 group 2 reports. Conclusions: Dedicated adherence to published thyroid ultrasound reporting guidelines can lead to improvements in report quality. This will reduce diagnostic ambiguity and improve clinician's decision-making, leading to overall reductions in unnecessary FNA biopsy and diagnostic surgery.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".