Sonographers' self‐reported visualization of normal postmenopausal ovaries on transvaginal ultrasound is not reliable: results of expert review of archived images from UKCTOCS
Bibliographic record
Abstract
OBJECTIVE: In the UK Collaborative Trial of Ovarian Cancer Screening (UKCTOCS), self-reported visualization rate (VR) of the ovaries by the sonographer on annual transvaginal sonographic (TVS) examinations was a key quality control (QC) metric. The objective of this study was to assess self-reported VR using expert review of a random sample of archived images of TVS examinations from UKCTOCS, and then to develop software for measuring VR automatically. METHODS: A single expert reviewed images archived from 1000 TVS examinations selected randomly from 68 931 TVS scans performed in UKCTOCS between 2008 and 2011 with ovaries reported as 'seen and normal'. Software was developed to identify the exact images used by the sonographer to measure the ovaries. This was achieved by measuring caliper dimensions in the image and matching them to those recorded by the sonographer. A logistic regression classifier to determine visualization was trained and validated using ovarian dimensions and visualization data reported by the expert. RESULTS: The expert reviewer confirmed visualization of both ovaries (VR-Both) in 50.2% (502/1000) of the examinations. The software identified the measurement image in 534 exams, which were split 2:1:1 providing training, validation and test data. Classifier mean accuracy on validation data was 70.9% (95% CI, 70.0-71.8%). Analysis of test data (133 exams) provided a sensitivity of 90.5% (95% CI, 80.9-95.8%) and specificity of 47.5% (95% CI, 34.5-60.8%) in detecting expert confirmed visualization of both ovaries. CONCLUSIONS: Our results suggest that, in a significant proportion of TVS annual screens, the sonographers may have mistaken other structures for normal ovaries. It is uncertain whether or not this affected the sensitivity and stage at detection of ovarian cancer in the ultrasound arm of UKCTOCS, but we conclude that QC metrics based on self-reported visualization of normal ovaries are unreliable. The classifier shows some potential for addressing this problem, though further research is needed. © 2017 The Authors. Ultrasound in Obstetrics & Gynecology published by John Wiley & Sons Ltd on behalf of the International Society of Ultrasound in Obstetrics and Gynecology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.027 | 0.128 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".