First Trimester Point of Care Ultrasound: Imaging Features and Review Behaviors Associated With Diagnostic Accuracy
Bibliographic record
Abstract
OBJECTIVES: We aimed to identify the most diagnostically challenging features in first-trimester point-of-care ultrasound (FT-POCUS) images. We also sought to determine the physician image review behaviors associated with increased diagnostic accuracy. METHODS: We conducted a multicenter prospective cross-sectional study in a convenience sample of emergency physicians in the United States and Canada. The web-based intervention included 400 FT-POCUS cases acquired via the transabdominal or transvaginal approach. Participants reviewed FT-POCUS cases to identify pregnancy-related imaging findings. We captured clickstream-level data with each case encounter, including the correctness of a participant's response and physician image review behaviors. RESULTS: We enrolled 317 participants, who collectively generated 16,295 case interpretations. The most diagnostically challenging imaging findings included eccentrically located gestational sac and endometrial collection/heterogeneous uterine material (p < 0.001 for all comparisons). Participants who reported "definite" certainty, as opposed to "probable," demonstrated a significantly higher odds of getting the diagnosis of intrauterine pregnancy (IUP) present or absent correct (OR = 4.48; 95% CI 4.00, 5.01) and a lower odds of time spent reviewing cases (OR = 0.46; 95% CI 0.40, 0.51). Those who reviewed a higher proportion of available views per case were more likely to accurately identify a fetal heartbeat (OR = 1.51; 95% 1.34, 1.69), multiple IUPs (OR = 1.33; 95% CI 1.10, 1.61), and adnexal structures (OR = 1.11; 95% CI 1.04, 1.17), but less likely to correctly identify an IUP (OR = 0.93; 95% CI 0.88, 0.99) and endometrial fluid collection/heterogeneous uterine material (OR = 0.96; 95% CI 0.92, 0.99). CONCLUSIONS: Emergency physicians interpreting FT-POCUS images encountered specific diagnostic challenges that may increase risks to patient safety. We found that higher diagnostic confidence correlated with greater diagnostic accuracy and efficiency. Reviewing a larger proportion of available images improved diagnostic accuracy for some findings, but not for others.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".