Determination of Endovaginal Ultrasound Proficiency and Learning Curve Among Emergency Medicine Trainees
Bibliographic record
Abstract
OBJECTIVES: Performing and interpreting endovaginal ultrasound is an important skill used during the evaluation of obstetric and gynecologic emergencies. This study aims to describe the level of proficiency and confidence achieved after performing 25 endovaginal examinations. METHODS: This is a prospective study at a single urban academic emergency department. Participants performed a minimum of 25 endovaginal ultrasounds under the supervision of a point-of-care ultrasound expert. Anatomical structures were identified by the expert under ultrasound prior to each session. Each examination was scored for agreement of findings between the participant and expert. The data were used to develop a performance curve identifying when proficiency was achieved, where experiential benefit diminished, and when participants felt confident. RESULTS: A total of 1117 endovaginal ultrasound examinations were performed by 50 participants. Agreement after 25 examinations was highest (>95%) for probe insertion and preparation, bladder and uterus identification, and directionality. Agreement was lowest for identification of the ovaries (76%). Experiential benefit plateaus occurred earliest (10 exams) for preparation and insertion followed by bladder identification and directionality. Surprisingly, ovarian experiential benefit plateaued at 16 exams. Participant confidence improved overall and was lowest for the identification of ovaries and abnormal pelvic anatomy. CONCLUSIONS: There is a significant learning curve when performing endovaginal ultrasound. Our data do not support the use of 25 examinations as a minimum standard for identification of the ovaries or abnormal ovarian pathology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.031 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".