Developing Clinical Artificial Intelligence for Obstetric Ultrasound to Improve Access in Underserved Regions: Protocol for a Computer-Assisted Low-Cost Point-of-Care UltraSound (CALOPUS) Study
Bibliographic record
Abstract
BACKGROUND: The World Health Organization recommends a package of pregnancy care that includes obstetric ultrasound scans. There are significant barriers to universal access to antenatal ultrasound, particularly because of the cost and need for maintenance of ultrasound equipment and a lack of trained personnel. As low-cost, handheld ultrasound devices have become widely available, the current roadblock is the global shortage of health care providers trained in obstetric scanning. OBJECTIVE: The aim of this study is to improve pregnancy and risk assessment for women in underserved regions. Therefore, we are undertaking the Computer-Assisted Low-Cost Point-of-Care UltraSound (CALOPUS) project, bringing together experts in machine learning and clinical obstetric ultrasound. METHODS: In this prospective study conducted in two clinical centers (United Kingdom and India), participating pregnant women were scanned and full-length ultrasounds were performed. Each woman underwent 2 consecutive ultrasound scans. The first was a series of simple, standardized ultrasound sweeps (the CALOPUS protocol), immediately followed by a routine, full clinical ultrasound examination that served as the comparator. We describe the development of a simple-to-use clinical protocol designed for nonexpert users to assess fetal viability, detect the presence of multiple pregnancies, evaluate placental location, assess amniotic fluid volume, determine fetal presentation, and perform basic fetal biometry. The CALOPUS protocol was designed using the smallest number of steps to minimize redundant information, while maximizing diagnostic information. Here, we describe how ultrasound videos and annotations are captured for machine learning. RESULTS: Over 5571 scans have been acquired, from which 1,541,751 label annotations have been performed. An adapted protocol, including a low pelvic brim sweep and a well-filled maternal bladder, improved visualization of the cervix from 28% to 91% and classification of placental location from 82% to 94%. Excellent levels of intra- and interannotator agreement are achievable following training and standardization. CONCLUSIONS: The CALOPUS study is a unique study that uses obstetric ultrasound videos and annotations from pregnancies dated from 11 weeks and followed up until birth using novel ultrasound and annotation protocols. The data from this study are being used to develop and test several different machine learning algorithms to address key clinical diagnostic questions pertaining to obstetric risk management. We also highlight some of the challenges and potential solutions to interdisciplinary multinational imaging collaboration. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): RR1-10.2196/37374.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".