Properties of Serial Ultrasound Clinical Diagnostic Pathway in Suspected Appendicitis and Related Computed Tomography Use
Bibliographic record
Abstract
OBJECTIVES: The primary objective was to determine the diagnostic accuracy of a serial ultrasound (US) clinical diagnostic pathway to detect appendicitis in children presenting to the emergency department (ED). The secondary objective was to examine the diagnostic performance of the initial and interval US and to compare the accuracy of the pathway to that of the initial US. METHODS: This was a prospective cohort study of 294 previously healthy children 4 to 17 years old with suspected appendicitis and baseline pediatric appendicitis scores of ≥2, who were managed with the serial US clinical diagnostic pathway. This pathway consisted of an initial US followed by a clinical reassessment in each patient and an interval US and surgical consultation in patients with equivocal initial US and persistent concern about appendicitis. The USs were interpreted by published criteria as positive, negative, or equivocal for appendicitis. Children in whom this pathway did not rule in or rule out appendicitis underwent computed tomography (CT). Cases with missed appendicitis, negative operations, and CTs after the pathway were considered inaccurate. The primary outcome was the diagnostic accuracy of the serial US clinical diagnostic pathway. The secondary outcomes included the test performance of the initial and interval US imaging studies. RESULTS: Of the 294 study children, 111 (38%) had appendicitis. Using the serial US clinical diagnostic pathway, 274 of 294 children (93%, 95% confidence interval [CI] = 90% to 96%) had diagnostically accurate results: 108 of the 111 (97%) appendicitis cases were successfully identified by the pathway without CT scans (two missed and one CT), and 166 of the 183 (91%) negative cases were ruled out without CT scans (14 negative operations and three CTs). The sensitivity of this pathway was 108 of 111 (97%, 95% CI = 94% to 100%), specificity 166 of 183 (91%, 95% CI = 87% to 95%), positive predictive value 108 of 125 (86%; 95% CI = 79% to 92%), and negative predictive value 166 of 169 (98%, 95% CI = 96% to 100%). The diagnostic accuracy of the pathway was higher than that of the initial US alone (274 of 294 vs. 160 of 294; p < 0.0001). Of 123 patients with equivocal initial US, concern about appendicitis subsided on clinical reassessment in 73 (no surgery and no missed appendicitis). Of 50 children with persistent symptoms, 40 underwent interval US and 10 had surgical consultation alone. The interval US confirmed or ruled out appendicitis in 22 of 40 children (55.0%) with equivocal initial US, with one false-positive interval US. CONCLUSIONS: The serial US clinical diagnostic pathway in suspected appendicitis has an acceptable diagnostic accuracy that is significantly higher than that of the initial US and results in few CT scans. This approach appears most useful in children with equivocal initial US, in whom the majority of negative cases were identified at clinical reassessment and appendicitis was diagnosed by interval US or surgical consultation in most study patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".