P190 Systematic review: Gastrointestinal ultrasound scoring indices for inflammatory bowel disease
Bibliographic record
Abstract
Abstract Background The management of inflammatory bowel disease (IBD) requires frequent monitoring and assessment of disease activity. Endoscopic assessment with biopsy remains the gold standard for disease activity. Gastrointestinal ultrasound (GIUS) is a non-invasive, accessible and affordable test used to assess and monitor IBD and has been shown to be similar to MRI for detecting disease. The aim of this study was to systematically review the literature to identify scoring indices used for GIUS measurement of disease activity in IBD and to appraise their operating characteristics. Methods A systematic search of Embase, Medline, Pubmed, Cochrane Central and Clinical Trials.gov from inception to July 2019 was conducted according to PRISMA guidelines. Included were all study types reporting GIUS indices used for grading activity of severity of IBD in comparison to an objective reference standard. Studies using an exclusive clinical reference standard were excluded. All study types and abstracts were considered. Study quality was assessed using the QUADAS tool. Results 27 eligible studies were identified investigating 1647 patients. Disease phenotype was Crohn’s disease (CD) (n = 13), ulcerative colitis (UC) (n = 10) and IBD (n = 4). The most common reference standard was colonoscopy (n = 23), histology (n = 2), and imaging (n = 2). Bowel wall thickness was an index parameter in 26 studies. The most frequent cut off was 3mm (n = 10), 4mm (n = 9), 5mm (n = 1), and not specified (n = 6). There was no noticeable difference in magnitude of cut off when stratified by disease phenotype. Colour Doppler was an index parameter in 16 studies and was based on the Limburg score (n = 7), binary (n = 7) or categorical (n = 2). Bowel wall stratification was an index parameter in 15 studies and was more frequently used in UC (70%) and IBD (75%) indices than in CD indices (38%). Other index parameters included bowel wall compressibility, presence of complications such as abscess or fistula, bowel wall echogenicity, mesenteric inflammatory, lymphadenopathy, contrast enhancement, ulceration, peristalsis, strictures, absence of haustra coli, and tissue sonoelastography. Twenty-three studies were identified as at risk of bias. Overall concordance was substantial to excellent and accuracy was good to excellent. Two studies demonstrated substantial inter-observer agreement. No studies reported intra-observer agreement. Conclusion The identified GIUS scoring indices demonstrate applicability to both CD and UC with good accuracy and concordance. Current evidence does not adequately address concerns about the intra- and inter-observer variability of GIUS. There is a need for robust validation of an evidence-based GIUS index before more widespread use in IBD as a surrogate for colonoscopy and in clinical trials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.059 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.010 | 0.009 |
| Bibliometrics | 0.014 | 0.015 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.015 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".