P336 Is endoscopy necessary to measure disease activity in ulcerative colitis: A systematic narrative review
Bibliographic record
Abstract
Optimal management of patients with ulcerative colitis (UC) requires the assessment of disease severity. Endoscopic evaluation is considered the gold standard approach but it is invasive and caries potential harmful consequences. We aimed to determine whether clinical symptoms correlate endoscopy for assessment of disease activity in UC patients. Systematic searches were performed (January 1980–September 2017) using MEDLINE, EMBASE, Scopus, CENTRAL and ISI Web of knowledge. Citation selection utilised a highly sensitive search strategy identifying studies with MeSH headings relating to: (i) ulcerative colitis, (ii) severity of illness index, (iii) invasive and non-invasive scoring systems. Recursive searches, cross-referencing, and subsequent hand-searches were completed. All fully published trials in French or English comparing patients reported outcomes with endoscopy were included. Trials comprising only paediatric or inpatients were excluded. Two investigators (S.R and C.-Y.C.) assessed citation eligibility with discrepancies resolved by an independent reviewer (P.L.L). The quality of trials was graded using the Quadas-2 tool. All data abstraction and entries were validated independently by two authors and a third authors re-validated few studies afterwards. Out of 3163 citations, 21 trials (1 RCT and 20 observational studies) fulfilled our inclusion criteria (n = 3172 patients, weighted average of mean age 41.3 ± 13.4, indication of endoscopy was an active disease in 18 trials). Clinical scales for UC severity assessment included SCCAI, pMAYO, Rachmilevich, Seo index, and patient-reported outcomes (diarrhoea and bleeding). Endoscopic scales as comparators encompassed UCCIS, MAYO, Rachmilevich, St Mark, Baron, and UCEIS. Comparison methods between clinical and endoscopic scores were conducted using correlation coefficient in 17 trials, while agreement coefficient and AUC in 1 and 2 trials, respectively. Dramatic heterogeneity precludes valid meta-analysis. Overall, mean correlation coefficients were very high when comparing clinical and endoscopic scores (partial Mayo 0.71 ± 0.14, range 0.49–0.97; Seo index 0.68 ± 0.06, range 0.64–0.72; SCCAI 0.67 ± 0.15 range 0.44–0.87; Rachmilevich 0.67 ± 0.11 range 0.52–0.79; patient reported outcomes 0.8 ± 0.12 range 0.71–0.88). AUC for the presence or absence of rectal bleeding was as high as 0.78. Mucosal healing can be predicted reliably by clinical symptoms, especially by the presence or absence of rectal bleeding. This is the first narrative review, suggesting that assessing and treating patients based on reported symptoms seems to be appropriate.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.079 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.010 | 0.009 |
| Bibliometrics | 0.011 | 0.012 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.005 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.003 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".