P198 Practical deep learning tool for the scoring of ulcerative colitis disease activity in central reading
Bibliographic record
Abstract
Abstract Background Central Reading Org.& the pharma industry employ subject matter experts(SMEs)to score videos from sites participating in clinical trials for Ulcerative Colitis(UC).As we are developing Artificial Intelligence(AI)models for scoring purposes,we need to build a new software interface that can incorporate these AI models to aid SMEs,making their determination of the scores for each segment of the video & for the video as a whole.We propose a system that reduces the time for SMEs to review & score videos,improving the accuracy of scoring,with the help of our AI models. Methods We built a web-based interface supported by our AI models which can read,write multiple databases & data stores to read & display videos to be scored by a central reader, as well as the associated metadata required to improve the process.User interface shows a timeline with markers for the segments of the colon,with sections that are blurry,poorly prepped,or unscorable highlighted in different colours.While we could also highlight sections of the video with the precise score assigned to it by the AI,this would bias the central reader’s opinion.We hide the precise score generated by our AI models & instead display 3 colours for low,medium or high disease activity.When a video is loaded to be read by the user,the playback marker is set to the first high disease activity section based on known medical indexes such as the Mayo Endoscopic Subscore(MES) & UCEIS(Ulcerative Colitis Endoscopic Index of Severity),usually consists of a few seconds of video & that video is played back continuously in a loop until the reader selects the appropriate score for that section.When the reader saves the section,the software immediately moves the video cursor to the highest scored section of the video.That way the central reader can review only the relevant portions of the video to confirm the score assigned to each segment.If the central reader’s scores do not align well with the AI scores then the software continues to show more sections of the video to the user,including sections it may have labeled as unscorable,that may be scorable. Results The review of the system by 3 key opinion leaders,user experience was positive.Not only does the system allow the reader’s attention to be more efficiently used,but the interface allows both AI & central reader scores to be saved,allowing for the latter to be used iteratively to re-train & improve the underlying AI model(s).Our tool was also used by a gastroenterologist specialist in order to perform video quality assessment & colon sections scoring. Conclusion We developed an AI tool that can be used to improve the efficiency & accuracy of the central reading process in clinical trials for UC.Further work is ongoing to improve the interface.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.041 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".