Improving Concordance Between Clinicians With Australian Guidelines for Bowel Cancer Prevention Using a Digital Application: Randomized Controlled Crossover Study
Bibliographic record
Abstract
Background Australia’s bowel cancer prevention guidelines, following a recent revision, are among the most complex in the world. Detailed decision tables outline screening or surveillance recommendations for 230 case scenarios alongside cessation recommendations for older patients. While these guidelines can help better allocate limited colonoscopy resources, their increasing complexity may limit their adoption and potential benefits. Therefore, tools to support clinicians in navigating these guidelines could be essential for national bowel cancer prevention efforts. Digital applications (DAs) represent a potentially inexpensive and scalable solution but are yet to be tested for this purpose. Objective This study aims to assess whether a DA could increase clinician adherence to Australia’s new colorectal cancer screening and surveillance guidelines and determine whether improved usability correlates with greater conformance to guidelines. Methods As part of a randomized controlled crossover study, we created a clinical vignette quiz to evaluate the efficacy of a DA in comparison with the standard resource (SR) for making screening and surveillance decisions. Briefings were provided to study participants, which were tailored to their level of familiarity with the guidelines. We measured the adherence of clinicians according to their number of guideline-concordant responses to the scenarios in the quiz using either the DA or the SR. The maximum score was 18, with higher scores indicating improved adherence. We also tested the DA’s usability using the System Usability Scale. Results Of 117 participants, 80 were included in the final analysis. Using the SR, the adherence of participants was rated a median (IQR) score of 10 (7.75-13) out of 18. The participants’ adherence improved by 40% (relative risk 1.4, P<.001) when using the DA, reaching a median (IQR) score of 14 (12-17) out of 18. The DA was rated highly for usability with a median (IQR) score of 90 (72.5-95) and ranked in the 96th percentile of systems. There was a moderate correlation between the usability of the DA and better adherence (rs=0.4; P<.001). No differences between the adherence of specialists and nonspecialists were found, either with the SR (10 vs 9; P=.47) or with the DA (13 vs 15; P=.24). There was no significant association between participants who were less adherent with the DA (n=17) and their age (P=.06), experience with decision support tools (P=.51), or academic involvement with a university (P=.39). Conclusions DAs can significantly improve the adoption of complex Australian bowel cancer prevention guidelines. As screening and surveillance guidelines become increasingly complex and personalized, these tools will be crucial to help clinicians accurately determine the most appropriate recommendations for their patients. Additional research to understand why some practitioners perform worse with DAs is required. Further improvements in application usability may optimize guideline concordance further.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".