Teleophthalmology-enabled and artificial intelligence-ready referral pathway for community optometry referrals of retinal disease (HERMES): a Cluster Randomised Superiority Trial with a linked Diagnostic Accuracy Study—HERMES study report 1—study protocol
Bibliographic record
Abstract
INTRODUCTION: Recent years have witnessed an upsurge of demand in eye care services in the UK. With a large proportion of patients referred to Hospital Eye Services (HES) for diagnostics and disease management, the referral process results in unnecessary referrals from erroneous diagnoses and delays in access to appropriate treatment. A potential solution is a teleophthalmology digital referral pathway linking community optometry and HES. METHODS AND ANALYSIS: The HERMES study (Teleophthalmology-enabled and artificial intelligence-ready referral pathway for community optometry referrals of retinal disease: a cluster randomised superiority trial with a linked diagnostic accuracy study) is a cluster randomised clinical trial for evaluating the effectiveness of a teleophthalmology referral pathway between community optometry and HES for retinal diseases. Nested within HERMES is a diagnostic accuracy study, which assesses the accuracy of an artificial intelligence (AI) decision support system (DSS) for automated diagnosis and referral recommendation. A postimplementation, observational substudy, a within-trial economic evaluation and discrete choice experiment will assess the feasibility of implementation of both digital technologies within a real-life setting. Patients with a suspicion of retinal disease, undergoing eye examination and optical coherence tomography (OCT) scans, will be recruited across 24 optometry practices in the UK. Optometry practices will be randomised to standard care or teleophthalmology. The primary outcome is the proportion of false-positive referrals (unnecessary HES visits) in the current referral pathway compared with the teleophthalmology referral pathway. OCT scans will be interpreted by the AI DSS, which provides a diagnosis and referral decision and the primary outcome for the AI diagnostic study is diagnostic accuracy of the referral decision made by the Moorfields-DeepMind AI system. Secondary outcomes relate to inappropriate referral rate, cost-effectiveness analyses and human-computer interaction (HCI) analyses. ETHICS AND DISSEMINATION: Ethical approval was obtained from the London-Bromley Research Ethics Committee (REC 20/LO/1299). Findings will be reported through academic journals in ophthalmology, health services research and HCI. TRIAL REGISTRATION NUMBER: ISRCTN18106677 (protocol V.1.1).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.025 |
| Meta-epidemiology (narrow) | 0.005 | 0.003 |
| Meta-epidemiology (broad) | 0.006 | 0.004 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.006 | 0.004 |
| Insufficient payload (model declined to judge) | 0.057 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".