Evaluation of a Community-Based AI-Assisted Visual Impairment Screening Model for Performance, Operational Efficiency, Acceptability, Feasibility, and Costs: Protocol for a 2-Arm Pragmatic Randomized Controlled Trial
Bibliographic record
Abstract
Background: Visual impairment (VI) affects more than 600 million people globally and significantly reduces quality of life. In Singapore, 20% of adults aged 60 years and older (~180,000 people) have VI, a figure expected to double by 2030 due to population aging. While about half of VI cases are due to uncorrected refractive errors, the rest are caused by age-related diseases. The current traditional screening model is a 2-visit, labor-intensive approach with low follow-up rates and frequent unnecessary referrals. Although AI for Disease-related Visual Impairment Screening Using Retinal Imaging, the deep learning model in this study, has demonstrated strong diagnostic performance in retrospective datasets (area under the curve=0.942), key aspects of real-world implementation such as operational efficiency, patient acceptability, workflow feasibility, and cost remain insufficiently studied. As a result, real-world evidence directly comparing artificial intelligence (AI)-assisted and traditional screening pathways is limited. Objective: This study aims to evaluate the referral accuracy, operational efficiency, acceptability, feasibility, and cost of an AI-assisted screening model compared with the current traditional screening model. Methods: This study aims to recruit 1000 participants aged 50 years and older using a 2-arm pragmatic randomized controlled trial design. Participants with presenting visual acuity worse than 6/12 (L2) will be randomized 1:1 into either the AI-assisted or traditional screening arms. In the AI-assisted arm, the AI model will analyze retinal photos on-site, with positive cases referred to an optometrist for secondary evaluation. The AI model, previously developed with promising diagnostic accuracy and further validated using community-acquired data, has been integrated with a custom user interface for use in this study. Traditional screening will include pinhole visual acuity, intraocular pressure, slit lamp examination, auto refraction, and retinal photography. All L2 participants will complete a patient-acceptance questionnaire and undergo assessments to determine ground truth. Results: The study was funded in 2022. Participant recruitment commenced in July 2024, with 487 participants enrolled as of September 14, 2024. Recruitment is ongoing, with study completion anticipated by March 2026 and data analysis expected to begin in April 2026. Conclusions: This study will provide critical evidence on the clinical utility, feasibility, and cost analysis of AI-assisted VI screening. Our findings may contribute real-world evidence to inform scalable, sustainable screening strategies that enhance efficiency, accuracy, and health system outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.043 | 0.054 |
| Meta-epidemiology (narrow) | 0.006 | 0.002 |
| Meta-epidemiology (broad) | 0.008 | 0.007 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.006 | 0.007 |
| Insufficient payload (model declined to judge) | 0.039 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".