Applying artificial intelligence in the prediction of cardiovascular risk from ophthalmic imaging: a systematic review
Bibliographic record
Abstract
Abstract Background The rising burden of cardiovascular disease (CVD) has spurred the development of innovative, non-invasive screening methods. Advances in artificial intelligence (AI) and deep learning (DL) applied to retinal imaging offer a promising avenue, as changes in the retinal vasculature can mirror systemic vascular health. Purpose To synthesize evidence from studies using artificial intelligence (AI) algorithms trained on retinal images to predict cardiovascular disease (CVD) risk. Methods A systematic literature search was conducted on MEDLINE and Embase databases up to February 2024 in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analysis (PRISMA) guideline using relevant search terms such as; "cardiovascular disease," "artificial intelligence," "deep learning," "retinal imaging," "colour fundus photography," etc. Inclusion criteria were the development of a DL model applied to any ophthalmic imaging modality for predicting CVD risk or outcome. Results Of 9880 studies which were screened, 13 studies were included. All studies included general population databases, while 7 (54%) studies used databases that included patients with pre-existing CVD risk factors. All studies used retinal fundus images as input for the DL models, and most models (92%) analysed characteristics of the retinal vasculature (e.g., vessel calibre, venular dilatation, arteriolar narrowing, microaneurysms) for their prediction. Overall, 18 different CVD risk factors were predicted through DL models, with age (n=13; 100%), sex (n=11; 85%) and smoking status (n=9; 69%) being the most common. Five (38%) studies analysed binary CVD outcomes including incident myocardial infarction, stroke, and coronary atherosclerotic disease; Four (31%) studies compared CVD risk prediction to traditional CVD risk scores (e.g., Framingham risk score, European Systematic Coronary Risk Evaluation, etc.). These studies were able to accurately stratify cumulative CVD events into low, moderate, and high-risk groups via fundus imaging, demonstrating stratification comparable to established CVD risk scores and cardiac imaging modalities. In total, 8 (62%) studies performed an external validation: the area under the receiver operating characteristic curve ranged between 68.2% and 85.9%. Accuracy, specificity and sensitivity were measured in 4 (31%) studies. Respectively, they ranged between 58.3%-82.0%, 40.4-66.0% and 81.0-89.1%. Only one AI model (Reti-CVD) was made publicly available for clinical use. Conclusions In conclusion, recent studies using DL applied to fundus imaging to predict CVD risk mostly examine retinal vasculature to make predictions and can stratify CVD risk comparably to other clinical risk scores. Though many report promising performance, the majority have not been used in real clinical settings. Additional research is required to enable their clinical implementation in a primary care context, or in an ophthalmological setting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".