DeepRetina Framework for Multi-retinal Diseases Classification
Bibliographic record
Abstract
Abstract Diagnosing retinal diseases is a fundamental challenge in developing robust multi-disease classification systems due to the limited availability of datasets and inconsistent quality. Therefore, this study presents $$\text{DeepRetina}$$ <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mtext>DeepRetina</mml:mtext> </mml:math> , a framework addressing these challenges through dataset harmonization. Five distinct retinal image datasets were unified using systematic preprocessing, resulting in a consolidated dataset of 29,966 high-resolution fundus images across eight disease categories. Harmonization corrects variations in image quality, color, and lighting resulting from different imaging devices or conditions. Data harmonization enhances the model’s ability to generalize across diverse datasets by standardizing the color and texture properties of images. The study compares the performance of custom CNN, EfficientNetV2, and MobileNetV3Large architectures for multi-disease classification. EfficientNetV2 achieved the highest accuracy of 79% with a precision of 54%. The proposed methodology significantly advances the field by (1) establishing a robust approach for harmonizing heterogeneous datasets, (2) presenting a large-scale, unified dataset for future research, and (3) presenting a comparative analysis of deep learning architectures optimized for retinal disease classification. $$\text{DeepRetina}$$ <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mtext>DeepRetina</mml:mtext> </mml:math> lays the foundation for scalable and accurate automated retinal disease diagnosis, contributing to improved detection and classification in ophthalmology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".