Comparing Telephone Survey Responses to Best-Corrected Visual Acuity to Estimate the Accuracy of Identifying Vision Loss: Validation Study
Bibliographic record
Abstract
BACKGROUND: Self-reported questions on blindness and vision problems are collected in many national surveys. Recently released surveillance estimates on the prevalence of vision loss used self-reported data to predict variation in the prevalence of objectively measured acuity loss among population groups for whom examination data are not available. However, the validity of self-reported measures to predict prevalence and disparities in visual acuity has not been established. OBJECTIVE: This study aimed to estimate the diagnostic accuracy of self-reported vision loss measures compared to best-corrected visual acuity (BCVA), inform the design and selection of questions for future data collection, and identify the concordance between self-reported vision and measured acuity at the population level to support ongoing surveillance efforts. METHODS: We calculated accuracy and correlation between self-reported visual function versus BCVA at the individual and population level among patients from the University of Washington ophthalmology or optometry clinics with a prior eye examination, randomly oversampled for visual acuity loss or diagnosed eye diseases. Self-reported visual function was collected via telephone survey. BCVA was determined based on retrospective chart review. Diagnostic accuracy of questions at the person level was measured based on the area under the receiver operator curve (AUC), whereas population-level accuracy was determined based on correlation. RESULTS: The survey question, "Are you blind or do you have serious difficulty seeing, even when wearing glasses?" had the highest accuracy for identifying patients with blindness (BCVA ≤20/200; AUC=0.797). The highest accuracy for detecting any vision loss (BCVA <20/40) was achieved by responses of "fair," "poor," or "very poor" to the question, "At the present time, would you say your eyesight, with glasses or contact lenses if you wear them, is excellent, good, fair, poor, or very poor" (AUC=0.716). At the population level, the relative relationship between prevalence based on survey questions and BCVA remained stable for most demographic groups, with the only exceptions being groups with small sample sizes, and these differences were generally not significant. CONCLUSIONS: Although survey questions are not considered to be sufficiently accurate to be used as a diagnostic test at the individual level, we did find relatively high levels of accuracy for some questions. At the population level, we found that the relative prevalence of the 2 most accurate survey questions were highly correlated with the prevalence of measured visual acuity loss among nearly all demographic groups. The results of this study suggest that self-reported vision questions fielded in national surveys are likely to yield an accurate and stable signal of vision loss across different population groups, although the actual measure of prevalence from these questions is not directly analogous to that of BCVA.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.034 | 0.110 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".