Measuring Infant Visual Acuity with Gaze Tracker Monitored Visual Fixation
Bibliographic record
Abstract
PURPOSE: To validate a method of measuring grating acuity with remote gaze tracking (GT) against a current clinical test of visual acuity (VA), the Teller Acuity Cards (TACs), as part of the development of an automated VA test for infants. METHODS: Visual acuity for computer-generated horizontal square-wave gratings was determined from relative fixation time on a grating area compared with the background. In experiment 1, binocular VA was based on eye movements with a GT in 15 uncorrected myopic adults and compared with VA measured with subjective responses with the same stimuli and with the TACs. In experiment 2, binocular VA was determined in 19 typically developing infants aged 3 to 11 months on two visits with both the GT and TACs. RESULTS: In adults, the mean difference between VA measured by the GT and TACs was 0.01 log cycles per degree (cpd) and the 95% limits of agreement were 0.11. One hundred percent of GT VA results were within 0.5 octave of the TACs' VAs. The mean difference between the GT and TACs for infants was 0.17 log cpd on both the first and second visit (95% limits of agreement, 0.42 and 0.47, respectively). The mean difference between test and retest for infant GT VA was 0.06 log cpd, and limits of agreement for repeatability were 0.48 log cpd. In infants, both the TACs and the GT had a reliability of 89% within less than or equal to 1 octave between visits. Gaze tracking VA improved with age and is in agreement with published norms. CONCLUSIONS: The agreement between the TACs and GT in adults and infants validates the method of measuring grating acuity with the remote GT. These results demonstrate its potential for an automated test of infant VA.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".