Psychophysics of variable fonts: Gaze measures of reading efficiency
Bibliographic record
Abstract
Reading is a demanding task, and while many studies have investigated the visual factors associated with reading, the recent development of variable fonts opens new avenues for this research. Variable fonts can be customized along a set of continuous parametric axes within a single font file (e.g., thin stroke, slant, etc.), lending themselves readily to psychophysical techniques. To understand how these settings can influence individual reading performance and which settings may improve reading efficiency, we recorded participants’ eye movements as they read short passages. For this study, we varied five font parameters within Roboto Flex: thick stroke, thin stroke, slant, weight, and width at five levels each. Participants read one passage per setting, displayed across four screens, and we measured saccade amplitude normalized to letter width as well as the number and duration of fixations. Our results demonstrate that increasing width and weight decrease reading efficiency since saccade amplitude decreased as letters became wider and visually heavier. Increasing thick stroke had the largest effect on reading efficiency, while thin stroke and slant had the smallest. We also found considerable individual variability in the degree to which these axes impacted individuals’ reading efficiency and the number of fixations they made. Our results suggest that the customizability of variable fonts and the sensitivity of our gaze measures may make it possible to quickly find the settings that are best for each reader and enable a new range of psychophysical investigations of the impact of font on reading.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".