Predicting cognitive performance using eye-movements, reaction time and difficulty level.
Bibliographic record
Abstract
Cognitively challenging tasks require complex coordination of information beyond visual input. Predicting accuracy on such tasks has potential applications in education and industry. Task difficulty is associated with increases in reaction time and variation in eye tracking indices. Critically, machine learning has not yet been used to predict accuracy on cognitive tasks with multiple difficulty levels. We report data on 57 (34 females; 20-30 years) participants who completed visuospatial tasks of mental attentional capacity with six levels of difficulty while their eye movements were recorded using EyeLink Portable Duo SR Research eye-tracker with 1ms temporal resolution (at 1000 Hz frequency) in remote head-free-to-move mode. Results show that task accuracy scores can be robustly predicted when all variables (e.g., eye-tracking, difficulty level and reaction time) are considered together (R2 = .80). Reaction time, difficulty level and eye tracking metrics are also effective independent predictors with R2 equaling .73, .58, and .36, respectively. Analyses for feature importance suggest eye-tracking indices with the most importance for the models include the number of fixations, number of saccades, duration of the current fixation and pupil size. Notably, our machine learning algorithms target a prediction question, rather than a classification one, and the current algorithm can be useful for future research and applications in other contexts where visuospatial processing is required. Theoretically, findings show common and distinct metrics that can inform theories of cognition and vision science.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".