Bibliographic record
Abstract
Background: Vision tests are noninvasive tests increasingly suggested for concussion assessment. To know if a test result is abnormal after a concussion, one needs to compare the result to a value obtained before a concussion (baseline). The test results may be affected by physiological and psychological changes of the patient over time and variations in how the clinician conducts the examination. Therefore, one needs to know how much the results may vary even without a concussion in order to determine if the change with a concussion (signal) is more than the expected change due to normal variability (noise). However, only a limited number of studies have assessed the test-retest reliability of vision tests, and none have evaluated test-retest time intervals longer than about 57 days. Because concussions may occur months after a yearly baseline test, appropriate interpretation of results can only be made if we understand one-year test-retest reliability.Objective: Our objective was to determine the one-year test-retest reliability of ten vision tests in a cohort of healthy Canadian elite athletes, who did not suffer a concussion during one-year. These ten tests measured different aspects of visual function, including Positive Fusional Vergence at 30cm and 3m, Negative Fusional Vergence at 30cm and 3m, Phoria at 30cm and 3m, Near Point of Convergence and Near Point of Convergence break, Gross Stereoscopic Acuity, and Saccades.Methods: We studied elite Canadian athletes followed at the Institut National du Sport du Quebec (INSQ), evaluated by a single INSQ sports medicine physician who was responsible for all vision testing referrals. The vision test data was abstracted from the medical records of a single clinician trained in orthoptic testing (APEXK) performing the vision tests. The two data sets were linked, cleaned and harmonized according to a pre-determined set of rules. After data verification, we included athletes who completed two baseline evaluations within 365±30 days. We excluded athletes with any concussion, vision training in between the annual evaluations, or medical conditions/treatment that might affect the tests. We evaluated test-retest reliability using Intraclass Correlation Coefficient (ICC) and 95% limits of agreement (95% LoA). We considered ICC of ≤0.5 as poor, 0.51-0.74 as moderate, 0.75-0.89 as good, and ≥0.90 as excellent reliability.Results: There were 16 athletes, nine females and seven males, out of 199 who met our inclusion criteria, with a mean age of 22.7 (SD 4.5) years. Among the vision tests, we observed excellent test-retest reliability in Positive Fusional Vergence at 30cm (ICC=0.93, 95% LoA=±41.9%), but the ICC dropped to 0.53 (95% LoA= ±43.5%) when an outlier was excluded in a sensitivity analysis. There was good to moderate reliability in Negative Fusional Vergence at 30cm (ICC=0.78, 95% LoA=±41.2%), Phoria at 30cm (ICC=0.68, 95% LoA=±119.2%), Near Point of Convergence break (ICC=0.65, 95% LoA=±49.4%) and Saccades (ICC=0.61, 95% LoA=±24.3%). The ICC for Positive Fusional Vergence at 3m (ICC=0.56, 95% LoA=±60.2%) also decreased to 0.21 after removing one outlier. We found poor reliability in Near Point of Convergence (ICC=0.47, 95% LoA=±73.9%), Gross Stereoscopic Acuity (ICC=0.03, 95% LoA=±92.5%) and Negative Fusional Vergence at 3m (ICC=0.0, 95% LoA=±48.4%). ICC for Phoria at 3m was not appropriate because scores were identical in 14/16 athletes.Conclusions: Five vision tests had good to moderate one-year test-retest reliability after removing outliers, and the remaining tests had poor reliability in this healthy athlete population. These results represent the “noise” of the tests. Future research is needed to investigate if concussions produce a signal that is greater than the noise. If this is the case, these vision tests will be useful clinically regardless of the absolute ICCs and LoA values
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".