Smartphone Pupillometry and Machine Learning for Detection of Acute Mild Traumatic Brain Injury: Cohort Study
Bibliographic record
Abstract
Background Quantitative pupillometry is used in mild traumatic brain injury (mTBI) with changes in pupil reactivity noted after blast injury, chronic mTBI, and sports-related concussion. Objective We evaluated the diagnostic capabilities of a smartphone-based digital pupillometer to differentiate patients with mTBI in the emergency department from controls. Methods Adult patients diagnosed with acute mTBI with normal neuroimaging were evaluated in an emergency department within 36 hours of injury (control group: healthy adults). The PupilScreen smartphone pupillometer was used to measure the pupillary light reflex (PLR), and quantitative curve morphological parameters of the PLR were compared between mTBI and healthy controls. To address the class imbalance in our sample, a synthetic minority oversampling technique was applied. All possible combinations of PLR parameters produced by the smartphone pupillometer were then applied as features to 4 binary classification machine learning algorithms: random forest, k-nearest neighbors, support vector machine, and logistic regression. A 10-fold cross-validation technique stratified by cohort was used to produce accuracy, sensitivity, specificity, area under the curve, and F1-score metrics for the classification of mTBI versus healthy participants. Results Of 12 patients with acute mTBI, 33% (4/12) were female (mean age 54.1, SD 22.2 years), and 58% (7/12) were White with a median Glasgow Coma Scale (GCS) of 15. Of the 132 healthy patients, 67% (88/132) were female, with a mean age of 36 (SD 10.2) years and 64% (84/132) were White with a median GCS of 15. Significant differences were observed in PLR recordings between healthy controls and patients with acute mTBI in the PLR parameters, that are (1) percent change (mean 34%, SD 8.3% vs mean 26%, SD 7.9%; P<.001), (2) minimum pupillary diameter (mean 34.8, SD 6.1 pixels vs mean 29.7, SD 6.1 pixels; P=.004), (3) maximum pupillary diameter (mean 53.6, SD 12.4 pixels vs mean 40.9, SD 11.9 pixels; P<.001), and (4) mean constriction velocity (mean 11.5, SD 5.0 pixels/second vs mean 6.8, SD 3.0 pixels/second; P<.001) between cohorts. After the synthetic minority oversampling technique, both cohorts had a sample size of 132 recordings. The best-performing binary classification model was a random forest model using the PLR parameters of latency, percent change, maximum diameter, minimum diameter, mean constriction velocity, and maximum constriction velocity as features. This model produced an overall accuracy of 93.5%, sensitivity of 96.2%, specificity of 90.9%, area under the curve of 0.936, and F1-score of 93.7% for differentiating between pupillary changes in mTBI and healthy participants. The absolute values are unable to be provided for the performance percentages reported here due to the mechanism of 10-fold cross validation that was used to obtain them. Conclusions In this pilot study, quantitative smartphone pupillometry demonstrates the potential to be a useful tool in the future diagnosis of acute mTBI.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".