Machine Learning Models Can Predict Tinnitus and Noise-Induced Hearing Loss
Bibliographic record
Abstract
OBJECTIVES: Despite the extensive use of machine learning (ML) models in health sciences for outcome prediction and condition classification, their application in differentiating various types of auditory disorders remains limited. This study aimed to address this gap by evaluating the efficacy of five ML models in distinguishing (a) individuals with tinnitus from those without tinnitus and (b) noise-induced hearing loss (NIHL) from age-related hearing loss (ARHL). DESIGN: We used data from a cross-sectional study of the Canadian population, which included audiologic and demographic information from 928 adults aged 30 to 100 years, diagnosed with either ARHL or NIHL due to long-term occupational noise exposure. The ML models applied in this study were artificial neural networks (ANNs), K-nearest neighbors, logistic regression, random forest (RF), and support vector machines. RESULTS: The study revealed that tinnitus prevalence was over twice as high in the NIHL group compared with the ARHL group, with a frequency of 27.85% versus 8.85% in constant tinnitus and 18.55% versus 10.86% in intermittent tinnitus. In pattern recognition, significantly greater hearing loss was found at medium- and high-band frequencies in NIHL versus ARHL. In both NIHL and ARHL, individuals with tinnitus showed better pure-tone sensitivity than those without tinnitus. Among the ML models, ANN achieved the highest overall accuracy (70%), precision (60%), and F1-score (87%) for predicting tinnitus, with an area under the curve of 0.71. RF outperformed other models in differentiating NIHL from ARHL, with the highest precision (79% for NIHL, 85% for ARHL), recall (85% for NIHL), F1-score (81% for NIHL), and area under the curve (0.90). CONCLUSIONS: Our findings highlight the application of ML models, particularly ANN and RF, in advancing diagnostic precision for tinnitus and NIHL, potentially providing a framework for integrating ML techniques into clinical audiology for improved diagnostic precision. Future research is suggested to expand datasets to include diverse populations and integrate longitudinal data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".