MétaCan
Menu
Back to cohort

Performance of Hearing Test Software Applications to Detect Hearing Loss

2025· article· en· W4408905605 on OpenAlexafffundabout
Meaghan Lunney, Natasha Wiebe, Tanis Howarth, Lorienne M. Jenstad, Alex DeBusschere, Gillian Crysdale, Sharon E. Straus, Kara Schick‐Makaroff, Maoliosa Donald, Stephanie Thompson, Jayna Holroyd‐Leduc, Marcello Tonelli

Bibliographic record

VenueJAMA Network Open · 2025
Typearticle
Languageen
FieldNeuroscience
TopicHearing Loss and Rehabilitation
Canadian institutionsUniversity of TorontoUniversity of British ColumbiaAlberta Health ServicesUniversity of AlbertaUniversity of Calgary
FundersCanadian Institutes of Health Research
KeywordsHearing lossAudiologyMedicineTest (biology)Reliability (semiconductor)Quality of life (healthcare)Hearing test

Abstract

fetched live from OpenAlex

Importance: Hearing loss is common and may impact health and quality of life if not properly managed. It is diagnosed following formal audiological assessment, which may not be available or practical. Hearing test software applications (apps) may help identify people who might benefit from audiological assessment, but their diagnostic accuracy has been incompletely studied. Objective: To measure and compare the validity and reliability of 2 commonly recommended apps (hearWHO and SHOEBOX) to detect moderately severe or greater hearing loss. Secondary objectives were to evaluate the apps' ability to detect less severe hearing loss and the diagnostic performance of 2 questionnaires for detecting both severities of hearing loss. Design, Setting, and Participants: This prospective diagnostic accuracy study compared the hearWHO and SHOEBOX apps with a 4-frequency pure-tone average audiological assessment reference standard. All consenting English-speaking patients aged 18 years or older and referred for routine audiological assessment at a publicly funded health center in Calgary, Canada, were included between May 17, 2023, and March 12, 2024. Main Outcome and Measures: The main outcome was the validity and reliability of 4 index tests, including the hearWHO app, SHOEBOX app, Revised Hearing Handicap Inventory-Screening (RHHI-S) questionnaire, and the Single-Item Self-Assessment (SISA) questionnaire, to detect moderate to severe hearing loss. All index test results were compared with an audiological assessment reference standard (hearing loss defined by a better ear hearing threshold of ≥50 dB [more severe denoted as HL50] or ≥20 dB [less severe denoted as HL20]). Test-retest reliability of the 2 apps and C statistics, sensitivity, specificity, and positive and negative predicted values of all index tests were measured. Results: A total of 130 participants were recruited (median [IQR] age, 58 [47-67] years; 82 female [63.1%]). Complete data for each comparison ranged from 123 to 129 participants. The prevalence of HL50 was 16.3% (21 or 130 participants). Neither the hearWHO nor the SHOEBOX app had high test-retest reliability (all κ-values <0.80), with the SHOEBOX having a κ of 0.64 (95% CI, 0.48-0.79) and hearWHO having a κ of 0.32 (95% CI, 0.18-0.46). All C statistics for HL50 were less than 0.80. When testing for HL50, diagnostic performance for both apps was better for the second measurement than the first measurement or the mean. Sensitivity and specificity for the second measurement of SHOEBOX were 0.26 (95% CI, 0.09-0.51) and 1.00 (95% CI, 0.97-1.00), respectively, and for the second measurement of hearWHO, 0.67 (95% CI, 0.43-0.85) and 0.71 (95% CI, 0.62-0.79), respectively. Sensitivity and specificity for the RHHI-S were 0.76 (95% CI, 0.53-0.92) and 0.42 (95% CI, 0.32-0.52), respectively, and for SISA, 0.10 (95% CI, 0.01-0.30) and 0.90 (95% CI, 0.83-0.95), respectively. Using a less stringent diagnostic threshold with SHOEBOX increased sensitivity for HL50 to at least 95% while retaining a specificity of 47% to 54%. Sensitivity and specificity for both apps were higher for HL20. Conclusions and Relevance: These findings suggest that both hearWHO and SHOEBOX have limited test-retest reliability, perhaps because of a learning effect. Both apps may be suitable if a sensitive strategy is desired for identifying people who may benefit from diagnostic audiological assessment, whereas the SHOEBOX app may be preferable if a specific strategy is desired. If neither app is available, the RHHI-S or the SISA could be used depending on whether sensitivity or specificity is desired.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.417
Threshold uncertainty score0.387

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.036
GPT teacher head0.314
Teacher spread0.279 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2025
Admission routes3
Has abstractyes

Explore more

Same venueJAMA Network OpenSame topicHearing Loss and RehabilitationFrench-language works237,207