CNSC-52. SPEECH-DERIVED LINGUISTIC AND SPECTRAL FEATURES PREDICT MOLECULAR SUBTYPES OF ADULT-TYPE DIFFUSE GLIOMAS
Bibliographic record
Abstract
Abstract INTRODUCTION Gliomas engage in synaptic integration with surrounding neural circuitry. Previous work demonstrated that glioma-infiltrated cortex retains the capacity to encode speech, though it relies on a more diffuse network. While behavioral variability across glioma subtypes has been reported, no prior study has evaluated whether speech characteristics can predict glioma subtype. Here, we conduct a comprehensive linguistic and spectral analysis of speech in patients with gliomas to determine whether preoperative features can predict molecular features including IDH mutation status. METHODS This study utilized a prospective registry of patients with gliomas who completed semantic naming tasks. Recordings were annotated for error types by a single blinded observer. Audio signals were filtered using notch, Butterworth, and low-pass filters before spectral decomposition. Power band and power spectral density (PSD) features were extracted. Non-parametric tests and Spearman correlations were used to compare IDH wild-type and mutant groups. A gradient boosting machine learning classifier was trained using a 70/30 train-test split. RESULTS A total of 100 patients were included. IDH wild-type patients were older (62.0 vs. 39.8 years, p<0.001), more often white (89.7% vs. 68.4%, p=0.007), and exhibited greater FLAIR volume, higher Ki67 indices, and lower MGMT promoter methylation. IDH wild-type patients exhibited higher anomia and circumlocution rates on picture naming (p = 0.018, p = 0.043) and increased overall errors, anomia, verbal paraphasias, and delays on auditory naming (p < 0.001–0.02). Spectral analysis revealed group differences in clarity and brilliance bands (p=0.029, 0.01), and PSD analysis revealed differences in bass, midrange, and brilliance power (p=0.041–0.05). The gradient boosting model achieved an AUC of 0.89, with top predictive features including age, brilliance band power, midrange PSD, and auditory naming error rate. CONCLUSIONS Pre-operative speech features differ significantly by IDH status. Spectral-linguistic profiling coupled with machine learning can non-invasively predict glioma molecular subtype.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".