Disentangling consonants and vowels in auditory cortices using an oscillation paradigm
Bibliographic record
Abstract
The auditory cortices contain tonotopic maps1. Phonemes may to some degree be organized similarly in phonotopic maps2,3. However, fMRI’s low temporal resolution challenges the localization of fast speed sounds. Here, we combined an oscillation-based protocol with continuous repetitive syllable presentation during fast fMRI acquisition to map phonemes in the brain. We aimed to disentangle the effects of vowels and consonants, despite being presented together in syllables, at two different oscillation frequencies4. We acquired fMRI scans of 23 healthy participants using a fast fMRI protocol at 3T (TR=371ms, multiband-EPI). Participants listened to continuous auditory stimuli with one 0.5s syllable (one consonant and one vowel) repeated twice per second (e.g., ba-ba-de-de-gi-gi), using two conditions (9v5c-9c5v). The 9v5c-condition combined nine Danish vowels with five consonants (45 unique syllables), with vowels (consonants) being repeated at every 9th (5th) trial, creating both new combinations and two highly predictable non-interfering oscillations. The 9c5v-condition combined nine consonants with five vowels. We ran 3 sessions of 18min with 6x4 blocks (6x9v5c-6x9c5v-6x9v5c-6x9c5v). To secure attention, participants had to respond to rare mismatches (two per 45s) (e.g., ba-ba-de-du-gi-gi). Preprocessing utilized the standard protocol in SPM12, including motion correction, normalization, and 4mm-FWHM smoothing. Oscillations in the data were modelled for each participant with sine and cosine waves at each presentation frequency (1/9Hz-1/5Hz). Using trigonometry, we converted the model’s beta estimates into amplitude maps for each condition and frequency, which indicated activation magnitude. A 2nd-level ANOVA-model across all subjects estimated the main effect of phoneme type, FWE-corrected (p<0.05). Highest amplitudes for the different conditions and frequencies localized to the auditory cortices. A main effect of phoneme type (consonants vs. vowels) was likewise observed in both auditory hemispheres. Consonants had a larger amplitude than vowels in both left [-42,-36,14] and right auditory cortex [56,-24,14], regardless of stimulus frequency, whereas vowels did not yield higher amplitude in any area. Perhaps because consonants cover a broader frequency spectrum and therefore activated a larger area. 1/9Hz oscillations yielded larger amplitudes than 1/5Hz across many brain areas, possibly because the natural rhythm of the BOLD signal is closer to 1/9Hz. We successfully differentiated between vowels and consonants despite the continuous stimulus with intermixed vowels and consonants. This oscillation-based method is a step towards faster fMRI protocols with more natural stimuli. However, several posterior areas were coincidentally activated at one of the chosen frequencies, making it imperative to control for oscillation power and frequency when comparing BOLD responses. References 1 Saenz, M., & Langers, D. R. (2014). Tonotopic mapping of human auditory cortex. Hearing Research, 307, 42-52, 10.1016/j.heares.2013.07.016. 2Formisano, E., De Martino, F., Bonte, M., & Goebel, R. (2008). “Who” is saying “what”? Brain-based decoding of human voice and speech. Science (New York, NY), 322, 970-973, 10.1126/science.1164318. 3Wallentin, M., Lund, T. E., Andersen, C. M., & Rocca, R. (2018). Fast phonotopic mapping with oscillation-based fMRI – Proof of concept In Society for the Neurobiology of Language. Quebec, 4Lewis, L. D., Setsompop, K., Rosen, B. R., & Polimeni, J. R. (2016). Fast fMRI can detect oscillatory neural activity in humans. Proceedings of the National Academy of Sciences of the United States of America, 113, E6679-E6685, 10.1073/pnas.1608117113.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".