Using EEG to decode semantics during an artificial language learning task
Bibliographic record
Abstract
BACKGROUND: As we learn a new nonnative language (L2), we begin to build a new map of concepts onto orthographic representations. Eventually, L2 can conjure as rich a semantic representation as our native language (L1). However, the neural processes for mapping a new orthographic representation to a familiar meaning are not well understood or characterized. METHODS: Using electroencephalography and an artificial language that maps symbols to English words, we show that it is possible to use machine learning models to detect a newly formed semantic mapping as it is acquired. RESULTS: Through a trial-by-trial analysis, we show that we can detect when a new semantic mapping is formed. Our results show that, like word meaning representations evoked by a L1, the localization of the newly formed neural representations is highly distributed, but the representation may emerge more slowly after the onset of the symbol. Furthermore, our mapping of word meanings to symbols removes the confound of the semantics to the visual characteristics of the stimulus, a confound that has been difficult to disentangle previously. CONCLUSION: We have shown that the L1 semantic representation conjured by a newly acquired L2 word can be detected using decoding techniques, and we give the first characterization of the emergence of that mapping. Our work opens up new possibilities for the study of semantic representations during L2 learning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".