EEG-GAN: A Generative EEG Augmentation Toolkit for Enhancing Neural Classification
Bibliographic record
Abstract
Abstract Electroencephalography (EEG) is a widely applied method for decoding neural activity, offering insights into cognitive function and driving advancements in neurotechnology. However, decoding EEG data remains challenging, as classification algorithms typically require large datasets that are expensive and time-consuming to collect. Recent advances in generative artificial intelligence have enabled the creation of realistic synthetic EEG data, yet no method has consistently demonstrated that such synthetic data can lead to improvements in EEG decodability across diverse datasets. Here, we introduce EEG-GAN, an open-source generative adversarial network (GAN) designed to augment EEG data. In the most comprehensive evaluation study to date, we assessed its capacity to generate realistic EEG samples and enhance classification performance across four datasets, five classifiers, and seven sample sizes, while benchmarking it against six established augmentation techniques. We found that EEG-GAN, when trained to generate raw single-trial EEG signals, produced signals that reproduce grand-averaged waveforms and time-frequency patterns of the original data. Furthermore, training classifiers on additional synthetic data improved their ability to decode held-out empirical data. EEG-GAN achieved up to a 16% improvement in decoding accuracy, with enhancements consistent across datasets but varying among classifiers. Data augmentations were particularly effective for smaller sample sizes (30 and below), significantly improving 70% of these classification analyses and only significantly impairing 4% of analyses. Moreover, EEG-GAN significantly outperformed all benchmark techniques in 69% of the comparisons across datasets, classifiers, and sample sizes and was only significantly outperformed in 3% of comparisons. These findings establish EEG-GAN as a robust toolkit for generating realistic EEG data, which can effectively reduce the costs associated with real-world EEG data collection for neural decoding tasks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".