Automated Multi-Class Seizure-Type Classification System Using EEG Signals and Machine Learning Algorithms
Bibliographic record
Abstract
Epilepsy is a chronic brain disorder characterized by recurrent unprovoked seizures. The treatment for epilepsy is influenced by the types of seizures. Therefore, developing a reliable, explainable, and automated system to identify seizure types is necessary. This study aims to automate the process of classification of five seizure types: focal non-specific, generalized, complex partial, absence, and tonic-clonic using electroencephalogram (EEG) signals and machine learning algorithms. The EEG signals of 2933 seizures from 327 patients were obtained from the publicly available Temple University Hospital dataset. Initially, the signals were preprocessed using a standard pipeline, and 110 features from the time, frequency, and time-frequency domain were computed from each seizure. Further, the features were ranked using the statistical test and extreme Gradient Boosting (XGBoost) algorithm to identify the significant features. We built binary and multiclass seizure-type classification systems using the identified features and machine learning algorithms. Our study revealed that the EEG band power between 11–13 Hz, 27–29 Hz, intrinsic mode function (IMF) band power 19–21 Hz, and delta band (1-4 Hz) played a crucial role in discriminating the seizures. We achieved an average accuracy of 88.21% and 69.43% for the binary and multiclass seizure-type classification, respectively, using the XGBoost classifier. We also found that the combination of features performed well compared to any single domain. This automated system has the potential to aid neurologists in making diagnosis of epileptic seizure types. The proposed methodology can be applied alongside the established clinical approach of visual evaluation for the classification of seizure-types.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".