Evaluation of Resampling Techniques in CNN-Based Heartbeat Classification
Bibliographic record
Abstract
This study investigates the efficacy of resampling techniques in ECG classification, addressing the challenge of data imbalance in heartbeat classification.Utilizing the PTB Diagnostic ECG database, the research focuses on the application of various Synthetic Minority Over-sampling Technique (SMOTE) variations, including SMOTE Borderline, ADASYN, Tomek, and ENN, alongside three algorithms: CNN, Transformer, and LSTM.The dataset, encompassing 549 patient records from 290 subjects, was bifurcated into training and testing segments, classifying heartbeats into normal and abnormal categories.The novelty of this work lies in its combined deep-structured learning model that integrates CNN, Transformer, and LSTM, further enhanced by an ensemble of these algorithms with original SMOTE and its variants for dataset balancing.The research revealed that the proposed method significantly ameliorates the classification of heartbeats, effectively addressing the class imbalance issue prevalent in ECG data.The results demonstrated that the transformer network, in particular, excelled in recognizing temporal continuities and extracting deep-seated features from ECG signals, thereby enhancing the model's performance beyond the capabilities of basic models.Key results indicate that CNN+SMOTE Borderline achieves the highest testing accuracy at 99.36%, while CNN+SMOTE Tomek leads in precision with 99.89%.Transformers excel in recall with a perfect score of 100%.The research concludes that CNNs effectively distinguish normal from abnormal heartbeats, with the highest accuracy using CNN+SMOTE at 99.06%.However, the study also acknowledges limitations, such as the dataset's restricted scope, and suggests further research with a more diverse dataset.Overall, the study demonstrates the effectiveness of CNN in ECG arrhythmia classification, offering a foundation for more advanced automatic diagnostic systems in cardiology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".