Heart Murmur Detection in Phonocardiogram Data Leveraging Data Augmentation and Artificial Intelligence
Bibliographic record
Abstract
Background/Objectives: With a 17.9 million annual mortality rate, cardiovascular disease is the leading global cause of death. As such, early detection and disease diagnosis are critical for effective treatment and symptom management. Cardiac auscultation, the process of listening to the heartbeat, often provides the first indication of underlying cardiac conditions. This practice allows for the identification of heart murmurs caused by turbulent blood flow. In this exploratory research paper, we propose an AI model to streamline this process to improve diagnostic accuracy and efficiency. Methods: We utilized data from the 2022 George Moody PhysioNet Heart Sound Classification Challenge, comprising phonocardiogram recordings of individuals under 21 years of age in Northeast Brazil. Only patients who had recordings from all four heart valves were included in our dataset. Audio files were synchronized across all recordings and converted to Mel spectrograms before being passed into a pre-trained Vision Transformer, and finally a MiniROCKET model. Additionally, data augmentation was conducted on audio files and spectrograms to generate new data, extending our total sample size from 928 spectrograms to 14,848. Results: Compared to the existing methods in the literature, our model yielded significantly enhanced quality assessment metrics, including Weighted Accuracy, Sensitivity, and F-Score, and resulted in a fast evaluation speed of 0.02 s per patient. Conclusions: The implementation of our method for the detection of heart murmurs can supplement physician diagnosis and contribute to earlier detection of underlying cardiovascular conditions, fast diagnosis times, increased scalability, and enhanced adaptability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".