Comparison of Silence Removal Methods for the Identification of Audio Cough Events
Bibliographic record
Abstract
Sensing technologies are embedded in our everyday lives. Smart homes typically use an Audio Virtual Assistant (AVA) (e.g. Alexa, Siri, and Google Home) interface that collects sensor information, which can provide security, assist in everyday activities and monitor health related information. One such measure is cough, changes of which can be a marker of worsening conditions for many respiratory diseases. Creating a reliable monitoring system utilizing technology that may already be present in the home (i.e. AVA) may provide an opportunity for early intervention and reductions in the number of long-term hospitalizations. This paper focuses on the optimization of the silence removal and segmentation step in an at home setting with low to moderate background noise to identify cough events. Three commonly used methods (Standard deviation (SD), Short-term Energy (SE), Zero-crossing rate (ZCR)) were compared to manual segmentations. Each method was applied to 209 audio files that were manually verified to contain at least one cough event and the average segmentation accuracy, over segmentation and under segmentation results were compared. The ZCR method had the highest accuracy (89%); however, it completely failed under moderate noise conditions. The SD method had the best combination of accuracy (86%), ability to perform under noisy conditions and low prevalence of over and under segmentation (22% and 15% respectively). Therefore, we recommend using an adaptive approach to silence removal among cough events based on the level of background noise (i.e use the ZCR method when the background noise is low and the SD method when it is higher) prior to implementation of a cough classification system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".