A Multi-task Deep Learning Algorithm for Sleep Stage Scoring and Sleep Arousal Detection
Bibliographic record
Abstract
Sleep is crucial for overall health and well-being. Analyzing sleep stages and the frequency of arousal can enhance our understanding of sleep quality and help protect individuals’ sleep health. The objective of this study is to explore the development and application of multi-task deep learning model that can simultaneously detect sleep arousal and score sleep stages. We employed a task-incremental approach to develop the multi-task deep learning model. For development and testing, 1069 polysomnography records were incorporated. Additionally, the model’s fairness was assessed across various subgroups, including sex, age, and ethnicity. Beyond performance evaluation, we investigated the intermediate features extracted by the model from the raw ECG signal. The multi-task deep learning model achieved a Cohen’s 𝜅 of 0.68 for sleep stage prediction and an area under the receiver operating characteristic (AUROC) of 0.94 for arousal detection. Additionally, the model showed consistent performance across different subgroups. The analysis of the interpretability of the intermediate features validated their applicability and adaptability to related tasks in sleep analysis. This study demonstrated the capacity of the multi-task deep learning model to score sleep stages and arousal simultaneously with substantial precision using a single-lead ECG. This research offers a promising avenue for advancing incremental learning and multi-task deep learning models in sleep analysis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".