Variability of sleep stage scoring in late midlife and early old age
Bibliographic record
Abstract
Sleep stage scoring can lead to important inter-expert variability. Although likely, whether this issue is amplified in older populations, which show alterations of sleep electrophysiology, has not been thoroughly assessed. Algorithms for automatic sleep stage scoring may appear ideal to eliminate inter-expert variability. Yet, variability between human experts and algorithm sleep stage scoring in healthy older individuals has not been investigated. Here, we aimed to compare stage scoring of older individuals and hypothesized that variability, whether between experts or considering the algorithm, would be higher than usually reported in the literature. Twenty cognitively normal and healthy late midlife individuals' (61 ± 5 years; 10 women) night-time sleep recordings were scored by two experts from different research centres and one algorithm. We computed agreements for the entire night (percentage and Cohen's κ) and each sleep stage. Whole-night pairwise agreements were relatively low and ranged from 67% to 78% (κ, 0.54-0.67). Sensitivity across pairs of scorers proved lowest for stages N1 (8.2%-63.4%) and N3 (44.8%-99.3%). Significant differences between experts and/or algorithm were found for total sleep time, sleep efficiency, time spent in N1/N2/N3 and wake after sleep onset (p ≤ 0.005), but not for sleep onset latency, rapid eye movement (REM) and slow-wave sleep (SWS) duration (N2 + N3). Our results confirm high inter-expert variability in healthy aging. Consensus appears good for REM and SWS, considered as a whole. It seems more difficult for N3, potentially because human raters adapt their interpretation according to overall changes in sleep characteristics. Although the algorithm does not substantially reduce variability, it would favour time-efficient standardization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.010 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".