O002 Machine learning applied to oximetry to detect paediatric sleep apnoea
Bibliographic record
Abstract
Abstract Introduction Polysomnography(PSG) is the gold-standard test for diagnosing paediatric obstructive sleep apnoea, but resource-intensiveness limits availability. Although oximetry testing is widely available, manual scoring methods such as McGill scoring have poor sensitivity, and decreased positive predictive value(PPV) for those with complex medical comorbidities. We hypothesised that computerised oximetry analysis could overcome these limitations. We developed and tested a novel support vector classifier(SVC) machine learning algorithm, with a component for discarding of likely movement/wake artefact, that classifies the overnight oximetry as being either low risk (predicted apnoea hypopnoea index[AHI] <5/hour) or high risk(≥5/hour). Methods Oxygen saturation(SpO2) and pulse rate(PR) data were extracted from 3476 PSGs performed at Queensland Children’s Hospital (QCH). 90% were selected for model training, with 10% reserved for final testing. A rules-based filter discarded periods of SpO2 < 60% or PR > 225bpm or < 40bpm. Thirteen SpO2/PR features were used for SVC model training; model performance was then evaluated on held-out data from QCH and other sources. Results Using oximetry data from the reserved 344 QCH PSGs, the algorithm showed a PPV of 93%, negative predictive value of 90%, specificity of 99%, and sensitivity of 65%. Using oximetry data from the Childhood Adenotonsillectomy Trial, Paediatric Adenotonsillectomy Trial for Snoring and British Columbia Children’s Hospital PSG datasets (n = 1215, n = 721 and n = 2895 respectively) model performance showed specificities of 99%, 98% and 84% respectively, with sensitivities of 35-62%. Conclusion A SVC algorithm can be used to classify oximetry into likely AHI ≥5 with a high degree of certainty, for children with and without complex comorbidities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".