0547 Predicting Response to Oral Appliance Therapy: Artificial Intelligence Versus Intuitive Approach
Bibliographic record
Abstract
We have developed a method to predict outcome with oral appliance therapy (OAT) using an artificial intelligence (AI) approach combined with a feedback controlled mandibular positioner (FCMP). We compared the performance of our method against an intuitive approach based on a one-night study used for prediction. Participants (n=101) with OSA (mean AHI: 30.4 hr-1, mean ODI: 30.6 hr-1, mean BMI: 32.2 kg/m2) received a 2- to 3-night FCMP test and were then treated by OAT. Their response to OAT was predicted by an AI-based approach embedded in MATRx plus (Zephyr Sleep Technologies). This was compared with an intuitive prediction where the ODI from a single night study, i.e., the 2nd night in the FCMP test, with a temporary oral appliance in place was used to predict the responders (ODI < 10 hr-1). Both predictions were compared to the efficacious response (ODI < 10 hr-1) in an outcome study with a custom OA in place. The sensitivity (Sens), specificity (Spec), positive and negative predictive values (PPV & NPV) were calculated for each of the two prediction methods. The mean ODI from the one-night study (18.1 hr-1) exceeded the mean outcome ODI (12.2 hr-1). Prediction using the intuitive approach resulted in Sens=68%, Spec=90%, PPV=96%, NPV=49%; while the AI-based method performance yielded Sens=88%, Spec=92%, PPV=97%, NPV=72%. Overall accuracy was 89% for AI and 74% for the intuitive approach. The sensitivity and overall predictive accuracy of the AI-based approach was greater than the intuitive approach, indicating that FCMP test outperformed the intuitive approach. Although it seems counter-intuitive, ODI from a one-night study is not a good estimate of the outcome ODI. Zephyr Sleep Technologies, Prosomnus Sleep Technologies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".