Wrist02 -- Reliable Peripheral Oxygen Saturation Readings from Wrist-Worn Pulse Oximeters
Bibliographic record
Abstract
Peripheral blood oxygen saturation Sp02 is a vital measure in healthcare. Modern off-the-shelf wrist-worn devices, such as the Apple Watch, FitBit, and Samsung Gear, have an onboard sensor called a pulse oximeter. While pulse oximeters are capable of measuring both Sp02 and heart rate, current wrist-worn devices use them only to determine heart rate, as Sp02 measurements collected from the wrist are believed to be inaccurate. Enabling oxygen saturation monitoring on wearable devices would make these devices tremendously more useful for health monitoring and open up new avenues of research. To the best of our knowledge, we present the first study of the reliability of Sp02 sensing from the wrist. Using a custom-built wrist-worn pulse oximeter, we find that existing algorithms designed for fingertip sensing are a poor match for this setting, and can lead to over 90% of readings being inaccurate and unusable. We further show that sensor placement and skin tone have a substantial effect on the measurement error, and must be considered when designing wrist-worn Sp02 sensors and measurement algorithms. Based on our findings, we propose \codename, an alternative approach for reliable Sp02 sensing. By selectively pruning data, \codename achieves an order of magnitude reduction in error compared to existing algorithms, while still providing sufficiently frequent readings for continuous health monitoring.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".