The Accuracy of Readily Available Consumer-Grade Oxygen Saturation Monitors in Pediatric Patients
Bibliographic record
Abstract
BACKGROUND: Pulse oximetry measurement is ubiquitous in acute health care settings in high-income countries and is familiar to any parent whose child has been treated in such a setting. Oximeters for home use are readily available online and are incorporated in several smartphones and smartwatches. METHODS: We wished to determine how accurate are oximeters available online that are designated for adult and pediatric use, and the saturation monitor integrated in a smartphone, when used in children, compared to reference, hospital-grade oximeters. We evaluated a fingertip oximeter marketed for children purchased online; an adult fingertip oximeter purchased online; the oximeter integrated in a smartphone; and reference, hospital-grade oximeters. Participants were < 18 y of age. Bland-Altman charts were generated, and the estimated root mean square error (EA RMS ) was calculated. Rates of failure to obtain a measurement, relationship between device and time to successful measurement, relationship between age and time to successful measurement, and relationship between error (vs the reference device) and age were evaluated for each consumer-grade device. RESULTS: We measured S pO 2 in 74 children between 0.1–17.0 y of age. Subjects weighing < 30 kg had a median (interquartile range [IQR]) age of 2 (1.0 month–1.4 y) months, and subjects weighing ≥ 30 kg had a median (IQR) age of 14.3 (11.9–16.2) y. Readings could not be obtained in 7.5, 0, and 38.8% of subjects using the pediatric, adult, and smartphone oximeters, respectively. The time to successful reading had a modest negative correlation with age with the inexpensive adult and pediatric oximeters. The inexpensive pediatric oximeter had an overall negative bias, with a mean difference from the reference device of −4.5% (SD 7.9%) and an error that ranged from > 8% to < 33% the reference device. The EA RMS was 7.92%. The inexpensive adult oximeter demonstrated no obvious trend in error in the limited saturation range evaluated of 87–99%. The overall mean difference was −0.7% (SD 2.5%). EA RMS was 2.5%. The smartphone oximeter underestimated S pO 2 at saturations < 94% and overestimated S pO 2 for saturations > 94%. Saturations could read as much as > 4%, or < 17%, than the reference oximeter. The mean difference was −2.9% (SD 5.2%). EA RMS was 5.1%. CONCLUSIONS: Our findings suggest that the performance of consumer-grade devices varies considerably by both subject age and device. The pediatric fingertip device and smartphone application we tested are poorly suited for use in infants. The adult fingertip device we tested performed quite well in larger children with relatively normal oxygen saturations, and the pediatric fingertip device performed moderately well in subjects > 1 y of age who weighed < 30 kg. Given the vast number of devices available online and ever-changing technology, research to evaluate nonclinical oximeters will continue to be required.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".