MétaCan
Menu
← Back to cohort
Record W7164532926 · doi:10.2196/87324

The representation of different populations in studies assessing the validity of consumer wearable PPG-based measurements: scoping review (Preprint)

2025· article· en· W7164532926 on OpenAlexvenueno aff
Rebecca M. Schipper, Fatime Oumar Djibrillah, Meyke Roosink, Laura H.H. Winkens, Arlene John, Eric J. Hazebroek, Annemieke Witteveen, Agnes A. M. Berendsen

Bibliographic record

VenueJMIR mhealth and uhealth · 2025
Typearticle
Languageen
FieldMedicine
TopicPhysical Activity and Health
Canadian institutionsnot available
Fundersnot available
KeywordsWearable computerRepresentation (politics)Wearable technologymHealthContext (archaeology)Identification (biology)

Abstract

fetched live from OpenAlex

Abstract Background Consumer wearables are increasingly being integrated into health research for data collection. Although they are attractive to use, the accuracy of their photoplethysmography (PPG)-based measurements can be influenced by user characteristics such as sex, age, BMI, and skin tone. However, our knowledge regarding the validity of these measurements in certain populations, such as those with darker skin tones, seems limited. This is concerning as uncorrected differences in measurement accuracy can lead to health disparities when consumer wearable measurements are used more frequently. A potential cause for the gap in our knowledge regarding consumer wearable validity is the underrepresentation of certain population groups in studies validating PPG-based consumer wearables. Objective This scoping review aimed to map the representation of different sex, age, BMI, and skin tone groups in studies assessing the validity of PPG-based pulse rate, heart rate variability, blood pressure, peripheral blood oxygen saturation (SpO 2 ), and respiratory rate measurements of consumer wearables. Methods A literature search was conducted in Scopus, PubMed, and IEEE Xplore in July 2025. Papers were eligible if they assessed the validity of consumer wearable PPG-based measurements, expressed as the agreement with a reference method. From the included papers, the study population distribution of sex, age, BMI, and Fitzpatrick scale was extracted. To evaluate the representation, percentages of people in specific age, BMI, and skin tone groups were estimated based on reported means and SDs. The median percentage of participants in each population group, as well as the total percentage, is reported. Results After the removal of duplicates, 734 papers were screened for eligibility. Following title and abstract screening, 238 papers remained, of which 186 passed full-text screening and were included in the review. Most of the studies (n=160) focused on pulse rate. Sex, age, BMI, and Fitzpatrick scale were reported by 179 (96.0%), 178 (96.0%), 101 (54.0%), and 35 (19.0%) out of 186 studies, respectively. While the median representation was 0% (IQR 0%-8%) for both older adults (>65 y) and individuals with obesity (BMI>30 kg/m 2 ; IQR 0%-13%), aggregate participation across all studies was higher (1290/6367, 20.0% and 473/3428, 14.0%, respectively). Individuals with underweight (BMI<18.5 kg/m 2 ) remained rare (median 3%, IQR 0%-7%), and the aggregate was 7.0% (225/3428). The median percentage of people with darker skin tones (Fitzpatrick type V and VI) participating in a study was 0%. Conclusions Based on our results, it can be concluded that older adults and people with underweight, obesity, or darker skin tones are generally underrepresented in studies assessing the validity of consumer wearable PPG-based measurements. Future validation studies should focus more on the representativeness of the study population. This can be achieved by setting a benchmark for representativeness and including study population representatives during the study design process.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.143
metaresearch head score (Gemma)0.445
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Review · Consensus signal: Review
Teacher disagreement score0.857
Threshold uncertainty score0.756

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1430.445
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0080.010
Bibliometrics0.0160.014
Science and technology studies0.0010.005
Scholarly communication0.0090.008
Open science0.0040.005
Research integrity0.0060.003
Insufficient payload (model declined to judge)0.0030.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.462
GPT teacher head0.558
Teacher spread0.096 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSystematic review
DomainMethods
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractno

Explore more

Same venueJMIR mhealth and uhealth→Same topicPhysical Activity and Health→French-language works237,207→