The representation of different populations in studies assessing the validity of consumer wearable PPG-based measurements: scoping review (Preprint)
Notice bibliographique
Résumé
Abstract Background Consumer wearables are increasingly being integrated into health research for data collection. Although they are attractive to use, the accuracy of their photoplethysmography (PPG)-based measurements can be influenced by user characteristics such as sex, age, BMI, and skin tone. However, our knowledge regarding the validity of these measurements in certain populations, such as those with darker skin tones, seems limited. This is concerning as uncorrected differences in measurement accuracy can lead to health disparities when consumer wearable measurements are used more frequently. A potential cause for the gap in our knowledge regarding consumer wearable validity is the underrepresentation of certain population groups in studies validating PPG-based consumer wearables. Objective This scoping review aimed to map the representation of different sex, age, BMI, and skin tone groups in studies assessing the validity of PPG-based pulse rate, heart rate variability, blood pressure, peripheral blood oxygen saturation (SpO 2 ), and respiratory rate measurements of consumer wearables. Methods A literature search was conducted in Scopus, PubMed, and IEEE Xplore in July 2025. Papers were eligible if they assessed the validity of consumer wearable PPG-based measurements, expressed as the agreement with a reference method. From the included papers, the study population distribution of sex, age, BMI, and Fitzpatrick scale was extracted. To evaluate the representation, percentages of people in specific age, BMI, and skin tone groups were estimated based on reported means and SDs. The median percentage of participants in each population group, as well as the total percentage, is reported. Results After the removal of duplicates, 734 papers were screened for eligibility. Following title and abstract screening, 238 papers remained, of which 186 passed full-text screening and were included in the review. Most of the studies (n=160) focused on pulse rate. Sex, age, BMI, and Fitzpatrick scale were reported by 179 (96.0%), 178 (96.0%), 101 (54.0%), and 35 (19.0%) out of 186 studies, respectively. While the median representation was 0% (IQR 0%-8%) for both older adults (>65 y) and individuals with obesity (BMI>30 kg/m 2 ; IQR 0%-13%), aggregate participation across all studies was higher (1290/6367, 20.0% and 473/3428, 14.0%, respectively). Individuals with underweight (BMI<18.5 kg/m 2 ) remained rare (median 3%, IQR 0%-7%), and the aggregate was 7.0% (225/3428). The median percentage of people with darker skin tones (Fitzpatrick type V and VI) participating in a study was 0%. Conclusions Based on our results, it can be concluded that older adults and people with underweight, obesity, or darker skin tones are generally underrepresented in studies assessing the validity of consumer wearable PPG-based measurements. Future validation studies should focus more on the representativeness of the study population. This can be achieved by setting a benchmark for representativeness and including study population representatives during the study design process.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,143 | 0,445 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,008 | 0,010 |
| Bibliométrie | 0,016 | 0,014 |
| Études des sciences et des technologies | 0,001 | 0,005 |
| Communication savante | 0,009 | 0,008 |
| Science ouverte | 0,004 | 0,005 |
| Intégrité de la recherche | 0,006 | 0,003 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,003 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».