MétaCan
Menu
← Back to cohort
Record W3212526843 · doi:10.2196/31618

Identifying Data Quality Dimensions for Person-Generated Wearable Device Data: Multi-Method Study

2021· article· en· W3212526843 on OpenAlexvenueno aff
Sylvia Cho, Chunhua Weng, Michael G. Kahn, Karthik Natarajan

Bibliographic record

VenueJMIR mhealth and uhealth · 2021
Typearticle
Languageen
FieldHealth Professions
TopicMobile Health and mHealth Applications
Canadian institutionsnot available
FundersNational Center for Advancing Translational SciencesNational Institutes of Health
KeywordsData qualityWearable computerComputer scienceWearable technologyFocus groupQuality (philosophy)Data scienceData collectionFacilitatorPsychologyEngineeringMetric (unit)

Abstract

fetched live from OpenAlex

BACKGROUND: There is a growing interest in using person-generated wearable device data for biomedical research, but there are also concerns regarding the quality of data such as missing or incorrect data. This emphasizes the importance of assessing data quality before conducting research. In order to perform data quality assessments, it is essential to define what data quality means for person-generated wearable device data by identifying the data quality dimensions. OBJECTIVE: This study aims to identify data quality dimensions for person-generated wearable device data for research purposes. METHODS: This study was conducted in 3 phases: literature review, survey, and focus group discussion. The literature review was conducted following the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guideline to identify factors affecting data quality and its associated data quality challenges. In addition, we conducted a survey to confirm and complement results from the literature review and to understand researchers' perceptions on data quality dimensions that were previously identified as dimensions for the secondary use of electronic health record (EHR) data. We sent the survey to researchers with experience in analyzing wearable device data. Focus group discussion sessions were conducted with domain experts to derive data quality dimensions for person-generated wearable device data. On the basis of the results from the literature review and survey, a facilitator proposed potential data quality dimensions relevant to person-generated wearable device data, and the domain experts accepted or rejected the suggested dimensions. RESULTS: In total, 19 studies were included in the literature review, and 3 major themes emerged: device- and technical-related, user-related, and data governance-related factors. The associated data quality problems were incomplete data, incorrect data, and heterogeneous data. A total of 20 respondents answered the survey. The major data quality challenges faced by researchers were completeness, accuracy, and plausibility. The importance ratings on data quality dimensions in an existing framework showed that the dimensions for secondary use of EHR data are applicable to person-generated wearable device data. There were 3 focus group sessions with domain experts in data quality and wearable device research. The experts concluded that intrinsic data quality features, such as conformance, completeness, and plausibility, and contextual and fitness-for-use data quality features, such as completeness (breadth and density) and temporal data granularity, are important data quality dimensions for assessing person-generated wearable device data for research purposes. CONCLUSIONS: In this study, intrinsic and contextual and fitness-for-use data quality dimensions for person-generated wearable device data were identified. The dimensions were adapted from data quality terminologies and frameworks for the secondary use of EHR data with a few modifications. Further research on how data quality can be assessed with respect to each dimension is needed.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.245
metaresearch head score (Gemma)0.369
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.755
Threshold uncertainty score0.931

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2450.369
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0050.009
Bibliometrics0.0120.014
Science and technology studies0.0020.003
Scholarly communication0.0060.006
Open science0.0030.005
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.578
GPT teacher head0.618
Teacher spread0.040 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designQualitative
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations26
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueJMIR mhealth and uhealth→Same topicMobile Health and mHealth Applications→French-language works237,207→