Establishing Linkages Between Distributed Survey Responses and Consumer Wearable Device Datasets: A Pilot Protocol
Bibliographic record
Abstract
BACKGROUND: As technology increasingly becomes an integral part of everyday life, many individuals are choosing to use wearable technology such as activity trackers to monitor their daily physical activity and other health-related goals. Researchers would benefit from learning more about the health of these individuals remotely, without meeting face-to-face with participants and avoiding the high cost of providing consumer wearables to participants for the study duration. OBJECTIVE: The present study seeks to develop the methods to collect data remotely and establish a linkage between self-reported survey responses and consumer wearable device biometric data, ultimately producing a de-identified and linked dataset. Establishing an effective protocol will allow for future studies of large-scale deployment and participant management. METHODS: A total of 30 participants who use a Fitbit will be recruited on Mechanical Turk Prime and asked to complete a short online self-administered questionnaire. They will also be asked to connect their personal Fitbit activity tracker to an online third-party software system, called Fitabase, which will allow access to 1 month's retrospective data and 1 month's prospective data, both from the date of consent. RESULTS: The protocol will be used to create and refine methods to establish linkages between remotely sourced and de-identified survey responses on health status and consumer wearable device data. CONCLUSIONS: The refinement of the protocol will inform collection and linkage of similar datasets at scale, enabling the integration of consumer wearable device data collection in cross-sectional and prospective cohort studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.230 | 0.218 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.006 | 0.004 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.004 | 0.006 |
| Insufficient payload (model declined to judge) | 0.048 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".