Measuring Criterion Validity of Microinteraction Ecological Momentary Assessment (Micro-EMA): Exploratory Pilot Study With Physical Activity Measurement
Bibliographic record
Abstract
BACKGROUND: Ecological momentary assessment (EMA) is an in situ method of gathering self-report on behaviors using mobile devices. In typical phone-based EMAs, participants are prompted repeatedly with multiple-choice questions, often causing participation burden. Alternatively, microinteraction EMA (micro-EMA or μEMA) is a type of EMA where all the self-report prompts are single-question surveys that can be answered using a 1-tap glanceable microinteraction conveniently on a smartwatch. Prior work suggests that μEMA may permit a substantially higher prompting rate than EMA, yielding higher response rates and lower participation burden. This is achieved by ensuring μEMA prompt questions are quick and cognitively simple to answer. However, the validity of participant responses from μEMA self-report has not yet been formally assessed. OBJECTIVE: In this pilot study, we explored the criterion validity of μEMA self-report on a smartwatch, using physical activity (PA) assessment as an example behavior of interest. METHODS: A total of 17 participants answered 72 μEMA prompts each day for 1 week using a custom-built μEMA smartwatch app. At each prompt, they self-reported whether they were doing sedentary, light/standing, moderate/walking, or vigorous activities by tapping on the smartwatch screen. Responses were compared with a research-grade activity monitor worn on the dominant ankle simultaneously (and continuously) measuring PA. RESULTS: Participants had an 87.01% (5226/6006) μEMA completion rate and a 74.00% (5226/7062) compliance rate taking an average of only 5.4 (SD 1.5) seconds to answer a prompt. When comparing μEMA responses with the activity monitor, we observed significantly higher (P<.001) momentary PA levels on the activity monitor when participants self-reported engaging in moderate+vigorous activities compared with sedentary or light/standing activities. The same comparison did not yield any significant differences in momentary PA levels as recorded by the activity monitor when the μEMA responses were randomly generated (ie, simulating careless taps on the smartwatch). CONCLUSIONS: For PA measurement, high-frequency μEMA self-report could be used to capture information that appears consistent with that of a research-grade continuous sensor for sedentary, light, and moderate+vigorous activity, suggesting criterion validity. The preliminary results show that participants were not carelessly answering μEMA prompts by randomly tapping on the smartwatch but were reporting their true behavior at that moment. However, more research is needed to examine the criterion validity of μEMA when measuring vigorous activities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.035 | 0.053 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".