A Novel Experience Sampling Method Tool Integrating Momentary Assessments of Cognitive Biases: Two Compliance, Usability, and Measurement Reactivity Studies
Bibliographic record
Abstract
BACKGROUND: Experience sampling methods (ESMs) are increasingly being used to study ecological emotion dynamics in daily functioning through repeated assessments taken over several days. However, most of these ESM approaches are only based on self-report assessments, and therefore, studies on the ecological trajectories of their underlying mechanisms are scarce (ie, cognitive biases) and require evaluation through experimental tasks. We developed a novel ESM tool that integrates self-report measures of emotion and emotion regulation with a previously validated app-based cognitive task that allows for the assessment of underlying mechanisms during daily functioning. OBJECTIVE: The objective of the study is to test this new tool and study its usability and the possible factors related to compliance with it in terms of latency and missing responses. Among the compliance predictors, we considered psychological and time-related variables, as well as usability, measurement reactivity, and participants' satisfaction with the tool. METHODS: We conducted 2 extensive ESM studies-study 1 (N=84; a total of 3 assessments per day for 5 days) and study 2 (N=135; a total of 3 assessments per day for 10 days). RESULTS: In both studies, participants found the tool highly usable (average usability score >81). By using mixed regression models, we found both common and specific results for the compliance predictors. In both study 1 and study 2, latency was significantly predicted by the day (P<.001 and P=.003, respectively). Participants showed slower responses to the notification as the days of the study progressed. In study 2 but not in study 1, latency was further predicted by individual differences in overload with the use of the app, and missing responses were accounted for by individual differences in stress reactivity to notifications (P=.04). Thus, by using a more extensive design, participants who experienced higher overload during the study were characterized by slower responses to notifications (P=.01), whereas those who experienced higher stress reactivity to the notification system were characterized by higher missing responses. CONCLUSIONS: The new tool had high levels of usability. Furthermore, the study of compliance is of enormous importance when implementing novel ESM methods, including app-based cognitive tasks. The main predictors of latency and missing responses found across studies, specifically when using extensive ESM protocols (study 2), are methodology-related variables. Future research that integrates cognitive tasks in ESM designs should take these results into consideration by performing accurate estimations of participants' response rates to facilitate the optimal quality of novel eHealth approaches, as in this study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".