The particular case of conducting addiction intervention research on Mechanical Turk
Bibliographic record
Abstract
As Mechanical Turk (MTurk) becomes a more popular platform for conducting addiction intervention studies, greater awareness of the limitations of this participant pool is needed. Specific challenges for testing interventions on MTurk include: limited intervention engagement, the sharing of crucial study information among MTurk workers and the collection of fraudulent responses. As more researchers turn to on-line convenience samples in an effort to reach hidden populations of substance users, reviews such as that by Mellis & Bickel [1] are both timely and necessary. Also, as the interest in developing efficacious on-line interventions for addictions grows [2, 3], Mechanical Turk (MTurk) presents a viable platform for the development and testing of these interventions due to its affordability and large participant pool. Nevertheless, its use in clinical research comes with challenges and limitations, such as non-naiveté, worker inattention and fraudulent responses [1, 4]. This commentary explores other challenges with conducting research on MTurk, including those specific to intervention testing, and expands upon what has been addressed by Mellis & Bickel [1]. One of the largest challenges that researchers face on MTurk is the collection of fraudulent respondes and duplicate participants. Indeed, this issue has been visited within the literature numerous times, and several strategies have been offered to minimize its occurrence (e.g. checking for common virtual private servers, response consistency, random responding and unique IP addresses) [1, 5, 6]. However, another facet of this problem is the sharing of valuable study information among MTurk workers on forums. This includes screenshots of human intelligence tasks (HITs), study eligibility criteria, embedded attention checks and survey links. Among all five intervention studies that our team has conducted on MTurk to date, we were able to find evidence of this information sharing on TurkOpticon, Reddit, MTurkCrowd, TurkView and mturkforum [7-10]. Therefore, it is recommended that researchers monitor these common websites regularly during recruitment. In addition, researchers who use survey link templates on MTurk can capture MTurk worker IDs in survey links using HTML code. The worker ID can then be captured by the survey software prior to data collection, and conditional logic can require a unique value for the survey to proceed. This can help to mitigate instances of participants completing a survey repeatedly to be found eligible, and cases where workers first complete the survey (via access to the survey link from another worker) and then attempt to find the HIT on MTurk for payment. Using this method, our team was able to identify nearly 5% of our eligible sample among four different trials as fraudulent or duplicate [11]. Another challenge specific to intervention research on MTurk is participant engagement. While this is an obstacle for the testing of any on-line intervention [12], our experience has been that more than one-time use of an intervention can be as low as 9% [9]. Although narrative intervention research using MTurk has found success in retaining participants’ attention through offering payment for intervention use [1], this is not feasible for all intervention studies. First, paying for intervention usage may not be ethical and can confound study results, particularly in the case of randomized controlled trials with a no-treatment control. In this case, only participants in the intervention group would be paid for use, and intervention usage may be artificially inflated. As intervention engagement has been found to be inseparable from intervention content, the target behaviour, and/or mechanisms of change within the program [13], paying participants for engagement can be problematic and perhaps overshadow important intervention weaknesses. Instead, we recommend that researchers find alternative ways of effectively engaging MTurk workers in testing interventions, such as recruiting workers who may be interested in receiving help for the behaviour being targeted, outlining the benefits of the intervention to the participant, and/or identifying the monetary value of access to the intervention. Additionally, researchers can provide greater compensations for providing feedback on the intervention which may improve engagement, rather than providing payment for interacting with the intervention itself. Despite the challenges of conducting intervention research on MTurk there continue to be advantages for its use, especially during the developmental stage of an intervention. Our inability to sometimes replicate the effectiveness of previously tested evidence-based interventions using MTurk participants suggests that this participant pool may not be appropriate for full intervention trials. However, it may be a fruitful platform for testing intervention components. Indeed, MTurk may be a good testing ground for manipulating small changes in interventions among multiple experimental groups. Regardless, it is imperative that researchers keep in mind that MTurk workers are mainly motivated to participate in HITs for pay, and not for the same intrinsic or altruistic reasons as typical research participants. Overall, MTurk can have a meaningful role in intervention research; however, researchers must pay special attention to the limitations of the platform when designing studies and reporting results. None.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".