The particular case of conducting addiction intervention research on Mechanical Turk
Notice bibliographique
Résumé
As Mechanical Turk (MTurk) becomes a more popular platform for conducting addiction intervention studies, greater awareness of the limitations of this participant pool is needed. Specific challenges for testing interventions on MTurk include: limited intervention engagement, the sharing of crucial study information among MTurk workers and the collection of fraudulent responses. As more researchers turn to on-line convenience samples in an effort to reach hidden populations of substance users, reviews such as that by Mellis & Bickel [1] are both timely and necessary. Also, as the interest in developing efficacious on-line interventions for addictions grows [2, 3], Mechanical Turk (MTurk) presents a viable platform for the development and testing of these interventions due to its affordability and large participant pool. Nevertheless, its use in clinical research comes with challenges and limitations, such as non-naiveté, worker inattention and fraudulent responses [1, 4]. This commentary explores other challenges with conducting research on MTurk, including those specific to intervention testing, and expands upon what has been addressed by Mellis & Bickel [1]. One of the largest challenges that researchers face on MTurk is the collection of fraudulent respondes and duplicate participants. Indeed, this issue has been visited within the literature numerous times, and several strategies have been offered to minimize its occurrence (e.g. checking for common virtual private servers, response consistency, random responding and unique IP addresses) [1, 5, 6]. However, another facet of this problem is the sharing of valuable study information among MTurk workers on forums. This includes screenshots of human intelligence tasks (HITs), study eligibility criteria, embedded attention checks and survey links. Among all five intervention studies that our team has conducted on MTurk to date, we were able to find evidence of this information sharing on TurkOpticon, Reddit, MTurkCrowd, TurkView and mturkforum [7-10]. Therefore, it is recommended that researchers monitor these common websites regularly during recruitment. In addition, researchers who use survey link templates on MTurk can capture MTurk worker IDs in survey links using HTML code. The worker ID can then be captured by the survey software prior to data collection, and conditional logic can require a unique value for the survey to proceed. This can help to mitigate instances of participants completing a survey repeatedly to be found eligible, and cases where workers first complete the survey (via access to the survey link from another worker) and then attempt to find the HIT on MTurk for payment. Using this method, our team was able to identify nearly 5% of our eligible sample among four different trials as fraudulent or duplicate [11]. Another challenge specific to intervention research on MTurk is participant engagement. While this is an obstacle for the testing of any on-line intervention [12], our experience has been that more than one-time use of an intervention can be as low as 9% [9]. Although narrative intervention research using MTurk has found success in retaining participants’ attention through offering payment for intervention use [1], this is not feasible for all intervention studies. First, paying for intervention usage may not be ethical and can confound study results, particularly in the case of randomized controlled trials with a no-treatment control. In this case, only participants in the intervention group would be paid for use, and intervention usage may be artificially inflated. As intervention engagement has been found to be inseparable from intervention content, the target behaviour, and/or mechanisms of change within the program [13], paying participants for engagement can be problematic and perhaps overshadow important intervention weaknesses. Instead, we recommend that researchers find alternative ways of effectively engaging MTurk workers in testing interventions, such as recruiting workers who may be interested in receiving help for the behaviour being targeted, outlining the benefits of the intervention to the participant, and/or identifying the monetary value of access to the intervention. Additionally, researchers can provide greater compensations for providing feedback on the intervention which may improve engagement, rather than providing payment for interacting with the intervention itself. Despite the challenges of conducting intervention research on MTurk there continue to be advantages for its use, especially during the developmental stage of an intervention. Our inability to sometimes replicate the effectiveness of previously tested evidence-based interventions using MTurk participants suggests that this participant pool may not be appropriate for full intervention trials. However, it may be a fruitful platform for testing intervention components. Indeed, MTurk may be a good testing ground for manipulating small changes in interventions among multiple experimental groups. Regardless, it is imperative that researchers keep in mind that MTurk workers are mainly motivated to participate in HITs for pay, and not for the same intrinsic or altruistic reasons as typical research participants. Overall, MTurk can have a meaningful role in intervention research; however, researchers must pay special attention to the limitations of the platform when designing studies and reporting results. None.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,003 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».