Exploring Methods to Mitigate Fraud in Web-Based Surveys: Multicase Study Analysis
Notice bibliographique
Résumé
BACKGROUND: Web-based surveys are a cost-effective technique to engage a large population of participants in research projects, including those who were previously difficult to reach due to geographic location, safety, and vulnerability. While web-based surveys have many advantages, they can be more susceptible to fraud, especially when a generic invitation link or a financial incentive is offered. There is a paucity of literature presenting experiences for mitigating this type of fraudulent study response, yet this important foundation is needed to inform the work of researchers and institutional review boards (IRBs) to support the collection of high-quality, appropriate data. OBJECTIVE: This study aims to analyze, compare, and contrast the range of strategies used to prevent, detect, and remove fraudulent responses by investigating 4 web-based surveys in Australia and Canada, each of which experienced fraudulent responses. METHODS: Our descriptive multiple case study presents 4 research projects from Australia and Canada that experienced survey fraud. These web-based surveys recruited patients of, or clinicians providing, family planning services. We describe each study's approach to preventing fraud (primary prevention; eg, CAPTCHA) and a screening protocol to detect fraudulent responses during data collection (secondary prevention). Once fraud was detected, each study team developed strategies to protect data integrity, in consultation with coinvestigators, ethics committees/ IRBs, and biostatisticians, to remove fraudulent respondents from the dataset (tertiary prevention). RESULTS: All studies recruited via a generic survey link and provided remuneration, which are common risk factors for fraud. Several studies also relied on social media for recruitment. All 4 studies implemented tertiary fraud detection strategies to identify and remove fraudulent responses and maintain data integrity (removing between 16% and 45% of respondents). Including personal identifiers during data collection provided 3 of the studies with a more robust option to identify and remove fraudulent respondents. Where personal identifiers could not be used (eg, to protect the identity of a vulnerable study population), investigators relied on a complex fraud detection algorithm verified by manual team review. CONCLUSIONS: Commonly used web-based anonymized survey methods, particularly those offering incentives for participation, are at substantial risk for fraud. Across these 4 studies, robust fraud detection methods were essential to ensure data reliability, with varying strategies, such as using personal identifiers, applied based on specific survey contexts. Fraud mitigation criteria explored in this multicase analysis can be adapted to other web-based surveys, survey topics, and populations. Implementing the fraud prevention and detection methods within survey design will assist researchers and IRBs in protecting data integrity. TRIAL REGISTRATION: Australian New Zealand Clinical Trials Registry ACTRN12622000655741; https://www.anzctr.org.au/Trial/Registration/TrialReview.aspx?id=383919 and ClinicalTrials.gov NCT05793944; https://clinicaltrials.gov/study/NCT05793944.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,551 | 0,251 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,003 | 0,005 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».