MétaCan
Menu
Back to cohort
Record W4416868232 · doi:10.2196/78671

Exploring Methods to Mitigate Fraud in Web-Based Surveys: Multicase Study Analysis

2025· article· en· W4416868232 on OpenAlexafffundabout
Madeleine Ennis, Regina Renner, Claudia Morando-Stokoe, Sharon James, Patricia A. Janssen, Sara Leckie, Sheila Dunn, Danielle Mazza, Wendy V. Norman

Bibliographic record

VenueJournal of Medical Internet Research · 2025
Typearticle
Languageen
FieldSocial Sciences
TopicSurvey Methodology and Nonresponse
Canadian institutionsWomen's College HospitalUniversity of TorontoUniversity of British Columbia
FundersNational Health and Medical Research CouncilMedical Research CouncilPublic Health AgencyCanadian Institutes of Health ResearchRoyal Australian College of General PractitionersAustralian GovernmentPublic Health Agency of CanadaAustralian Commission on Safety and Quality in Health Care
KeywordsData collectionMEDLINEReliability (semiconductor)Risk assessmentPublic healthContext (archaeology)

Abstract

fetched live from OpenAlex

BACKGROUND: Web-based surveys are a cost-effective technique to engage a large population of participants in research projects, including those who were previously difficult to reach due to geographic location, safety, and vulnerability. While web-based surveys have many advantages, they can be more susceptible to fraud, especially when a generic invitation link or a financial incentive is offered. There is a paucity of literature presenting experiences for mitigating this type of fraudulent study response, yet this important foundation is needed to inform the work of researchers and institutional review boards (IRBs) to support the collection of high-quality, appropriate data. OBJECTIVE: This study aims to analyze, compare, and contrast the range of strategies used to prevent, detect, and remove fraudulent responses by investigating 4 web-based surveys in Australia and Canada, each of which experienced fraudulent responses. METHODS: Our descriptive multiple case study presents 4 research projects from Australia and Canada that experienced survey fraud. These web-based surveys recruited patients of, or clinicians providing, family planning services. We describe each study's approach to preventing fraud (primary prevention; eg, CAPTCHA) and a screening protocol to detect fraudulent responses during data collection (secondary prevention). Once fraud was detected, each study team developed strategies to protect data integrity, in consultation with coinvestigators, ethics committees/ IRBs, and biostatisticians, to remove fraudulent respondents from the dataset (tertiary prevention). RESULTS: All studies recruited via a generic survey link and provided remuneration, which are common risk factors for fraud. Several studies also relied on social media for recruitment. All 4 studies implemented tertiary fraud detection strategies to identify and remove fraudulent responses and maintain data integrity (removing between 16% and 45% of respondents). Including personal identifiers during data collection provided 3 of the studies with a more robust option to identify and remove fraudulent respondents. Where personal identifiers could not be used (eg, to protect the identity of a vulnerable study population), investigators relied on a complex fraud detection algorithm verified by manual team review. CONCLUSIONS: Commonly used web-based anonymized survey methods, particularly those offering incentives for participation, are at substantial risk for fraud. Across these 4 studies, robust fraud detection methods were essential to ensure data reliability, with varying strategies, such as using personal identifiers, applied based on specific survey contexts. Fraud mitigation criteria explored in this multicase analysis can be adapted to other web-based surveys, survey topics, and populations. Implementing the fraud prevention and detection methods within survey design will assist researchers and IRBs in protecting data integrity. TRIAL REGISTRATION: Australian New Zealand Clinical Trials Registry ACTRN12622000655741; https://www.anzctr.org.au/Trial/Registration/TrialReview.aspx?id=383919 and ClinicalTrials.gov NCT05793944; https://clinicaltrials.gov/study/NCT05793944.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.191
metaresearch head score (Gemma)0.408
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Research integrity
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.999
Threshold uncertainty score0.997

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1910.408
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.004
Bibliometrics0.0150.016
Science and technology studies0.0020.002
Scholarly communication0.0040.004
Open science0.0020.004
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0050.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.687
GPT teacher head0.659
Teacher spread0.028 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designQualitative
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes3
Has abstractyes

Explore more

Same venueJournal of Medical Internet ResearchSame topicSurvey Methodology and NonresponseFrench-language works237,207