Exploring Methods to Mitigate Fraud in Web-Based Surveys: Multicase Study Analysis
Bibliographic record
Abstract
BACKGROUND: Web-based surveys are a cost-effective technique to engage a large population of participants in research projects, including those who were previously difficult to reach due to geographic location, safety, and vulnerability. While web-based surveys have many advantages, they can be more susceptible to fraud, especially when a generic invitation link or a financial incentive is offered. There is a paucity of literature presenting experiences for mitigating this type of fraudulent study response, yet this important foundation is needed to inform the work of researchers and institutional review boards (IRBs) to support the collection of high-quality, appropriate data. OBJECTIVE: This study aims to analyze, compare, and contrast the range of strategies used to prevent, detect, and remove fraudulent responses by investigating 4 web-based surveys in Australia and Canada, each of which experienced fraudulent responses. METHODS: Our descriptive multiple case study presents 4 research projects from Australia and Canada that experienced survey fraud. These web-based surveys recruited patients of, or clinicians providing, family planning services. We describe each study's approach to preventing fraud (primary prevention; eg, CAPTCHA) and a screening protocol to detect fraudulent responses during data collection (secondary prevention). Once fraud was detected, each study team developed strategies to protect data integrity, in consultation with coinvestigators, ethics committees/ IRBs, and biostatisticians, to remove fraudulent respondents from the dataset (tertiary prevention). RESULTS: All studies recruited via a generic survey link and provided remuneration, which are common risk factors for fraud. Several studies also relied on social media for recruitment. All 4 studies implemented tertiary fraud detection strategies to identify and remove fraudulent responses and maintain data integrity (removing between 16% and 45% of respondents). Including personal identifiers during data collection provided 3 of the studies with a more robust option to identify and remove fraudulent respondents. Where personal identifiers could not be used (eg, to protect the identity of a vulnerable study population), investigators relied on a complex fraud detection algorithm verified by manual team review. CONCLUSIONS: Commonly used web-based anonymized survey methods, particularly those offering incentives for participation, are at substantial risk for fraud. Across these 4 studies, robust fraud detection methods were essential to ensure data reliability, with varying strategies, such as using personal identifiers, applied based on specific survey contexts. Fraud mitigation criteria explored in this multicase analysis can be adapted to other web-based surveys, survey topics, and populations. Implementing the fraud prevention and detection methods within survey design will assist researchers and IRBs in protecting data integrity. TRIAL REGISTRATION: Australian New Zealand Clinical Trials Registry ACTRN12622000655741; https://www.anzctr.org.au/Trial/Registration/TrialReview.aspx?id=383919 and ClinicalTrials.gov NCT05793944; https://clinicaltrials.gov/study/NCT05793944.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.551 | 0.251 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".