Got Bots? Practical Recommendations to Protect Online Survey Data from Bot Attacks
Bibliographic record
Abstract
The Internet has been a popular source of data amongst academic researchers for many years, and for good reason. Online data collection is fast, provides access to hard-to-reach populations, and is often less expensive than in-lab recruitment. With these benefits also come risks, such as duplicate responses or participant inattention, which can significantly reduce data quality. Very recently, researchers have become aware of another concern associated with online data collection. Bots, also known as automatic survey-takers or fraudsters, have begun infiltrating scientific surveys, largely threatening the integrity of academic research conducted online. The aim of this paper is to warn researchers of the threat posed by bots and to highlight practical strategies that can be used to detect and prevent these bots. We first discuss strategies recommended in the literature that we implemented to identify bot responses from online survey data we collected in the past six months. We then share which strategies proved to be most and least effective in detecting bots. Finally, we discuss the implications of bot-generated data for the integrity of online research and the imminent future of bots in online data collection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.163 | 0.472 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.007 | 0.004 |
| Science and technology studies | 0.005 | 0.008 |
| Scholarly communication | 0.015 | 0.034 |
| Open science | 0.008 | 0.008 |
| Research integrity | 0.016 | 0.011 |
| Insufficient payload (model declined to judge) | 0.021 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".