Bibliographic record
Abstract
Background: There is a recent and dangerous threat to online research and data collection: bots and bad actors. Bots (artificial intelligence (AI)) and bad actors (human participants who do not truthfully complete surveys) are growing in their strength and numbers when it comes to infiltrating studies. In the field of autism research, these present additional barriers to an already underrepresented population. As these fraudulent participants continue to evolve past current methods which aim to combat them, it is imperative researchers consider all options of prevention, detection, and elimination, without compromising their survey’s integrity or increasing participation barriers for valid participants. Objective: To synthesize current reviews of bot and bad actor prevention, detection, and elimination from online research surveys. Methods: A search of major databases was conducted for independent studies and reviews, resulting in 19 papers found to be most applicable. The literature was then summarized by what methods of bot and bad actor prevention, detection, and elimination were explored, how authors used each method, and the effectiveness of these. Results: The search resulted in 19 articles, 12 of which were independent studies which explained authors firsthand experiences dealing with bots and bad actors [1-12]. The remaining 7 were reviews which assessed common strategies for bot/bad actor prevention, detection, and elimination [13-19]. Of the independent studies, 2 focused on dealing with bad actors, 1 focused on bots, and 9 focused on both. For the reviews, none focused solely on bad actors, 1 focused on bots, and 6 discussed dealing with both. Conclusion: Current strategies employed to tackle bots and bad actors in online autism research is a complex and nuanced landscape. A synthesis of 19 relevant studies revealed several distinct approaches. However, none of these methods are completely effective in isolation. This emphasizes the necessity to combine multiple strategies to enhance their overall efficacy. Another recurring concern surfaces throughout the discussion: the imminent obsolescence of current strategies in the face of rapidly evolving AI capabilities. While these methods can have effective outcomes, the relentless progress of AI technologies poses a formidable challenge to their sustainability. Thus, it becomes evident that a dynamic and adaptable approach is needed. Researchers across disciplines must collaborate to find novel method combinations and novel strategies. As the use of online questionnaires in all research—especially studies on autism and other underrepresented populations—continues to grow, it is increasingly important to keep participation barriers low while collecting valid data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.010 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.005 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".