Methods for Authenticating Participants in Fully Web-Based Mobile App Trials from the iReach Project: Cross-sectional Study
Bibliographic record
Abstract
BACKGROUND: Mobile health apps are important interventions that increase the scale and reach of prevention services, including HIV testing and prevention counseling, pre-exposure prophylaxis, condom distribution, and education, of which all are required to decrease HIV incidence rates. The use of these web-based apps as well as fully web-based intervention trials can be challenged by the need to remove fraudulent or duplicate entries and authenticate unique trial participants before randomization to protect the integrity of the sample and trial results. It is critical to ensure that the data collected through this modality are valid and reliable. OBJECTIVE: The aim of this study is to discuss the electronic and manual authentication strategies for the iReach randomized controlled trial that were used to monitor and prevent fraudulent enrollment. METHODS: iReach is a randomized controlled trial that focused on same-sex attracted, cisgender males (people assigned male at birth who identify as men) aged 13-18 years in the United States and on enrolling people of color and those in rural communities. The data were evaluated by identifying possible duplications in enrollment, identifying potentially fraudulent or ineligible participants through inconsistencies in the data collected at screening and survey data, and reviewing baseline completion times to avoid enrolling bots and those who did not complete the baseline questionnaire. Electronic systems flagged questionable enrollment. Additional manual reviews included the verification of age, IP addresses, email addresses, social media accounts, and completion times for surveys. RESULTS: The electronic and manual strategies, including the integration of social media profiles, resulted in the identification and prevention of 624 cases of potential fraudulent, duplicative, or ineligible enrollment. A total of 79% (493/624) of the potentially fraudulent or ineligible cases were identified through electronic strategies, thereby reducing the burden of manual authentication for most cases. A case study with a scenario, resolution, and authentication strategy response was included. CONCLUSIONS: As web-based trials are becoming more common, methods for handling suspicious enrollments that compromise data quality have become increasingly important for inclusion in protocols. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): RR2-10.2196/10174.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.003 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".