MétaCan
Menu
Back to cohort
Record W2914472042 · doi:10.2196/12344

Fraud Detection Protocol for Web-Based Research Among Men Who Have Sex With Men: Development and Descriptive Evaluation

2019· article· en· W2914472042 on OpenAlexvenueno aff
April Ballard, Trey Cardwell, April M. Young

Bibliographic record

VenueJMIR Public Health and Surveillance · 2019
Typearticle
Languageen
FieldMedicine
TopicHIV/AIDS Research and Interventions
Canadian institutionsnot available
FundersNational Institute of Environmental Health SciencesNational Institute on Drug Abuse
KeywordsMen who have sex with menThe InternetInternet privacyWeb applicationProtocol (science)PsychologyWorld Wide WebComputer scienceMedicineHuman immunodeficiency virus (HIV)Family medicineAlternative medicine

Abstract

fetched live from OpenAlex

BACKGROUND: Internet is becoming an increasingly common tool for survey research, particularly among "hidden" or vulnerable populations, such as men who have sex with men (MSM). Web-based research has many advantages for participants and researchers, but fraud can present a significant threat to data integrity. OBJECTIVE: The purpose of this analysis was to evaluate fraud detection strategies in a Web-based survey of young MSM and describe new protocols to improve fraud detection in Web-based survey research. METHODS: This study involved a cross-sectional Web-based survey that examined individual- and network-level risk factors for HIV transmission and substance use among young MSM residing in 15 counties in Central Kentucky. Each survey entry, which was at least 50% complete, was evaluated by the study staff for fraud using an algorithm involving 8 criteria based on a combination of geolocation data, survey data, and personal information. Entries were classified as fraudulent, potentially fraudulent, or valid. Descriptive analyses were performed to describe each fraud detection criterion among entries. RESULTS: Of the 414 survey entries, the final categorization resulted in 119 (28.7%) entries identified as fraud, 42 (10.1%) as potential fraud, and 253 (61.1%) as valid. Geolocation outside of the study area (164/414, 39.6%) was the most frequently violated criterion. However, 33.3% (82/246) of the entries that had ineligible geolocations belonged to participants who were in eligible locations (as verified by their request to mail payment to an address within the study area or participation at a local event). The second most frequently violated criterion was an invalid phone number (94/414, 22.7%), followed by mismatching names within an entry (43/414, 10.4%) and unusual email addresses (37/414, 8.9%). Less than 5% (18/414) of the entries had some combination of personal information items matching that of a previous entry. CONCLUSIONS: This study suggests that researchers conducting Web-based surveys of MSM should be vigilant about the potential for fraud. Researchers should have a fraud detection algorithm in place prior to data collection and should not rely on the Internet Protocol (IP) address or geolocation alone, but should rather use a combination of indicators.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.296
metaresearch head score (Gemma)0.410
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.704
Threshold uncertainty score0.869

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2960.410
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0030.004
Bibliometrics0.0170.011
Science and technology studies0.0040.003
Scholarly communication0.0060.006
Open science0.0040.005
Research integrity0.0030.006
Insufficient payload (model declined to judge)0.0110.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.122
GPT teacher head0.433
Teacher spread0.311 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations119
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Public Health and SurveillanceSame topicHIV/AIDS Research and InterventionsFrench-language works237,207