Do You Know Who You’re Talking To? Methodological Reflections on Maintaining Inclusivity and Research Integrity When Responding to Inauthentic Encounters in Online Qualitative Research
Bibliographic record
Abstract
There is an ongoing debate around how to design online synchronous qualitative research studies, and respond in the moment, when researchers suspect that they are engaging with ‘impostor’ or ‘fraudulent’ participants. Initial literature framed ineligible participants as a threat to data quality and the integrity of the research itself, calling for reactionary approaches to potential participants. This paper contributes to the growing literature cautioning that strict screening approaches may negatively harm genuine participants and undermine inclusion efforts. This paper explores the concept of ‘knowing’ research participants in qualitative research, focusing on methods that enhance how we genuinely come to know the participants we seek to include, particularly in reclaiming interactions that may have become curtailed through the expediency of online research. Through consideration of researchers’ ethical responsibilities in relation to what is presumed or learned, we offer methodological reflections on how researchers’ skilful attention to the research encounter may be all that is required to ensure continued research integrity within the context of inauthentic participants. Taking actions to better know participants upholds our ethical responsibilities to them and also has the effect of identifying inauthentic participants who intentionally falsify their accounts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchResearch integrity Domain: Methods · Genre: Methods About the Canadian research system: no · About a Canadian topic: no | Qualitative | low |
| gpt | MetaresearchResearch integrity Domain: Methods · Genre: Methods About the Canadian research system: no · About a Canadian topic: no | Qualitative | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.754 | 0.695 |
| Meta-epidemiology (narrow) | 0.002 | 0.005 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.007 | 0.007 |
| Science and technology studies | 0.036 | 0.115 |
| Scholarly communication | 0.035 | 0.034 |
| Open science | 0.015 | 0.036 |
| Research integrity | 0.014 | 0.024 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".