Sampling in Qualitative Research: Insights from an Overview of the Methods Literature
Bibliographic record
Abstract
The methods literature regarding sampling in qualitative research is characterized by important inconsistencies and ambiguities, which can be problematic for students and researchers seeking a clear and coherent understanding. In this article we present insights about sampling in qualitative research derived from a systematic methods overview we conducted of the literature from three research traditions: grounded theory, phenomenology, and case study. We identified and selected influential methods literature from each tradition using a purposeful and transparent procedure, abstracted textual data using structured abstraction forms, and used a multistep approach for deriving conclusions from the data. We organize the findings from this review into eight topic sections corresponding to the major domains of sampling identified in the review process: definitions of sampling, usage of the term sampling strategy, purposeful sampling, theoretical sampling, sampling units, saturation, sample size, and the timing of sampling decisions. Within each section we summarize how the topic is characterized in the corresponding literature, present our comparative analysis of important differences among research traditions, and offer analytic comments on the findings for that topic. We identify several specific issues with the available guidance on certain topics, representing opportunities for future methods authors to improve our collective understanding.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.312 | 0.304 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.010 | 0.015 |
| Science and technology studies | 0.007 | 0.012 |
| Scholarly communication | 0.011 | 0.013 |
| Open science | 0.003 | 0.008 |
| Research integrity | 0.004 | 0.005 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".