Pair-Coding as a Method to Support Intercoder Agreement in Qualitative Research
Bibliographic record
Abstract
The goal of this WiP paper is to provide an overview of our adaptation of pair programming from agile software development to support intercoder agreement in qualitative analysis. In pair programming, two programmers work together on the same program: one writes while the other observes and guides. Pair programming interleaves development and inspection activities and has been shown to produce higher quality software in a shorter time than individuals working alone. Our adaptation of agile software development pair programming to qualitative research, which we have called pair-coding, uses a similar approach to merge the qualitative research activities of coding and consensus-building by having two team members work on the same text simultaneously. In using pair-coding in qualitative research, the active team member highlights passages and assigns codes while the other team member observes and guides. Our research team anecdotally found that pair-coding of qualitative data provided benefits that were similar to the benefits of pair programming. Specifically, consensus and consistency were built continuously rather than at discrete stages. As well, differences in biases and worldviews were overtly revealed in the act of coding, thus improving the bracketing of biases. Finally, we found that consensus was better understood when built in the moment of coding rather than in comparison after the fact. We believe pair-coding could be effective in supporting the trustworthiness and credibility of qualitative analysis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".