Some assembly required: tracing the interpretative work of Clinical Competency Committees
Bibliographic record
Abstract
OBJECTIVES: This qualitative study describes the social processes of evidence interpretation employed by Clinical Competency Committees (CCCs), explicating how they interpret, grapple with and weigh assessment data. METHODS: Over 8 months, two researchers observed 10 CCC meetings across four postgraduate programmes at a Canadian medical school, spanning over 25 hours and 100 individual decisions. After each CCC meeting, a semi-structured interview was conducted with one member. Following constructivist grounded theory methodology, data collection and inductive analysis were conducted iteratively. RESULTS: Members of the CCCs held an assumption that they would be presented with high-quality assessment data that would enable them to make systematic and transparent decisions. This assumption was frequently challenged by the discovery of what we have termed 'problematic evidence' (evidence that CCC members struggled to meaningful interpret) within the catalogue of learner data. When CCCs were confronted with 'problematic evidence', they engaged in lengthy, effortful discussions aided by contextual data in order to make meaning of the evidence in question. This process of effortful discussion enabled CCCs to arrive at progression decisions that were informed by, rather than ignored, problematic evidence. CONCLUSIONS: Small groups involved in the review of trainee assessment data should be prepared to encounter evidence that is uncertain, absent, incomplete, or otherwise difficult to interpret, and should openly discuss strategies for addressing these challenges. The answer to the problem of effortful processes of data interpretation and problematic evidence is not as simple as generating more data with strong psychometric properties. Rather, it involves grappling with the discrepancies between our interpretive frameworks and the inescapably subjective nature of assessment data and judgement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".