Linking medical licensing examination scores with longitudinal physician practice data using a privacy preserving protocol
Bibliographic record
Abstract
IntroductionMedical education and regulatory bodies do not often share performance data due to privacy concerns. Innovative approaches are needed to facilitate research while preserving security and privacy. To this end, a privacy preserving protocol was employed linking medical examination and regulatory data to examine future physician competence across the career.
 Objectives and ApproachThis study extends previous work linking de-identified Canadian medical licensing examination data with medical regulatory outcomes to answer the following question: is there a predictive relationship between licensing examination scores and post-licensure practice outcomes? A privacy preserving protocol using a third party organization was employed to link data between two disparate organizations - a medical licensing examination organization (MLE) and a medical regulatory authority (MRA). Multiple years of licensing examinations were linked to thirteen years of regulatory assessment outcomes (2004 – 2016) without identifiable data being shared to either party.
 ResultsMedical Identification Number for Canada (MINC) was used as a common identifying variable between the two organizations. First, the analytic cohort was created by linking identifying variables of the physicians of interest from both parties, thereby creating a common cohort. The third-party organization then created an encryption key using the common cohort and the MLE examination data. The key was given to the MRA and the encrypted, de-identified examination data was given back to the MLE. Lastly, the MRA data was de-identified, encrypted and transferred to the MLE for analysis. This ensured neither party had access to each other’s encrypted data and the key simultaneously.
 Conclusion/ImplicationsPrivacy preserving protocols enhance opportunities for novel research questions and data linkages within and across sectors; here, results from this analysis may enhance the utility of medical licensing exams by providing evidence for secondary uses. Furthermore, it will offer other physician organizations evidence to support physicians across their career trajectory.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.038 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.013 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".