How Are Linkage Results Using Privacy-Preserving Record Linkage Different?
Bibliographic record
Abstract
IntroductionPrivacy-Preserving Record Linkage (PPRL) presents opportunities to improve privacy protection when performing record linkage on the most sensitive data. Currently our linkage agency performs all linkages in clear text, but expansion of data sources is now including extremely sensitive data, such as justice data. Understanding that specific circumstances may demand different approaches to linkage, we evaluated a PPRL algorithm implemented through the LinXmart software. This is the first real-world evaluation of PPRL in British Columbia and among the first in Canada. Objectives and ApproachOur standard linkage method is probabilistic and relies on rules established by analysts to determine accepted links. Datasets are linked to a population spine (N=8,440,442) containing all current and past residents of the province. LinXmart was configured to link to the top weighted candidate above a predetermined confidence threshold. We evaluated performance by comparing the standard method to PPRL for three increasingly complex (messy) datasets. Initial results on the simplest/cleanest dataset informed an iterative process to improve implementation of PPRL. ResultsOverall linkage rates were lower for standard linkage (81%) compared to PPRL (90%). Records with a unique ID linked at similarly high rates in clear-text and PPRL, while the performance of PPRL with records without the unique ID varied depending on the exact parameters chosen for the match threshold and field comparisons. Conclusion / ImplicationsThis work suggests that for datasets that include a well-populated unique identifier, PPRL can be implemented in real-world linkages without a substantial drop-off in linkage quality. Messier data require careful tuning of linkage parameters to match the performance of clear linkage. PPRL may best be used in cases where clear text identifiers cannot be shared, and where some degradation in linkage rates is acceptable.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.040 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.006 | 0.009 |
| Open science | 0.011 | 0.004 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".