Secure federated Boolean count queries using fully-homomorphic cryptography
Bibliographic record
Abstract
Abstract Biomedical data is often distributed between a network of custodians, causing challenges for researchers wishing to securely compute aggregate statistics on those data without centralizing everything—the prototypical ‘count query’ asks how many patients match some multifaceted set of conditions across a network of hospitals. Difficulty arises from two sources: (1) the need to deduplicate patients who may be present in the records of multiple hospitals and (2) the need to unify partial records for the same patient which may be split across hospitals. Although cryptographic tools for secure computation promise to enable collaborative studies with formal privacy guarantees, existing approaches either are computationally impractical or support only simplified analysis pipelines. To the best of our knowledge, no existing practical secure method addresses both of these difficulties simultaneously. Here, we introduce secure federated Boolean count queries using a novel 2-stage probabilistic sketching and sampling protocol that can be efficiently implemented in off-the-shelf federated homomorphic encryption libraries (Palisade and Lattigo), provably ensuring data security. To this end, we needed several key technological innovations, including re-encoding the LogLog union-cardinality sketch and designing an appropriate sampling for intersection cardinalities. Our benchmarking shows that we can answer federated Boolean count queries in less than 2 CPU-minutes with absolute errors in the range of 6% of the total number of touched records, while revealing only the final answer and the total number of touched records. With modern core-parallelism, we can thus answer queries on the order of seconds. Our study demonstrates that by computing on compressed and encrypted data, it is possible to securely answer federated Boolean count queries in real-time.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.004 | 0.001 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".