Privacy-preserving querying mechanism on privately encrypted personal health records
Bibliographic record
Abstract
The affordability of cloud data storage has made it simpler for users to store and access data online from any location or operating system. These services may be used by users to store sensitive data, such as personal health records or financial data. Many service providers offer features such as analyzing the users’ private data to generate useful reports for medical data. Storing such sensitive data on the cloud raises many privacy concerns. While encryption can ensure data confidentiality, it introduces the challenge of analyzing the privately encrypted data while preserving the privacy of the users and the querying entity. In this paper, we address this problem by proposing a network protocol that would allow a third party, such as a health organization, to query privately encrypted data without relying on a trusted entity. The protocol we propose preserves the privacy of the users and the querying entity. The protocol relies on homomorphic, threshold cryptography, and randomization to allow for secure, distributed, and privacy-preserving queries. We evaluate the performance of our protocol and report on the results of the implementation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".