EPPQ: Efficient and Privacy-Preserving NN Query Processing for Outsourced High-Dimensional Data
Bibliographic record
Abstract
Extensive schemes have been conducted on the development of efficient and privacy-preserving$k$NN query algorithms in data outsourcing scenarios. However, existing researches primarily address low-dimensional data, posing scalability challenges in higher dimensions. To tackle this issue, we propose an efficient and privacy-preserving$k$NN query scheme for outsourced high-dimensional data (EPPQ), emphasizing the complete lifecycle from secure dimensionality reduction of high-dimensional data to secure$k$NN query on the reduced-dimensional data. Specifically,in the secure dimensionality reduction phase: on the one hand, EPPQ integrates principal component analysis (PCA) for dimensionality reduction to minimize computational overhead. On the other hand, to address privacy concerns during the process of PCA, by incorporating differential privacy (DP), we propose the Privacy-Preserving Data Dimensionality Reduction Algorithm based on PCA (PDDRP).In the secure$k$NN query phase: for one thing, EPPQ facilitates the index of the reduced-dimensional data by k-d tree. To enhance index efficiency, we innovatively propose plaintexts-based distance calculation definitions (PDC definitions) and construct an efficient variant of k-d tree (Ek-d tree), for the first time. For another, the Paillier homomorphic encryption (PHE) technique is leveraged to safeguard privacy when outsourcing Ek-d tree to untrusted cloud servers. Additionally, for ciphertexts-based distance calculations and comparisons, we design the Secure Precomputed Distance protocol (SPCD) and Secure Comparison protocol (SCOM). Finally, we creatively present the Privacy-Preserving$k$NN Query Algorithm based on Ek-d tree (PKQKT) for efficient and secure$k$NN query. Comprehensive security analysis demonstrates that the EPPQ scheme meets the required security properties under thehonest-but-curiousmodel. Extensive experiments confirms that EPPQ achieves high computational efficiency and query accuracy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.001 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.007 |
| Open science | 0.004 | 0.009 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".