An Efficient Private Set Frequency Query Scheme Under Local Differential Privacy
Bibliographic record
Abstract
Crowdsourcing has become a widely utilized method for data collection and analysis; however, privacy concerns remain a significant challenge. In this paper, we introduce a novel and efficient private set frequency (PSF) query scheme designed for crowdsourcing scenarios. Our proposed scheme is based on edge computing and leverages local differential privacy (LDP) and Bloom filter techniques to ensure both query privacy and high communication efficiency. Specifically, we employ two non-colluding edge devices to assist the server in achieving highaccuracy query result estimation while preserving the privacy of both the server's query set and users' sensitive data. A comprehensive security analysis confirms that the query value remains confidential, and users' privacy is guaranteed under <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\varepsilon$</tex>-LDP. Additionally, performance evaluations demonstrate the efficiency and improved accuracy of our proposed scheme.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".