Private data sharing between decentralized users through the privGAN\n architecture
Bibliographic record
Abstract
More data is almost always beneficial for analysis and machine learning\ntasks. In many realistic situations however, an enterprise cannot share its\ndata, either to keep a competitive advantage or to protect the privacy of the\ndata sources, the enterprise's clients for example. We propose a method for\ndata owners to share synthetic or fake versions of their data without sharing\nthe actual data, nor the parameters of models that have direct access to the\ndata. The method proposed is based on the privGAN architecture where local GANs\nare trained on their respective data subsets with an extra penalty from a\ncentral discriminator aiming to discriminate the origin of a given fake sample.\nWe demonstrate that this approach, when applied to subsets of various sizes,\nleads to better utility for the owners than the utility from their real small\ndatasets. The only shared pieces of information are the parameter updates of\nthe central discriminator. The privacy is demonstrated with white-box attacks\non the most vulnerable elments of the architecture and the results are close to\nrandom guessing. This method would apply naturally in a federated learning\nsetting.\n
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.002 | 0.006 |
| Open science | 0.002 | 0.006 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".