Enhancing privacy protection of physical examination data through synthetic algorithms based on differential privacy
Notice bibliographique
Résumé
BACKGROUND: Health physical examinations play a crucial role in early detection of cancer and chronic disease. However, privacy concerns limit the utilization of this kind of data for health interventions and research. Synthetic data methods based on differential privacy are increasingly used to create complete datasets that protect privacy while enabling data analysis and result interpretation. Hence, the use of synthetic algorithms based on differential privacy for privacy protection of physical examination data is a promising research direction. METHODS: Three synthetic algorithms, PrivBayes, PeGS, and DP-Gibbs were used to generate complete synthetic datasets that adhere to differential privacy standards using physical examination data composed of categorical data, which compared with the existing algorithm Private-PGM. RESULTS: Compared with the existing algorithm, DP-Gibbs can provide privacy preserving capacity of 4.686 (ε = 0.5), while the existing algorithm only with 2.012. In addition, DP-Gibbs provides 0.620 of precision, 0.539 of F1-score, 0.342 of Kappa Coefficient, and 0.765 of AUC-score. The corresponding statistical results of existing algorithm are 0.520, 0.321, 0.188 and 0.695. CONCLUSIONS: The main contributions of this study are the exploration of combination models incorporating different noise forms and Bayesian synthetic algorithms, alongside a comparative analysis against existing algorithms. This study explored the balance between privacy protection and data utility under different levels of privacy protection, and DP-Gibbs offers more stable technical support for de-identifying physical examination data prior to sharing and analysis, which realized the mining and application of a wider range of medical data under the requirements of privacy protection. By leveraging this effective privacy protection technique, clinical researchers can extract valuable insights on diseases and population health from the physical examination data without the risk of leaking private information.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,036 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,011 | 0,037 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».