A Comparative Study on Various ML Models Using Synthetic Data for Privacy Preservation
Bibliographic record
Abstract
In the contemporary world, data holds paramount importance across various sectors, such as healthcare, finance, and government. The confidentiality of information is a critical concern, leading to a surge in research focused on data privacy. Over the past two decades, numerous techniques have been employed to safeguard data, with synthetic data generation emerging as a prominent method. Organizations, including those in the medical, banking, and government sectors, prioritize data privacy. Consequently, researchers have developed techniques for generating synthetic data based on real datasets. This study aims to assess the accuracy of synthetic data in comparison to real datasets using fundamental machine learning models such as Logistic Regression, Random Forest Classifier, Decision Tree, and Support Vector Machine. The investigation involves three authentic datasets related to cancer patients. The datasets are employed to train and test the machine learning models, and performance metrics such as accuracy, precision, and recall are computed for the real data. Simultaneously, synthetic data is generated using the authentic dataset, and the synthetic dataset is trained and tested using the same machine learning models. Performance metrics are then computed for the synthetic dataset. The study concludes with a comprehensive comparison of accuracy, precision, and recall values between synthetic and real data. The comparison is conducted across the four machine learning models to determine which model yields the most favorable results for each dataset.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.051 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".