A Quantitative Comparison of Structural and Distributional Properties of Synthetic Tabular Data in Parkinson’s Disease
Bibliographic record
Abstract
Abstract Background Parkinson’s disease (PD) research relies heavily on patient data, but access is often limited by privacy concerns, data scarcity, and collection costs. Synthetic data generation offers a potential solution, but its utility hinges on rigorously evaluated fidelity to real-world data. This study quantitatively assesses the structural and distributional fidelity of synthetic tabular data designed to represent PD patients. Methods We compared a synthetically generated dataset (N=500 hypothetical entries) against an anonymized real-world dataset (N=57 PD patients) containing demographics, clinical scores (UPDRS, MoCA), and mobility data (6MWT-related variables). The evaluation focused on three key quantitative metrics: (1) Column Correlation Stability, measured by the average absolute difference between Pearson correlation matrices, assessed overall and for clinically relevant variable subgroups (6MWT, UPDRS, MoCA); (2) Principal Component Analysis (PCA), evaluating the variance captured by the top principal components in both datasets; and (3) Jensen-Shannon Distance (JSD), quantifying the distributional similarity between real and synthetic variables across different groups. Results The overall average absolute correlation difference between the real and synthetic datasets was 0.049, indicating moderate preservation of pairwise variable relationships globally. However, stability varied across subgroups, with the 6MWT group showing higher fidelity (difference ∼0.044) compared to the UPDRS (∼0.080) and MoCA (0.081) groups. PCA revealed that the first two principal components captured 21.36% and 16.36% of the variance, respectively, with visual analysis showing partial overlap between real and synthetic data clusters. Average JSD values indicated moderate distributional similarity overall, with the MoCA group exhibiting the highest fidelity (JSD = 0.0573), while Demographics (0.1167), Clinical (0.1256), and 6MWT (0.1175) groups showed lower distributional similarity. Conclusion Synthetic data generation techniques can replicate univariate distributional properties of PD patient data with moderate success, particularly for certain variable types like cognitive assessments (MoCA). However, accurately capturing the complex multivariate correlation structures, crucial for understanding symptom interactions and building predictive models, remains a significant challenge, especially within specific clinical domains like UPDRS. While synthetic data holds promise for addressing data access issues in PD research, particularly for tasks less sensitive to correlation structure, its application requires careful, context-specific validation. Further development is needed to enhance the structural fidelity of synthetic tabular data for high-stakes, multivariate clinical research applications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".