MétaCan
Menu
Back to cohort
Record W4410074846 · doi:10.1101/2025.05.02.25326887

Evaluating the Potential of AI-Generated Synthetic Diaries in Parkinson’s Disease Research

2025· preprint· en· W4410074846 on OpenAlexafffund
Farhan Raza, Shahryar Wasif, Dhruvil Patel, Taylor Chomiak, Bin Hu

Bibliographic record

VenuemedRxiv · 2025
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBiomedical Text Mining and Ontologies
Canadian institutionsUniversity of Calgary
FundersCanadian Institutes of Health Research
KeywordsParkinson's diseaseDiseasePsychologyData scienceMedicineComputer scienceInternal medicine

Abstract

fetched live from OpenAlex

Abstract The integration of Artificial Intelligence (AI), particularly large language models like GPT-4o, into Parkinson’s Disease (PD) research presents a novel approach for generating synthetic patient diaries. These technologies offer potential benefits, including addressing data privacy concerns, overcoming limited sample sizes, and accelerating research timelines by providing alternative data sources. By leveraging its internal knowledge, GPT-4o demonstrated the capability to replicate overall symptom prevalence distributions observed in a real PD patient dataset without significant statistical deviation. Despite these advantages, the widespread utility of AI-generated diaries based solely on internal knowledge is hindered by significant limitations identified in this case study. Key challenges include the failure to capture complex inter-variable correlations essential for understanding symptom co-occurrence, and a lack of the narrative richness, contextual depth, and linguistic nuance found in authentic patient reports. These findings underscore the constraints of current models in replicating real-world patient experiences without specific domain grounding. Addressing these challenges requires a multifaceted approach, including domain-specific fine-tuning, enhanced prompt engineering, and potentially hybrid data strategies to improve fidelity for high-stakes research applications. This case study explored the baseline capabilities and limitations of using GPT-4o’s internal knowledge for synthetic PD diary generation. It emphasizes the need for a balanced approach, acknowledging the potential for exploratory uses while highlighting the necessity for rigorous validation and further development before deployment in contexts requiring high fidelity. By fostering continued research and methodological refinement, AI-driven synthetic data generation can be better harnessed to support PD research and ultimately improve patient understanding.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.018
metaresearch head score (Gemma)0.105
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.018
Threshold uncertainty score0.098

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0180.105
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0030.002
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.085
GPT teacher head0.411
Teacher spread0.326 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes2
Has abstractyes

Explore more

Same venuemedRxivSame topicBiomedical Text Mining and OntologiesFrench-language works237,207