Attitudes Toward Common Data Models Among Chinese Biomedical Professionals: Cross-Sectional Survey
Bibliographic record
Abstract
BACKGROUND: In the rapidly evolving landscape of health informatics, adopting a standardized common data model (CDM) is a pivotal strategy for harmonizing data from diverse sources within a cohesive framework. Transitioning regional databases to a CDM is important because it facilitates integration and analysis of vast and varied health datasets. This is particularly relevant in China, where unique demographic and epidemiologic profiles present a rich yet complex data landscape. The significance of this research from the perspective of the Chinese population lies in its potential to bridge gaps among disparate data sources, enabling more comprehensive insights into health trends and outcomes. OBJECTIVE: This study aimed to understand biomedical professionals' and trainees' acceptance of the CDM in medical data management in China and to explore potential advantages and challenges associated with its promotion, implementation, and development in the country. METHODS: We conducted a questionnaire survey using Sojump and distributed it on WeChat to evaluate the Chinese population's acceptance of transitioning from local databases to a standardized CDM. The survey assessed participants' understanding of the CDM and the Observational Medical Outcomes Partnership CDM, as well as their views on the importance of CDM for regional databases in China. Analysis of the survey results revealed the current state, challenges, and trends in CDM application within Chinese health care, providing a foundation for future efforts in data standardization and sharing. The reliability of the questionnaire data was assessed using Cronbach α and Guttman Lambda 6 to determine internal consistency. RESULTS: Our survey of 418 participants revealed that 41.9% (175/418) were aware of the CDM. Recognition of CDM increased with higher education levels and was notably higher among professionals in contract research organizations and the pharmaceutical industry. Knowledge of CDM was primarily gained through literature and conferences, with formal education less common. Logistic regression analysis indicated that individuals with doctoral degrees, researchers, executives, medical professionals, data engineers, Centers for Disease Control and Prevention staff, and statisticians were more likely to be aware of CDM. Subgroup analyses showed higher awareness among doctoral versus nondoctoral and Beijing-based versus non-Beijing respondents, while perceived necessity was broadly comparable across subgroups. Overall, 94.7% (396/418) of respondents believed CDM integration in China is necessary for standardization and efficiency. Despite 60.7% (254/418) optimism for the Observational Medical Outcomes Partnership as the preferred CDM, challenges such as mapping traditional Chinese medicine or Chinese medical insurance remain. CONCLUSIONS: A large proportion of respondents expressed a favorable view of implementing the CDM in regional databases in China, with notable endorsement from the doctoral group and professionals in contract research organizations or pharmaceutical sectors; subgroup differences were concentrated in awareness rather than perceived necessity. Participants suggested enhancing CDM-related education and establishing clear data-sharing regulations to support CDM advancement in China.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.015 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".