AI and Medical Big Data Education in Clinical Medicine Undergraduates in China: Awareness, Need, and Curriculum Preferences (Preprint)
Bibliographic record
Abstract
Background: The rapid integration of artificial intelligence (AI) and medical big data into health care is transforming diagnosis, treatment planning, and research. However, formal education in these areas remains limited in undergraduate medical curricula, particularly in China. Objective: This study aimed to investigate clinical medicine undergraduates' familiarity with AI and medical big data, their perceived need for related courses, and their preferred curriculum design and assessment methods. Methods: A cross-sectional, web-based survey was conducted at Zunyi Medical University, Guizhou, China, from January 10 to 17, 2025. In the institutional context of this study, "clinical medicine" included related clinical-track specialties such as pediatrics and psychiatry. All eligible students (N=1094) were invited, and 871 (79.6%) were included in the final analysis. The self-administered questionnaire was developed based on a literature review and expert consultation, with content validity quantified using the content validity index. Descriptive statistics were used to summarize response distributions. For ordinal outcomes (items 1-14), adjusted ordinal logistic regression models were applied, with gender and grade as predictors and major as a covariate. Given the small number of third- and fourth-year students, grade was modeled as an ordered trend variable. For nominal outcomes (items 15-16), group differences were assessed using chi-square tests or Fisher exact tests, as appropriate. Results: A total of 871 students were analyzed, of whom 62.6% (n=545) were women. Overall familiarity with AI and medical big data was limited: 34.8% (303/871) agreed or strongly agreed that they were familiar with the topic, and only 33% (287/871) reported having at least some prior learning experience. In contrast, the perceived educational need was high: 94% (819/871) considered such a course at least somewhat necessary, 57% (497/871) reported that the course was needed or very needed, 75.5% (658/871) indicated that they would likely or definitely enroll, and 56.5% (492/871) reported that they would likely or definitely engage in self-directed learning. Personalized teaching based on textbooks (566/871, 65%) or open-book examinations (633/871, 72.7%) was the most preferred instructional and assessment format. Preferences for course materials and assessment methods differed by grade but not by gender. Conclusions: Early-stage clinical medicine undergraduates demonstrated limited familiarity with AI and medical big data but expressed a strong demand for related education. Students preferred structured yet flexible instructional formats and open-book assessments. Although the findings are based predominantly on first- and second-year students, they support the development of staged, practice-oriented AI and medical big data curricula tailored to the needs of early-stage clinical medicine undergraduates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.016 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".