Exploring Physician Perspectives on Using Real-world Care Data for the Development of Artificial Intelligence–Based Technologies in Health Care: Qualitative Study
Bibliographic record
Abstract
BACKGROUND: Development of artificial intelligence (AI)-based technologies in health care is proceeding rapidly. The sharing and release of real-world data are key practical issues surrounding the implementation of AI solutions into existing clinical practice. However, data derived from daily patient care are necessary for initial training, and continued data supply is needed for the ongoing training, validation, and improvement of AI-based solutions. Data may need to be shared across multiple institutions and settings for the widespread implementation and high-quality use of these solutions. To date, solutions have not been widely implemented in Germany to meet the challenge of providing a sufficient data volume for the development of AI-based technologies for research and third-party entities. The Protected Artificial Intelligence Innovation Environment for Patient-Oriented Digital Health Solutions (pAItient) project aims to meet this challenge by creating a large data pool that feeds on the donation of data derived from daily patient care. Prior to building this data pool, physician perspectives regarding data donation for AI-based solutions should be studied. OBJECTIVE: This study explores physician perspectives on providing and using real-world care data for the development of AI-based solutions in health care in Germany. METHODS: As a part of the requirements analysis preceding the pAItient project, this qualitative study explored physician perspectives and expectations regarding the use of data derived from daily patient care in AI-based solutions. Semistructured, guide-based, and problem-centered interviews were audiorecorded, deidentified, transcribed verbatim, and analyzed inductively in a thematically structured approach. RESULTS: Interviews (N=8) with a mean duration of 24 (SD 7.8) minutes were conducted with 6 general practitioners and 2 hospital-based physicians. The mean participant age was 54 (SD 14.1; range 30-74) years, with an average experience as a physician of 25 (SD 13.9; range 1-45) years. Self-rated affinity toward modern information technology varied from very high to low (5-point Likert scale: mean 3.75, SD 1.1). All participants reported they would support the development of AI-based solutions in research contexts by donating deidentified data derived from daily patient care if subsequent data use was made transparent to them and their patients and the benefits for patient care were clear. Contributing to care optimization and efficiency were cited as motivation for potential data donation. Concerns regarding workflow integration (time and effort), appropriate deidentification, and the involvement of third-party entities with economic interests were discussed. The donation of data in reference to psychosomatic treatment needs was viewed critically. CONCLUSIONS: The interviewed physicians reported they would agree to use real-world care data to support the development of AI-based solutions with a clear benefit for daily patient care. Joint ventures with third-party entities were viewed critically and should focus on care optimization and patient benefits rather than financial interests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".