Conditionally positive: a qualitative study of public perceptions about using health data for artificial intelligence research
Bibliographic record
Abstract
OBJECTIVES: Given widespread interest in applying artificial intelligence (AI) to health data to improve patient care and health system efficiency, there is a need to understand the perspectives of the general public regarding the use of health data in AI research. DESIGN: A qualitative study involving six focus groups with members of the public. Participants discussed their views about AI in general, then were asked to share their thoughts about three realistic health AI research scenarios. Data were analysed using qualitative description thematic analysis. SETTINGS: Two cities in Ontario, Canada: Sudbury (400 km north of Toronto) and Mississauga (part of the Greater Toronto Area). PARTICIPANTS: Forty-one purposively sampled members of the public (21M:20F, 25-65 years, median age 40). RESULTS: Participants had low levels of prior knowledge of AI and mixed, mostly negative, perceptions of AI in general. Most endorsed using data for health AI research when there is strong potential for public benefit, providing that concerns about privacy, commercial motives and other risks were addressed. Inductive thematic analysis identified AI-specific hopes (eg, potential for faster and more accurate analyses, ability to use more data), fears (eg, loss of human touch, skill depreciation from over-reliance on machines) and conditions (eg, human verification of computer-aided decisions, transparency). There were mixed views about whether data subject consent is required for health AI research, with most participants wanting to know if, how and by whom their data were used. Though it was not an objective of the study, realistic health AI scenarios were found to have an educational effect. CONCLUSIONS: Notwithstanding concerns and limited knowledge about AI in general, most members of the general public in six focus groups in Ontario, Canada perceived benefits from health AI and conditionally supported the use of health data for AI research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.037 | 0.053 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.011 | 0.015 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.002 | 0.007 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".