Artificial intelligence after the bedside: co-design of AI-based clinical informatics workflows to routinely analyse patient-reported experience measures in hospitals
Bibliographic record
Abstract
OBJECTIVE: To co-design artificial intelligence (AI)-based clinical informatics workflows to routinely analyse patient-reported experience measures (PREMs) in hospitals. METHODS: The context was public hospitals (n=114) and health services (n=16) in a large state in Australia serving a population of ~5 million. We conducted a participatory action research study with multidisciplinary healthcare professionals, managers, data analysts, consumer representatives and industry professionals (n=16) across three phases: (1) defining the problem, (2) current workflow and co-designing a future workflow and (3) developing proof-of-concept AI-based workflows. Co-designed workflows were deductively mapped to a validated feasibility framework to inform future clinical piloting. Qualitative data underwent inductive thematic analysis. RESULTS: Between 2020 and 2022 (n=16 health services), 175 282 PREMs inpatient surveys received 23 982 open-ended responses (mean response rate, 13.7%). Existing PREMs workflows were problematic due to overwhelming data volume, analytical limitations, poor integration with health service workflows and inequitable resource distribution. Three potential semiautomated, AI-based (unsupervised machine learning) workflows were developed to address the identified problems: (1) no code (simple reports, no analytics), (2) low code (PowerBI dashboard, descriptive analytics) and (3) high code (Power BI dashboard, descriptive analytics, clinical unit-level interactive reporting). DISCUSSION: The manual analysis of free-text PREMs data is laborious and difficult at scale. Automating analysis with AI could sharpen the focus on consumer input and accelerate quality improvement cycles in hospitals. Future research should investigate how AI-based workflows impact healthcare quality and safety. CONCLUSION: AI-based clinical informatics workflows to routinely analyse free-text PREMs data were co-designed with multidisciplinary end-users and are ready for clinical piloting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.070 | 0.122 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.004 | 0.006 |
| Scholarly communication | 0.007 | 0.005 |
| Open science | 0.003 | 0.008 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".