Human-computer interactions and compassionate healthcare: A Wizard of Oz study using a self-administered AI-assisted cognitive assessment
Bibliographic record
Abstract
Abstract Introduction Artificial intelligence (AI) is increasingly transforming healthcare, however, the interaction between the user, AI, and the user’s environment is poorly understood. To elucidate this interplay and support the delivery of compassionate remote care in Occupational Therapy (OT) practice, we created a self-administered remote AI-powered cognitive assessment, facilitated using the Wizard of Oz method, and measured the experiences of patients, controls, caregivers, and healthcare providers. The Wizard of Oz method uses human-computer interaction (HCI) and user experience (UX) research to simulate the functionality of a system before it is fully developed by creating the illusion of a functioning system by having a human behind the scenes, controlling the system’s responses. Research Questions: 1. How are AI-assisted cognitive assessments experienced by patients, caregivers, healthcare providers, and controls? 2. How can we maintain compassionate care while incorporating technology into healthcare? 3. What are the concerns of healthcare providers using chatbots to administer cognitive assessments? Methods 6 participants with progressive cognitive decline, 6 healthy controls, 6 caregivers, and 6 healthcare providers were invited to complete a virtual AI-powered cognitive assessment followed by a survey about their experience and other demographical information. Survey questions were separated into 4 scales: Trust, Compassion, Usability, and Care Experience. Results No statistically significant difference in mean survey scores between participant categories was observed. Factors such as sex, device type, chatbot familiarity, and education had no statistically significant effects. Participants scored statistically significantly lower on the scale Trust (8.09) than on Compassion (8.72). Additionally, those who used the chatbot during the assessment scored statistically significantly lower on the Usability scale compared to those who did not (7.33 vs. 9.20). Conclusion The findings help to evaluate user experience with virtual AI-based cognitive assessments and provide insights that can inform important design characteristics to improve user experience and compassionate care delivery. Author Summary As more technological advancements are being achieved in our modern lives today, we continue to see an increasing amount of these technologies in our healthcare system as well. AI is a popular category of these technological advancements and appropriately implementing this powerful tool into day-to-day medicine may prove to benefit our healthcare system. However, we first need to understand our current views on how AI can affect our delivery of care. In this paper, we explored user experience specifically in the context of a cognitive screening tool using wizard of Oz methodology. In brief, our research explores how various stakeholders (patients, control, caregivers, healthcare providers) experience and feel about using a virtual AI-assisted platform for conducting cognitive assessments. We believe that further exploration of AI in medicine and how it can be improved provides an overview of our attitudes towards implementing artificial intelligence into our healthcare system and will also inspire further research for artificial intelligence in medicine.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.006 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".