MétaCan
Menu
Back to cohort
Record W4405967116 · doi:10.1101/2024.12.17.24317999

Human-computer interactions and compassionate healthcare: A Wizard of Oz study using a self-administered AI-assisted cognitive assessment

2024· preprint· en· W4405967116 on OpenAlexaff
Martin C. W. Yu, Weina Jin, David Stitt, Cristina Conati, Giuseppe Carenini, Tal Jarus, Thalia S. Field

Bibliographic record

VenuemedRxiv · 2024
Typepreprint
Languageen
FieldComputer Science
TopicAI in Service Interactions
Canadian institutionsSimon Fraser UniversityUniversity of British Columbia
Fundersnot available
KeywordsWizardWizard of ozCognitionHealth carePsychologyCognitive Assessment SystemMedicineHuman–computer interactionComputer scienceCognitive impairmentWorld Wide WebNeuroscience

Abstract

fetched live from OpenAlex

Abstract Introduction Artificial intelligence (AI) is increasingly transforming healthcare, however, the interaction between the user, AI, and the user’s environment is poorly understood. To elucidate this interplay and support the delivery of compassionate remote care in Occupational Therapy (OT) practice, we created a self-administered remote AI-powered cognitive assessment, facilitated using the Wizard of Oz method, and measured the experiences of patients, controls, caregivers, and healthcare providers. The Wizard of Oz method uses human-computer interaction (HCI) and user experience (UX) research to simulate the functionality of a system before it is fully developed by creating the illusion of a functioning system by having a human behind the scenes, controlling the system’s responses. Research Questions: 1. How are AI-assisted cognitive assessments experienced by patients, caregivers, healthcare providers, and controls? 2. How can we maintain compassionate care while incorporating technology into healthcare? 3. What are the concerns of healthcare providers using chatbots to administer cognitive assessments? Methods 6 participants with progressive cognitive decline, 6 healthy controls, 6 caregivers, and 6 healthcare providers were invited to complete a virtual AI-powered cognitive assessment followed by a survey about their experience and other demographical information. Survey questions were separated into 4 scales: Trust, Compassion, Usability, and Care Experience. Results No statistically significant difference in mean survey scores between participant categories was observed. Factors such as sex, device type, chatbot familiarity, and education had no statistically significant effects. Participants scored statistically significantly lower on the scale Trust (8.09) than on Compassion (8.72). Additionally, those who used the chatbot during the assessment scored statistically significantly lower on the Usability scale compared to those who did not (7.33 vs. 9.20). Conclusion The findings help to evaluate user experience with virtual AI-based cognitive assessments and provide insights that can inform important design characteristics to improve user experience and compassionate care delivery. Author Summary As more technological advancements are being achieved in our modern lives today, we continue to see an increasing amount of these technologies in our healthcare system as well. AI is a popular category of these technological advancements and appropriately implementing this powerful tool into day-to-day medicine may prove to benefit our healthcare system. However, we first need to understand our current views on how AI can affect our delivery of care. In this paper, we explored user experience specifically in the context of a cognitive screening tool using wizard of Oz methodology. In brief, our research explores how various stakeholders (patients, control, caregivers, healthcare providers) experience and feel about using a virtual AI-assisted platform for conducting cognitive assessments. We believe that further exploration of AI in medicine and how it can be improved provides an overview of our attitudes towards implementing artificial intelligence into our healthcare system and will also inspire further research for artificial intelligence in medicine.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.593
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.000
Open science0.0010.006
Research integrity0.0000.002
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.082
GPT teacher head0.422
Teacher spread0.340 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2024
Admission routes1
Has abstractyes

Explore more

Same venuemedRxivSame topicAI in Service InteractionsFrench-language works237,207