MétaCan
Menu
Back to cohort
Record W4412137476 · doi:10.2196/76496

Acceptability and Usability of a Socially Assistive Robot Integrated With a Large Language Model for Enhanced Human-Robot Interaction in a Geriatric Care Institution: Mixed Methods Evaluation

2025· article· en· W4412137476 on OpenAlexvenueno aff
Xavier Alameda-Pineda, Séverin Lemaignan, Francesco Tonini, Maribel Pino

Bibliographic record

VenueJMIR Human Factors · 2025
Typearticle
Languageen
FieldPsychology
TopicSocial Robot Interaction and HRI
Canadian institutionsnot available
Fundersnot available
KeywordsPreprintUsabilityRobotHuman–robot interactionHuman–computer interactionInstitutionComputer sciencePsychologyGerontological nursingSociologyMedicineNursingWorld Wide WebArtificial intelligenceSocial science

Abstract

fetched live from OpenAlex

BACKGROUND: Socially assistive robots (SARs) hold promise for supporting older adults (OAs) in hospital settings by promoting social engagement, reducing loneliness, and enhancing emotional well-being. They may also assist health care professionals by delivering information, managing routines, and alleviating workload. However, their acceptability and usability remain major challenges, particularly in dynamic real-world care environments. OBJECTIVE: This study aimed to evaluate the acceptability and usability of a SAR in a geriatric day care hospital (DCH) and to identify key factors influencing its adoption by OAs and their informal caregivers. METHODS: Over the course of 1 year, 97 participants (n=65, 67%, OA patients and n=32, 33%, informal caregivers) took part in a mixed methods evaluation of ARI, a socially assistive humanoid robot developed by PAL Robotics. ARI was deployed in the waiting area of a geriatric day care robot in Paris (France), where it interacted with users through voice-based dialogue. After each session, participants completed 2 standardized assessments, the Acceptability E-scale (AES) and the System Usability Scale (SUS), administered orally to ensure accessibility. Open-ended qualitative feedback was also collected to capture subjective experiences and contextual perceptions. RESULTS: Acceptability scores significantly increased across waves (wave 1: mean 15.4/30, SD 5.81; wave 2: mean 20.9/30, SD 5.25; wave 3: mean 22.5/30, SD 4.23; P<.001). Usability scores also improved (wave 1: mean 47.9/100, SD 24.18; wave 2: mean 57.4/100, SD 22.46; wave 3: mean 69.3/100, SD 16.03; P<.001). A strong positive correlation was observed between acceptability and usability scores (r=0.664, P<.001). Qualitative findings indicated improved ease of use, clarity, and user satisfaction over time, particularly following the integration of a large language model (LLM) in wave 2, leading to more coherent, natural, and context-aware interactions. CONCLUSIONS: Successive system enhancements, most notably the integration of an LLM, led to measurable gains in usability and acceptability among patients and informal caregivers. These findings underscore the importance of iterative, user-centered design in deploying SARs in geriatric care environments. TRIAL REGISTRATION: Approved by the French national ethics committee (CPP Ouest II, IRB: 2021/20) as it did not involve randomization or clinical intervention.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.718
Threshold uncertainty score0.869

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.069
GPT teacher head0.499
Teacher spread0.430 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations11
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Human FactorsSame topicSocial Robot Interaction and HRIFrench-language works237,207