Acceptability and Usability of a Socially Assistive Robot Integrated With a Large Language Model for Enhanced Human-Robot Interaction in a Geriatric Care Institution: Mixed Methods Evaluation
Notice bibliographique
Résumé
BACKGROUND: Socially assistive robots (SARs) hold promise for supporting older adults (OAs) in hospital settings by promoting social engagement, reducing loneliness, and enhancing emotional well-being. They may also assist health care professionals by delivering information, managing routines, and alleviating workload. However, their acceptability and usability remain major challenges, particularly in dynamic real-world care environments. OBJECTIVE: This study aimed to evaluate the acceptability and usability of a SAR in a geriatric day care hospital (DCH) and to identify key factors influencing its adoption by OAs and their informal caregivers. METHODS: Over the course of 1 year, 97 participants (n=65, 67%, OA patients and n=32, 33%, informal caregivers) took part in a mixed methods evaluation of ARI, a socially assistive humanoid robot developed by PAL Robotics. ARI was deployed in the waiting area of a geriatric day care robot in Paris (France), where it interacted with users through voice-based dialogue. After each session, participants completed 2 standardized assessments, the Acceptability E-scale (AES) and the System Usability Scale (SUS), administered orally to ensure accessibility. Open-ended qualitative feedback was also collected to capture subjective experiences and contextual perceptions. RESULTS: Acceptability scores significantly increased across waves (wave 1: mean 15.4/30, SD 5.81; wave 2: mean 20.9/30, SD 5.25; wave 3: mean 22.5/30, SD 4.23; P<.001). Usability scores also improved (wave 1: mean 47.9/100, SD 24.18; wave 2: mean 57.4/100, SD 22.46; wave 3: mean 69.3/100, SD 16.03; P<.001). A strong positive correlation was observed between acceptability and usability scores (r=0.664, P<.001). Qualitative findings indicated improved ease of use, clarity, and user satisfaction over time, particularly following the integration of a large language model (LLM) in wave 2, leading to more coherent, natural, and context-aware interactions. CONCLUSIONS: Successive system enhancements, most notably the integration of an LLM, led to measurable gains in usability and acceptability among patients and informal caregivers. These findings underscore the importance of iterative, user-centered design in deploying SARs in geriatric care environments. TRIAL REGISTRATION: Approved by the French national ethics committee (CPP Ouest II, IRB: 2021/20) as it did not involve randomization or clinical intervention.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».