Evaluation of a Language Translation App in an Undergraduate Medical Communication Course: Proof-of-Concept and Usability Study
Notice bibliographique
Résumé
BACKGROUND: Language barriers in medical encounters pose risks for interactions with patients, their care, and their outcomes. Because human translators, the gold standard for mitigating language barriers, can be cost- and time-intensive, mechanical alternatives such as language translation apps (LTA) have gained in popularity. However, adequate training for physicians in using LTAs remains elusive. OBJECTIVE: A proof-of-concept pilot study was designed to evaluate the use of a speech-to-speech LTA in a specific simulated physician-patient situation, particularly its perceived usability, helpfulness, and meaningfulness, and to assess the teaching unit overall. METHODS: Students engaged in a 90-min simulation with a standardized patient (SP) and the LTA iTranslate Converse. Thereafter, they rated the LTA with six items-helpful, intuitive, informative, accurate, recommendable, and applicable-on a 7-point Likert scale ranging from 1 (don't agree at all) to 7 (completely agree) and could provide free-text responses for four items: general impression of the LTA, the LTA's benefits, the LTA's risks, and suggestions for improvement. Students also assessed the teaching unit on a 6-point scale from 1 (excellent) to 6 (insufficient). Data were evaluated quantitatively with mean (SD) values and qualitatively in thematic content analysis. RESULTS: Of 111 students in the course, 76 (68.5%) participated (59.2% women, age 20.7 years, SD 3.3 years). Values for the LTA's being helpful (mean 3.45, SD 1.79), recommendable (mean 3.33, SD 1.65) and applicable (mean 3.57, SD 1.85) were centered around the average of 3.5. The items intuitive (mean 4.57, SD 1.74) and informative (mean 4.53, SD 1.95) were above average. The only below-average item concerned its accuracy (mean 2.38, SD 1.36). Students rated the teaching unit as being excellent (mean 1.2, SD 0.54) but wanted practical training with an SP plus a simulated human translator first. Free-text responses revealed several concerns about translation errors that could jeopardize diagnostic decisions. Students feared that patient-physician communication mediated by the LTA could decrease empathy and raised concerns regarding data protection and technical reliability. Nevertheless, they appreciated the LTA's cost-effectiveness and usefulness as the best option when the gold standard is unavailable. They also reported wanting more medical-specific vocabulary and images to convey all information necessary for medical communication. CONCLUSIONS: This study revealed the feasibility of using a speech-to-speech LTA in an undergraduate medical course. Although human translators remain the gold standard, LTAs could be valuable alternatives. Students appreciated the simulated teaching and recognized the LTA's potential benefits and risks for use in real-world clinical settings. To optimize patients' and health care professionals' experiences with LTAs, future investigations should examine specific design options for training interventions and consider the legal aspects of human-machine interaction in health care settings.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,009 | 0,013 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,000 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».