Evaluation of a Language Translation App in an Undergraduate Medical Communication Course: Proof-of-Concept and Usability Study
Bibliographic record
Abstract
BACKGROUND: Language barriers in medical encounters pose risks for interactions with patients, their care, and their outcomes. Because human translators, the gold standard for mitigating language barriers, can be cost- and time-intensive, mechanical alternatives such as language translation apps (LTA) have gained in popularity. However, adequate training for physicians in using LTAs remains elusive. OBJECTIVE: A proof-of-concept pilot study was designed to evaluate the use of a speech-to-speech LTA in a specific simulated physician-patient situation, particularly its perceived usability, helpfulness, and meaningfulness, and to assess the teaching unit overall. METHODS: Students engaged in a 90-min simulation with a standardized patient (SP) and the LTA iTranslate Converse. Thereafter, they rated the LTA with six items-helpful, intuitive, informative, accurate, recommendable, and applicable-on a 7-point Likert scale ranging from 1 (don't agree at all) to 7 (completely agree) and could provide free-text responses for four items: general impression of the LTA, the LTA's benefits, the LTA's risks, and suggestions for improvement. Students also assessed the teaching unit on a 6-point scale from 1 (excellent) to 6 (insufficient). Data were evaluated quantitatively with mean (SD) values and qualitatively in thematic content analysis. RESULTS: Of 111 students in the course, 76 (68.5%) participated (59.2% women, age 20.7 years, SD 3.3 years). Values for the LTA's being helpful (mean 3.45, SD 1.79), recommendable (mean 3.33, SD 1.65) and applicable (mean 3.57, SD 1.85) were centered around the average of 3.5. The items intuitive (mean 4.57, SD 1.74) and informative (mean 4.53, SD 1.95) were above average. The only below-average item concerned its accuracy (mean 2.38, SD 1.36). Students rated the teaching unit as being excellent (mean 1.2, SD 0.54) but wanted practical training with an SP plus a simulated human translator first. Free-text responses revealed several concerns about translation errors that could jeopardize diagnostic decisions. Students feared that patient-physician communication mediated by the LTA could decrease empathy and raised concerns regarding data protection and technical reliability. Nevertheless, they appreciated the LTA's cost-effectiveness and usefulness as the best option when the gold standard is unavailable. They also reported wanting more medical-specific vocabulary and images to convey all information necessary for medical communication. CONCLUSIONS: This study revealed the feasibility of using a speech-to-speech LTA in an undergraduate medical course. Although human translators remain the gold standard, LTAs could be valuable alternatives. Students appreciated the simulated teaching and recognized the LTA's potential benefits and risks for use in real-world clinical settings. To optimize patients' and health care professionals' experiences with LTAs, future investigations should examine specific design options for training interventions and consider the legal aspects of human-machine interaction in health care settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".