MétaCan
Menu
Back to cohort
Record W2725076736 · doi:10.2196/diabetes.7446

Machine or Human? Evaluating the Quality of a Language Translation Mobile App for Diabetes Education Material

2017· article· en· W2725076736 on OpenAlexvenueno aff
Xuewei Chen, Sandra Acosta, Adam E. Barry

Bibliographic record

VenueJMIR Diabetes · 2017
Typearticle
Languageen
FieldHealth Professions
TopicInterpreting and Communication in Healthcare
Canadian institutionsnot available
FundersCollege of Education and Human Development, Texas A and M University
KeywordsMobile appsComputer scienceQuality (philosophy)Translation (biology)Machine translationDiabetes mellitusArtificial intelligenceWorld Wide WebMedicineChemistryEndocrinologyBiochemistry

Abstract

fetched live from OpenAlex

BACKGROUND: Diabetes is a major health crisis for Hispanics and Asian Americans. Moreover, Spanish and Chinese speakers are more likely to have limited English proficiency in the United States. One potential tool for facilitating language communication between diabetes patients and health care providers is technology, specifically mobile phones. OBJECTIVE: Previous studies have assessed machine translation quality using only writing inputs. To bridge such a research gap, we conducted a pilot study to evaluate the quality of a mobile language translation app (iTranslate) with a voice recognition feature for translating diabetes patient education material. METHODS: The pamphlet, "You are the heart of your family…take care of it," is a health education sheet for diabetes patients that outlines three recommended questions for patients to ask their clinicians. Two professional translators translated the original English sentences into Spanish and Chinese. We recruited six certified medical translators (three Spanish and three Chinese) to conduct blinded evaluations of the following versions: (1) sentences interpreted by iTranslate, and (2) sentences interpreted by the professional human translators. Evaluators rated the sentences (ranging from 1-5) on four scales: Fluency, Adequacy, Meaning, and Severity. We performed descriptive analyses to examine the differences between these two versions. RESULTS: Cronbach alpha values exhibited high degrees of agreement on the rating outcomes of both evaluator groups: .920 for the Spanish raters and .971 for the Chinese raters. The readability scores generated using MS Word's Flesch-Kincaid Grade Level for these sentences were 0.0, 1.0, and 7.1. We found iTranslate generally provided translation accuracy comparable to human translators on simple sentences. However, iTranslate made more errors when translating difficult sentences. CONCLUSIONS: Although the evidence from our study supports iTranslate's potential for supplementing professional human translators, further evidence is needed. For this reason, mobile language translation apps should be used with caution.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.039
metaresearch head score (Gemma)0.179
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.039
Threshold uncertainty score0.207

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0390.179
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0010.001
Scholarly communication0.0030.002
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.178
GPT teacher head0.581
Teacher spread0.403 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations40
Published2017
Admission routes1
Has abstractyes

Explore more

Same venueJMIR DiabetesSame topicInterpreting and Communication in HealthcareFrench-language works237,207