Human compared to AI: A single-blind evaluation of Spanish language translation quality.
Abstract
e13678 Background: Informed consent forms (ICF) are valuable decision-making aids and information resources for research participants. Regulations require that ICFs be available in a language understandable to a participant. However, translating ICFs requires substantial time, cost, and resources, limiting trial access for those with limited English proficiency. Artificial intelligence (AI), including large language models, may aid ICF translations. This feasibility study evaluated the quality of Spanish language ICFs translated by humans (HT) compared to AI. Methods: We used three ICFs with both English and Spanish translations approved for interventional studies between 2021-2024. AI translation of ICFs was done using Amazon Translate. Translations were blinded and evaluated by native Spanish-speaking healthcare professionals, each a qualified translator or interpreter. Evaluators were assigned an English ICF and its corresponding Spanish translation (either HT or AI). Each sentence was individually reviewed and scored using a five-point Likert scale across seven quality metrics: flow, correctness, accuracy, representativeness, preservation, spelling, and punctuation. All scores were dichotomized as “acceptable” (score ≥ 3) or “not acceptable” (score < 3). Results: Combined, the three ICFs represented 787 unique English sentences yielding 1,574 Spanish translations available for scoring. Evaluations were analyzed by sentence, section, and whole document. At the sentence level, the acceptable HT and AI translations were highly aligned, 97.9% precision, 97.9% recall, and 95.9% accuracy. However, some HT consent sections were rated “not acceptable” primarily because not all English words were translated. All AI translated sections were rated as “acceptable”. Post-hoc analysis suggested that AI language quality and consistency might be further improved through retrieval augmented generation, especially for sections that used standardized language. Whole document median scores across all criteria were "very good" (Likert = 5) for both HT and AI. Interquartile ranges for all HT criteria and most AI were 0 (AI correctness and accuracy were slightly higher at 1). Conclusions: This feasibility study demonstrated AI translated ICFs are comparable in quality to human translators. Utilization of AI for translating ICFs has potential to reduce time, cost, and resources associated language translation. AI could also significantly improve clinical practice efficiency and facilitate access to research studies for those with limited English proficiency. Further research should quantify the efficiency gained by using AI translation and validate results. It is also crucial to refine and standardize translation quality evaluation criteria. AI translation may help facilitate recruitment of Spanish-speaking participants and help address language as critical barrier to research participation.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Tanna Nelson
NIH/National Cancer Institute, Rockville, MD
Shaalan Beg
UT Southwestern Medical Center, Coppell, TX
Umit Topaloglu
James L Gulley
Center for Immuno-Oncology, Center for Cancer Research, National Cancer Institute, National Institutes of Health, Bethesda, MD
Liliana Robles
UT Southwestern Medical Center, Dallas, TX
Fabian Enrique Robles
UT Southwestern Medical Center, Dallas, TX
Anne-Marie Meyer
NIH/National Cancer Institute, Rockville, MD