By Tania Strahan, with Hannah Feely, Jessica Fernando, Sophia Chan and Lucy Andresen, in collaboration with La Trobe University
The translation of medical texts is a serious matter, in some cases it can be literally a matter of life and death. In other cases, getting it wrong can involve huge financial costs, or can cause unwarranted fear or other emotional strain.
Maybe more than any other kind of translation, medical translations need to be accurate – lexically and culturally.
Medical translators are highly trained specialists, and not just in medical terminology, but also in understanding the end user and thus the reason for the translation. Is the translation designed for a pamphlet accompanying a new medicine, or is it an information brochure targeting immigrants?
Around 20% of humanitarian migrants to Australia are not literate in their first language, according to the Australian Department of Social Services. This group is particularly vulnerable to poor translations about critical information.
The lack of health literacy in CALD (culturally and linguistically diverse) communities in Australia during the COVID-19 pandemic allowed misinformation about the virus and vaccines to circulate within CALD communities, that contributed to vaccine hesitancy and impeded effective vaccine rollout, as documented by Australian Policy Online. Easy access to accurate translations is critical to keeping vulnerable groups healthy.
LLM-chatbots need to give accurate medical advice
Increasingly, people are turning to LLMs for medical advice. A study published in PubMed Central in March 2025 showed that nearly 50% of respondents who “both use AI and self-report mental health challenges are utilizing major LLMs for therapeutic support”. For health and medical information obtained from LLM-based chatbots, “48.4% (30/63) of participants followed the advice”, with 90% of the study’s respondents citing accessibility as the reason for choosing LLM-based chatbots for their mental health questions.
Similarly, a Sentio University study in California in March 2025 declared that “ChatGPT may be the largest provider of mental health support in the United States” (Sentio University), and by August 2025, Rolling Stone had reported that nearly 40% of Americans trust LLM chatbots “in navigating healthcare decisions”, although this may be driven by “dissatisfaction or concerns about the state of healthcare in U.S.”. Rolling Stone also reported that 48% of men believe that LLMs are “a reliable source of health information”, rising to 52% of men aged 45-54.
With this level of trust already granted to LLMs, the medical advice they dispense needs to be accurate.
Going multilingual? LXT covers 1,000+ language locales.
From low-resource languages to regional dialects, LXT’s native-speaker network delivers annotation quality that generic crowdsourcing platforms can’t match.
Your flywheel is only as fast as your annotation pipeline.
LXT embeds into your retraining loop, delivering labeled data at the cadence your model demands, not months later.
Medical translations in 2023
In October 2023, researchers at the Georgia Institute of Technology asked two LLM-chatbots (Gemini & ChatGPT) 2,000 questions that regular people pose about “diseases, medical procedures, medications, and other general health topics” , as reported by Scientific American, and then translated these questions into the most widely spoken languages on Earth: Mandarin Chinese, Hindi and Spanish, using the LLMs. Chinese and Spanish answers to these translated questions were incorrect one-fifth of the time, while Hindi responses were inaccurate two-thirds of the time, as reported by Scientific American.
Medical translations in 2026
LXT, in partnership with LaTrobe University, conducted its own investigation to see if medical translations and health advice from LLM-chatbots has improved over the past few years. Taking into account recent advances including fine-tuning and new guardrails intended to prevent unethical and dangerous responses from LLM-chatbots, and the gap between high- and low-resource languages in language model development, we tested how a few LLM-chatbots translated a small set of short medical texts to get a feel for the current state of the technology.
Text 1. In the absence of abnormal clinical results, the patients’ symptoms indicate that a diagnosis of Fibromyalgia would be appropriate. Patient reports widespread body pain, fatigue, and sleep disturbances. Its cause is unknown but may involve genetic factors and is linked to an oversensitive pain system, where the body perceives normal signals as painful.
Text 2. Neuropathy is when nerve damage leads to pain, weakness, numbness or tingling in one or more parts of your body.
Text 3. The patient is bedridden and has been feeling symptoms of shortness of breath, kidney problems, and paresthesia. We recommend the treatment plan to be 3-5 weeks long including hemodialysis and close monitoring.
We used English as the original language for our medical texts, as it is often used as a “pivot” or “bridge” language, meaning we expect translations from English into other languages to be more accurate than LLM translations between other language pairs.
We used several LLM-chatbots to translate these short medical texts into six languages:
- One high-resource language: Mandarin Chinese (often used as a pivot language for Japanese, Korean, and the other languages spoken in China)
- Three medium-resource languages: Norwegian, Danish and Icelandic (all quite closely related to each other and English)
- Three low-resource language: Sinhalese (Sri Lanka, low English proficiency University of Moratuwa), Romanesco (central Italian dialect, spoken not written), Cantonese (classified as low-resource due to a lack of written resources, ACM Digital Library).
Results
High-resource language (Mandarin Chinese)
Mandarin Chinese: The translations were consistently clear and accurate, including the appropriate use of culturally and contextually precise terms such as 监测 jiāncè for ‘close monitoring [of symptoms]’. This reflects a pattern noted in the literature: high-resource languages are reproduced with a high degree of fidelity, even when this results in syntactically complex output. Minor inaccuracies such as 报告 bàogào for ‘[the patient] reports’ do not affect the clarity of the translations. In medical contexts, this faithful reproduction of clinical detail is particularly valuable.
Medium-resource languages (Norwegian, Danish, Icelandic)
On the whole, all of the Norwegian, Danish and Icelandic texts were translated cleanly and clearly, with one startling exception.
Norwegian: Norwegian has two official written languages, one (Bokmål ‘book language’, based on Danish) more widely used than the other (Nynorsk ‘New Norwegian’, based on the grammatically conservative dialects).
These written forms differ grammatically and lexically, but are considered mutually intelligible. Some differences (which were not always reflected in the LLM translations) include:
- Nynorsk always has three genders, but Bokmål has mostly collapsed these to two
- Bokmål has a productive genitive construction with -s (like in English) pasientens symptomer ‘patient.the.GENITIVE symptom.PLURAL’, while Nynorsk prefers to reverse the order of the nouns and link them with a preposition symptoma til pasienten ‘symptom.the-PLURAL to patient.the’.
While some lexical and grammatical outputs could be disputed, and Claude’s Nynorsk contained various Bokmål features, no unnatural or misleading artefacts obfuscated the intended meaning in any material way.
Danish: Like Norwegian, no artefacts prevented the original meaning from coming through in the translated texts.
Icelandic: The most egregious translation error in these medium-resource languages came from Claude Sonnet 4.5, which, on 2026-01-30, translated fibromyalgia as “taugaveikiverk”.
While Claude produced an accurate breakdown of the translation it generated, it is also highly misleading. Tauga does mean ‘nerve’, and veiki does mean ‘weakness’, but together, taugaveiki is ‘typhus’, or, more colloquially, a ‘hysterical or neurotic breakdown’. Verk does mean ‘pain’, but in this context, it implies ‘pain from having a nervous breakdown’. The actual translation of fibromyalgia into Icelandic is vefjagigt, which means ‘muscle/tendon’ + ‘arthritis/rheumatism’.
Claude’s translation is so misleading, as well as confidently stated (without a source), that LXT re-did the translation to confirm the result. Four days after the initial test, Claude confidently translated fibromyalgia as taugatrefjagigt (which Google’s Gemini AI Overview also incorrectly back-translates as fibromyalgia).
This translation is NOT accurate.
In both instances, the frame for the translation request was in English, the same as the source text: “Translate this from English to Icelandic”.
When the frame itself was in Icelandic (Nennirðu þýða þetta fyrir mig, svo ég skil hvað stendur hér: ‘Please translate this for me, so I understand what is written here:’), the result was starkly different, and the correct term vöðvagigt ‘muscle-arthritis’ was given, along with a reasonable explanation of the translated definition.
LXT would like to draw the conclusion that the language framing a prompt matters, and that LLMs may do better translating into the frame language rather than from it. However, Qwen did not produce an accurate translation for fibromyalgia, even with an Icelandic frame. Instead, it used the English term (which would not be inaccurate for an Icelandic reader), but suggested an Icelandic-looking word fjölmyalgia ‘many’ + ‘myalgia’, which is not at all accurate or even comprehensible.
Gemini used the correct term when asked in English for a translation, as did ChatGPT.
Low-resource languages (Sinhalese, Romanesco, Cantonese)
Sinhalese: Our red-teaming revealed a persistent reliance on literary Sinhalese, which can alienate younger speakers or those less educated. Concerningly, Gemini raised the bar on the degree of fatigue required for fatigue, translating it as ‘severe fatigue’. The addition of ‘severe’ is common in literary and medical jargon, but could easily prevent a person from seeking medical assistance if they interpreted the words literally, leading them to believed that their symptoms were not that great.
Both Qwen3 and Gemini replaced ‘nerve damage leads to pain’ with ‘nerve damage causes pain’, which is probably close enough for a casual user but which may dissuade a person from seeking further medical advice if they believe the entire causal chain is clear.
Like with Icelandic, Qwen mistranslated a key term in the source text, this time translating fatigue as අධික අස්වැන්න adhika asvænna ‘severe’ + ‘harvests’, which is pure confabulation. This may have come about because අස්වැන්න asvænna means ‘yield, output, effort, harvests’, thus the two words could be literally translated as ‘excessive output’ instead of ‘chronic fatigue’.
Romanesco: These translations were generally clear, although tended to use standard Italian words and phrases rather than Romanesco.
Cantonese: These translations were generally clear, but had a tendency to use Mandarin function words instead of Cantonese. As with Sinhalese, the LLMs confabulated symptom severity, including the phrases 長期 coeng4kei4 ‘long-term’ and 一直 jat1zik6 ‘persistently’, neither of which were present in the source text.
Like the Icelandic translations, the LLMs produced very different results depending on whether the request was given in English or Cantonese. However, the outcome was the opposite.
When the English text was prefaced with 呢句係咩意思? ni1 geoi3 hai6 me1 ji3si1? ‘What does this mean?’, the LLMs took this literally, and only Claude provided a straight translation. Qwen and Gemini both explained the English text line-by-line without providing full translations, but providing extra information about the subject. ChatGPT started with a line-by-line explanation but then summarised the second half of the text instead of explaining it.
All of these LLMs produced reasonable, although Mandarin-flavoured, translations when the request was framed in English.
Lower-resource languages still lag behind high-resource languages
LLMs have made remarkable progress in medical translation, particularly for high-resource languages, but our red-teaming shows that critical risks remain for low-resource languages and vulnerable populations. Subtle lexical shifts, confabulated severity and culturally inappropriate phrasing can materially change health outcomes. As reliance on LLMs for medical information continues to grow, ongoing human-in-the-loop quality checks need to remain part of the overall infrastructure, to ensure equitable access to reliable medical information across all languages.




