A court notice arrives in a language its recipient cannot read. The page may contain a hearing date, the conditions of release, a charge, or a deadline for appeal. A translation button can turn the shapes into fluent sentences within seconds. That is already a substantial improvement: the document is no longer completely opaque. Yet one apparently small choice—whether a condition is mandatory, whether a deadline runs from sending or receipt, whether a person “may” or “must” attend—can change what the recipient is able to do next.
This is the central tension in legal machine translation. It can lower the first barrier to a court, help professionals process multilingual material and give people an earlier sense of what a document concerns. But where the translation itself carries testimony, consent, legal advice or the authoritative meaning of a decision, fluent output is not enough. A legal right needs an accountable translation process: a qualified person must be able to inspect the source, correct the result and take responsibility for the version on which someone is expected to act.

Translation is not one legal task
“Legal translation” sounds like a single activity, but it covers decisions with very different consequences. A visitor may want the gist of a court website. A clerk may need to find potentially relevant documents in a large multilingual bundle. A lawyer-linguist may use a machine draft as one input while producing a publication-quality judgment. A suspect may need to understand a charge well enough to instruct counsel. A witness may be answering questions whose exact wording becomes evidence.
The same system can be useful in the first two settings and inadequate in the last two. The important variable is not whether the text concerns law. It is what the translation is being asked to do. Orientation tolerates some uncertainty because the user can return to the original or ask for help. A consequential translation may determine whether a person understands an allegation, gives valid consent, meets a deadline or can challenge the state. In those settings, an error is not merely linguistic friction; it can change effective access to a right.
That distinction is already visible in European law. Directive 2010/64/EU requires interpretation without delay for suspects or accused persons who do not understand the language of criminal proceedings. It also requires written translations of essential documents, including decisions depriving a person of liberty, charges or indictments, and judgments. The required quality is not expressed as a benchmark percentage. It must be sufficient to safeguard fairness, knowledge of the case and the ability to exercise the right of defence. The person must be able to challenge a finding that interpretation or translation is unnecessary and to complain when its quality is insufficient.
This is a demanding standard because it starts from human capability. The question is not simply, “Did the software produce a plausible target-language sentence?” It is, “Could this person understand the case, communicate with counsel and act in time?”
Why fluent output can still be wrong
Modern neural machine translation does not usually replace each source word with a target word. A model converts pieces of the source text into numerical representations, uses surrounding context to estimate relationships between them, and generates a likely sequence in the target language. Multilingual models can share patterns across languages, allowing stronger languages to assist those with less direct training data. This has dramatically expanded coverage and improved fluency.
The No Language Left Behind research published in Nature in 2024, for example, demonstrated a single model across 200 languages and evaluated more than 40,000 translation directions. The work is an important research demonstration of broader coverage. It is not evidence that every language pair, dialect or legal task has reached the same level of reliability, still less that an output is fit to serve as a court-certified translation.
A model selects probable language; it does not determine legal effect. That gap matters because legal errors are often small in form and large in consequence. Negation may be weakened. A modal verb may turn discretion into obligation. A defined term may be rendered inconsistently across a document. A familiar word may refer to different institutions in different legal systems. A name, date, amount or statutory reference may survive most of the sentence while one character changes. The output can remain grammatical enough that a monolingual reader has no signal that anything is wrong.
Context also sets a hard boundary. A witness may use an idiom, hesitate, correct themselves or switch language mid-sentence. A lawyer may rely on a term that has no direct equivalent in the target legal system. An interpreter can ask for clarification and put ambiguity on the record; a one-way translation interface tends to resolve ambiguity silently. Current Council of Europe CEPEJ guidelines on quality interpreting in judicial proceedings expressly say that ambiguity should be addressed on the record rather than silently resolved. They treat AI-supported systems as tools requiring transparent, risk-sensitive use, human oversight and mechanisms for verification or correction, with particular caution for spoken and interactive communication.
The uneven map of language performance
Machine translation quality is not a property of a product in the abstract. It varies by source language, target language, direction, domain, document type and even the quality of the scan or transcript supplied to it. English to a widely represented European language may draw on extensive parallel material. A regional variety, code-switched conversation or low-resource language may have far less suitable data. Transfer learning helps, but it does not manufacture the missing legal vocabulary, cultural context or representative examples.
This unevenness is partly an evidence problem. The FLORES benchmarks developed alongside multilingual systems made it possible to compare many more language directions using professionally translated material. Yet a general benchmark samples broad language competence; it does not reproduce a bail hearing, an asylum narrative or the terminology of a particular jurisdiction. Performance averaged across sentences can also hide rare but severe mistakes.
Legal benchmarks are beginning to narrow the gap. The 2025 SwiLTra-Bench study assembled more than 180,000 aligned Swiss legal translation pairs across laws, court-decision headnotes and press releases. Its results showed that relative performance depended on document type: translation-specific systems were particularly strong on laws but weaker on headnotes, while frontier general models performed differently across tasks. That is useful research evidence. It also illustrates why a court cannot import a headline score into an operational guarantee. Swiss official-language material is not testimony in a low-resource language, and a correct statute translation does not prove that the same system will preserve an ambiguous answer under pressure.
A serious deployment therefore needs evidence for the actual language pair and task. It should test terminology, names, dates, numbers, negation, long dependencies and legally significant modal verbs. Evaluation should count the severity of errors, not only their frequency. A mistranslated adjective in background correspondence is not equivalent to a reversed condition of release. Testing must also include poor audio, scanned documents, dialect, rapid speech and the other conditions in which courts actually operate.
What eTranslation demonstrates
The European Commission’s eTranslation service is a useful example because it is operational infrastructure, not a laboratory benchmark. The Commission describes it as a neural machine translation tool trained on EU professional translation data. Eligible public administrations, organisations and other users can translate text and documents or integrate the service through an API. The Commission’s current service information says the tools support all 24 official EU languages and several others, and that data is processed under EU data-protection rules rather than used to train commercial AI models.
Scale shows how valuable such infrastructure can be. The Commission reports that eTranslation delivered 890 million pages in 2025. But its own description of the workflow preserves an important distinction: eTranslation produces raw machine translations, while Commission translators may use those outputs as an optional resource. The machine product and the professional product are not quietly treated as the same thing.
For justice systems, that suggests a productive role. Secure machine translation can help staff discover what a document is about, route it to the right language team, search a multilingual bundle, prepare terminology and create a first draft for revision. It can reduce the time a person waits before receiving basic orientation. An interface can display the machine version alongside the source, preserve paragraph alignment and flag names or numbers for checking. None of these benefits requires pretending that the first output is authoritative.
Security still depends on the whole workflow. Court documents may contain health information, addresses, criminal allegations, privileged communications or the identity of vulnerable people. A secure institutional service is materially different from pasting text into an unapproved public tool. But the system owner must still decide who may upload material, where it is processed, how long files and logs are retained, who can retrieve them, and how incidents are investigated. The Directive’s requirement that interpreters and translators observe confidentiality cannot be reduced to encryption in transit; it is also a duty attached to people, organisations and access controls.
Where the machine should assist
The strongest case for machine translation is not substitution but earlier and broader assistance. A person who cannot read a court portal should not have to wait days merely to discover which office to contact. A lawyer should be able to identify potentially relevant foreign-language documents before commissioning full translations of a smaller set. An interpreter may benefit from an approved glossary prepared from case documents. A qualified translator can often revise a well-produced draft faster than starting with a blank page, provided that the workflow allows genuine comparison with the source rather than superficial proofreading of fluent output.
There is a serious counterargument to strict human gates. Qualified legal interpreters and translators are scarce, some language combinations are especially difficult to cover, and delay itself can defeat access. Human work is neither perfectly consistent nor immune to fatigue. If every low-risk page requires the same process as a judgment, scarce expertise is consumed where it adds least value and people remain excluded for longer.
The answer is to allocate human responsibility according to consequence, not to insist that every word receive identical treatment. Machine output can be clearly marked and used for orientation, discovery and draft preparation. Human revision becomes mandatory when the text will be relied upon for a decision, waiver, instruction, testimony or official record. Live interaction requires an interpreter who can clarify ambiguity and respond to the participants. The more difficult the language pair, the more consequential the setting and the less opportunity the person has to detect an error, the stronger the human safeguard must be.
Where responsibility must remain human
Some tasks should not be handed to an unattended translation system. A machine should not conduct the interpretation through which a suspect learns the accusation and instructs counsel. It should not be the sole channel for testimony, a guilty plea, informed consent, a waiver of rights or an asylum account. It should not create the authoritative version of a judgment or legislation without accountable legal-linguistic revision. These are not exclusions based on nostalgia for a pre-digital court. They follow from the need to ask questions, preserve ambiguity, understand legal systems and identify who can correct the record.
The Court of Justice of the European Union makes the institutional stakes visible. Its legal translation service works across 24 official languages and 552 possible language combinations. Because the material is highly technical, its translators are lawyer-linguists: people with legal qualifications as well as deep knowledge of other languages. That model does not mean every court must reproduce the same organisation. It shows that authoritative legal translation joins linguistic and legal judgment. The problem is not merely finding a sentence that sounds natural; it is carrying a concept between legal systems without changing what the court has decided.
Human responsibility must also be visible to the affected person. They should be told when a version is machine-generated, what it may safely be used for and how to request a qualified interpreter or corrected translation. The original should remain available. Corrections should be versioned rather than silently overwritten. When a translated passage matters to a decision, the record should identify which version was used and who reviewed it. A person must be able to challenge not only the underlying legal outcome but also a translation that prevented effective participation.
A right needs a correction path
The decisive design question is therefore not whether machine translation is “accurate enough” in general. It is whether a particular system, for a defined language pair and legal task, has a correction path proportionate to the harm an error could cause. That path needs a qualified reviewer with time to compare source and output, authority to stop the process, access to case context and a record of what changed. It also needs an institutional promise that requesting human help will not make the person miss the deadline the translation was meant to explain.
If you value careful reporting on where public technology should assist and where human responsibility must remain visible, you can subscribe to Alkemata for the next article in this series.
Machine translation can carry words across a language barrier. It can carry a legal right only when the surrounding institution carries the uncertainty: by choosing the task carefully, testing the actual language pair, protecting confidential material, providing accountable human interpretation and preserving a route to challenge and correction. The open question for a court is not whether to use the technology. It is which translations it is prepared to let people rely on—and who will answer when a probable sentence proves to be the wrong one.