You receive a long email in English from an overseas client. You don't have time to write a reply. For now, you throw it into DeepL or, more recently, ChatGPT to turn it into Japanese and read it on your screen. You understand the content. You can make a judgment. Or at least, you feel like you have.
The reverse is also true. You need to send out an English announcement internally, so you run a draft written in Japanese through machine translation, scan the resulting English, think "it looks okay," and hit send. The reply from the other party doesn't seem particularly strange. So you think, "It must have gone well this time too."
The problem is that no one on this side has verified whether it actually went well. The process of lining up the source and target texts and judging "this part is OK," "this part is risky," or "this part is fatal" is missing from modern translation pipelines.
This isn't limited to personal emails. Marketing copy, internal notices, draft contracts, IR materials, support FAQs—all kinds of documents have been placed on machine translation in the last few years. The volume has exploded. Yet, the layer for checking quality has hardly increased. Some people think they are checking. They call "reading it and feeling no discomfort" a check. It is not. That is not a check; it is wishful thinking.
And what many people overlook here is that this isn't just about machine translation. Even for translations ordered from professional translators, there is currently almost no way for the client to independently verify whether the translation is consistent with their intent throughout the entire piece. They read the delivery, think "it reads like Japanese" or "it passes as English," and that's it. No one is performing the task of verifying, item by item, whether the work was done according to the brief for the entire manuscript due to cost issues. It might be more accurate to say there is no way to do it.
The field of translation evaluation has taken on the job of solving this problem. However, this field is not doing as well as those outside the industry might think.
Machine Translation Is Not as Perfect as Everyone Thinks
First, one myth needs to be dispelled.
The output of machine translation and LLMs in 2026 is certainly at a different level than it was a few years ago. This is a fact. For 80% of daily uses, humans no longer need to redo the work.
However, what follows is misunderstood. In the remaining 20%—especially in situations where the translation has a specific "aim"—current systems still do not work consistently. Specifically, two things happen.
One is the fluctuation of translation intent. Even if you throw the same source text into the same model twice and instruct it to "translate as a speech" both times, the two resulting translations will not have a consistent tone. One might be inflammatory while the other is somewhat restrained. The landing point relative to the aim is only determined probabilistically.
The other is consistency in long texts. A document might translate "contract" as "keiyaku" in the first half, but it turns into "keiyakusho" in the second half. A tone that was formal in the first half might shift subtly to casual in the second half. Proper nouns might appear in three different notations. Individually, these are not fatal. But for the document as a whole, they chip away at quality like body blows.
Furthermore, these problems are hard to notice just by reading the output. If you look at each sentence, they are all "proper sentences." The source of the discomfort occurs at the document level, not the sentence level. Readers who don't feel discomfort feel safe hitting the send button. Whether that is a state they should feel safe in is another matter.
The same structural problems occur with human translators. Fatigue or fluctuations in judgment creep in during a long document. Details of genre or register may not perfectly align between the first and last pages. Veterans are aware of this, which is why they spend time on self-review. However, there is also no way to verify from the outside whether that self-review is complete.
In short, whether it's a machine or a human, the industry as a whole lacks a mechanism to independently verify whether the quality of a translation is consistent "relative to the aim" and "throughout the entire piece." This is the current situation.
The Metrics We Have Don't Measure What We Care About
It's not that there are no verification mechanisms. There are. The problem is what they are measuring.
The two main automatic evaluation metrics in machine translation research are BLEU and COMET. BLEU compares the output with a reference translation and counts overlapping word sequences. COMET uses a pre-trained model to score semantic similarity. Both are useful for their intended purposes. And both share the same premise that engineering cannot escape: the premise that a "correct answer" exists, and that correct answer is the reference translation placed in the test set.
Translation as an activity does not work that way.
Let's borrow a sentence from Japanese children's literature: "Ano otokonoko wa marude Momotaro mitai da." For someone raised in Japan, Momotaro is a hero of folklore who defeats ogres with a dog, a monkey, and a pheasant. Calling a child "like Momotaro" carries the nuance that they are brave for their age, have guts, and shouldn't be underestimated even if they are small.
If you translate this for an English reader who has never heard of Momotaro, there are multiple options.
You could write it literally: "That boy is just like Momotaro." The surface of the original is perfectly preserved. If the reference translation happens to be the same, BLEU will be delighted. But almost nothing will be conveyed to the English reader. The meaning of the sentence is entirely locked inside a name they don't know.
You could also write: "That boy is so brave for his age." You discard the cultural specificity, but the meaning lands instantly.
Or you could take a bold step: "That boy's another little Hulk." This is a daring choice. It replaces a culturally incomprehensible reference with another culturally comprehensible one. It is an adaptation that transplants the "function" rather than the "content" of the reference. Depending on the brief (the order details), this could be the best move or an inappropriate one.
Which is the correct answer? It depends on the conditions. It depends on the reader, the medium, the allowed range of discretion, the register of the surrounding text, and about twelve other factors. And all those factors exist above the sentence level, invisible to metrics that only look at sentences.
BLEU praises choices that happen to match the reference translation and punishes everything else, regardless of which is superior. COMET ranks by semantic distance and skips the entire question of what effect the translation should have on the reader. Neither tells you "which choice fits this job." In the first place, the idea of replacing it with the Hulk is a leap a human translator might make, but a metric trained on proximity to a reference translation will never reward it.
Let me say one important thing about translation. There is no single correct answer. For every line, valid options open up like a fan. Which one to take is the translator's judgment. "Like Momotaro," "brave for his age," "a little Hulk"—all are defensible choices. The job of evaluation is not to guess the one and only correct translation. It is to look at the judgment the translator made and frankly state whether it functions for this job.
What Happens When You Ask an LLM Directly
If you want to do it the modern way, you can skip metrics and ask an LLM directly. "Here's the source, here's the translation—how is it?"
This works better than you'd think, and worse than you'd hope.
It works better than you'd think because LLMs can, in principle, reason about things like register, audience, cultural references, and rhetorical effects—things BLEU and COMET cannot reach. It can notice when a sentence has lost its rhythm. It can point out that the Momotaro translation is opaque.
The reasons it works worse than you'd hope are twofold, and in practice, they have a compounding effect.
First, the evaluation axes shift. If you ask the same question to the same model twice, it will return answers with different dimensional configurations. The first time might focus on fluency, the second on accuracy, and the third might invent a new category halfway through. You cannot compare evaluations across multiple documents because what is being measured isn't consistent to begin with. While you are measuring, the instrument itself is moving.
Second is sycophancy (flattery). LLMs are trained quite strongly to please their conversation partner. If you hand over a translation and ask "Is this good?", there is a high probability it will say "It's good." If you push back, it will agree with your pushback. The model is optimizing for "pretending to listen and respect the user," not for "being a cold-blooded evaluator." For low-risk uses, this is fine. But for someone shipping translation as a product, someone grading translation, or someone buying translation with real money, this trait is the exact opposite of what is needed.
As a result, a strange situation has persisted. The thing we wanted—a means for non-experts to actually verify translations—did not properly exist. If you outsourced translation, you just had to trust the vendor. If you used machine translation, you just had to cross your fingers and pray. The only people who could reliably check a translation were senior reviewers who had already mastered both languages to a level where they didn't need the translation in the first place.
What CATER Is Trying to Do
There is a tool called CATER. I am on the development side of it. I believe this is the first serious attempt at a solution to this problem, so let me explain what it does.
CATER evaluates translation on six explicit axes: Grammatical Precision (GP), Semantic Integrity (SI), Factual Consistency (FC), Terminology Consistency (TC), Discourse Coherence (DC), and Communicative & Stylistic Appropriateness (CSA). The axes do not move from run to run. They are the same every time. If you compare Evaluation A and Evaluation B, that comparison has meaning.
The scoring itself is not left to the LLM. The model is responsible for identifying and characterizing errors, and a deterministic pipeline calculates numerical values based on severity, mandatory strength, and axis sensitivity. If you run the same input twice, you get the same numbers. This might sound like a minor implementation detail. It isn't. This is the line that separates an "instrument" from an "atmosphere."
And what I want to spend a little time talking about is the Translation Brief.
The brief tells CATER what the translation is for. Who is the reader, what is the medium, what effect should it produce, what can be sacrificed, and what must absolutely be preserved? With a brief, evaluation is no longer a question of "how close is this to some abstract correct answer?" It becomes the question: "Is this translation doing the job it was hired to do?" As far as I know, this is the only question that truly matters.
When you provide a brief, the Momotaro example is resolved. If the brief is "Children's book for American kids, readability is top priority," then "That boy is just like Momotaro" gets flagged on the CSA axis because the reference doesn't land. "That boy's another little Hulk" could get a high score. If the brief is "Literary translation for an academic anthology, preserve cultural specificity," the judgment is reversed. Keeping the Momotaro line as in the original is correct, and replacing it with the Hulk is excessive domestication. Same source, same options, different brief, different correct answer. The evaluator can see this because the brief is a first-class input.
If you don't provide a brief, CATER supplements it with inference. In actual practice, translations with a written brief are in the minority. But there is always an implicit brief. Genre, register, and target audience constrain the boundaries of a "good translation." A skilled reviewer reads while automatically reconstructing the brief in their head. CATER does the same thing explicitly and displays the inferred brief on the screen. If it's wrong, a human can correct it.
Applying CATER to a Nobel-Prize-Level Translation
Let me give a concrete example. This is not machine translation output. I'll shift the talk to human literary translation from long before machine translation appeared.
Yasunari Kawabata's "Snow Country," translated by Edward Seidensticker. First translated in 1956 and later revised, this Seidensticker translation was what the selection committee referred to when Kawabata won the Nobel Prize in Literature in 1968. It can be called the 20th-century pinnacle of English translation of Japanese literature. It is quite difficult to find a human translation that surpasses this.
I ran its opening paragraph through CATER. No explicit brief was provided; I let the evaluator infer it.
Overall score: 58.8/100. The judgment was "Major rework required." Here is the breakdown by axis:
Please don't jump to conclusions. CATER is not saying "Seidensticker is bad." Grammar, facts, and logical structure all get perfect scores. What is broken are the core meaning and literary effects—limited but critical parts. And CATER points exactly to where that breakdown is.
"Yoru no soko ga shiroku natta" (The bottom of the night turned white) is translated by Seidensticker as "The earth lay white under the night sky." CATER's diagnosis is this: the metaphorical and perceptual expression "bottom of the night" has been replaced by a different scene, "the ground is white under the night sky." The original's semantic effect, where whiteness rises from within the night, is not preserved. As a minimum correction, "The bottom of the night turned white" is suggested. This is the main reason the SI axis dropped to 25.
One more. A scene where a girl opening a window calls out "as if shouting into the distance, 'Station master! Station master!'" Seidensticker's translation is: "Leaning far out the window, the girl called to the station master as though he were a great distance away." The direct speech itself has been replaced by a summary description. CATER's comment: "The resonance of the call itself and the immediacy of the scene are lost. Because the spoken content was erased, the sense of presence as a literary reproduction is weakened." This brought the CSA axis to 0.
The correct way to take these points is that Seidensticker must have had his own reasons for these choices. As a literary translation for English-speaking readers in the 1950s, he consciously chose this within the translation norms of the time. CATER's comment itself does not call this a "mistranslation" but treats it as "an issue of interpretation where the judgment changes depending on the brief." However, if you commissioned the same source text for a modern Japanese literature translation project and wrote a brief saying "Prioritize the reproduction of the original's poetic atmosphere," the judgment would be that this translation is not doing its job. Same translation, different brief, different conclusion.
What I want you to take away from this is not the score itself, but three things.
First: Even a translation that leaves its mark on world literary history can receive a "rework" judgment under a specific brief. This is proof that there is no absolute ceiling to translation quality, and at the same time, a demonstration that evaluation is a task relative to a brief.
Second: CATER doesn't just say "this is bad," it provides specific alternatives for "fix it like this." Diagnosis and prescription are integrated. Because of this, the side that sees the evaluation results can take the next step.
Third: Even if grammar, facts, and logical structure are perfect, the job as a literary translation can fail. Failures that cannot be caught by "reading it and feeling no discomfort" properly show up on different axes. How weak a verification it is to check machine translation output by saying "it was okay because I didn't feel discomfort"—this becomes visible by calculating backward from the fact that even a Seidensticker-level translation wavers when broken down into explicit axes.
Why This Is Important Even for Non-Translation Researchers
Most of what I've said so far might sound like inside baseball to those outside the localization industry. Here is why it is still important.
Translation is currently one of the primary uses of LLMs. The number of companies and the volume of traffic passing customer support, product copy, internal notices, legal documents, and contracts through machine translation pipelines are at a scale that didn't exist three years ago. The economic activity riding on these pipelines is enormous. The verification layer sitting on top of them is almost zero. People are shipping translations they cannot verify. They are sending out legal clauses praying the model hasn't fabricated numbers, hasn't weakened "must" to "should," and hasn't dropped the tone that made the marketing copy work.
And I repeat, this is not a problem with machine translation. Translations outsourced to professional translators are also being delivered in the same absence of verification. The client reads the delivered translation and confirms "it reads like Japanese." That is not a quality check of the translation; it is a confirmation that the delivery is in Japanese. The long distance that should exist between the two has never been traversed by anyone until now.
Automatic translation in 2026 is at a level that is "practically sufficient" for most uses. This is true. But "sufficient for most uses" is not a verification strategy. In situations where mistakes can be fatal—clinical documents, contracts, public statements, literary translations—"I didn't feel discomfort reading it" or "I asked a reliable translator" are no longer acceptable answers.
What we have long lacked is a layer for non-experts to verify output. Not a single number. Not a rubber stamp. A structured, readable diagnosis that says, "The strength of this translation is here, the risk is here, and if you value X, fix this part."
That is CATER. You can try it for free at cater.erudaite.ai. The methodology and the theory behind it are available at about.erudaite.ai.
Try throwing in a translation you have on hand—for example, an email before sending, internal materials, a draft contract, or a manuscript just delivered from an outsource—and see what comes out.
You should be able to confirm its quality and the characteristics of the translation with CATER.





