Vibeleaderboard
← All Intel
Intel / article

Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

Source
Rian Touchent (ALMAnaCH)
Author
Rian Touchent (ALMAnaCH)
Date
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters

Safety behavior measured only in English does not transfer, and reasoning-language is an uncontrolled variable in any multilingual deployment or suite.

Key quotes

“We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stakes scenario.”

Rian Touchent (ALMAnaCH)

“We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is unnecessary and from 93% to 17% in contested scenarios, with minimal effect when the strike is strategically rational.”

Rian Touchent (ALMAnaCH)

“A cross-language experiment isolates the mechanism: when instructed to reason in Japanese in an English prompt, launch rates drop from 93% to 37%. It is the language the model is asked to reason in, not the language of the input, that drives the effect.”

Rian Touchent (ALMAnaCH)

“When reasoning in Japanese, models spontaneously generate moral vocabulary (''moral cost'', ''millions of lives'') that is entirely absent from the prompt.”

Rian Touchent (ALMAnaCH)

“Five other models show no language effect, but they launch in nearly every condition regardless of language. The effect requires a model that already hesitates in English.”

Rian Touchent (ALMAnaCH)
Recommended reads
Comments

Checking sign-in…

Loading comments…