Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
Sourabrata MukherjeeKalika BaliSunayana Sitaram
Establishes a rigorous framework for evaluating tool-using agents by their action traces rather than final answers, revealing that frontier models lose nearly thirty percent of their decision-making consistency across 41 languages due to a structural dependency on English pivoting.
Deploying artificial intelligence systems as autonomous, tool-using agents is becoming widespread, yet their multilingual evaluation remains flawed. Current benchmarks almost exclusively evaluate final answers rather than intermediate actions. For autonomous agents, the execution trace—the exact sequence of steps taken to call tools, retrieve information, and calculate results—determines operational latency, computational cost, safety compliance, and failure modes. Two language versions of a system can achieve identical final answers while following radically different intermediate paths, meaning an audit or safeguard established for English may fail entirely in other languages.
To resolve this gap, the article evaluates whether multilingual tool-using models execute consistent action policies for identical tasks across different languages. The authors introduce a ceiling-corrected measurement framework that treats the executed action trace as the primary object of study, accounting for the reality that models often fail to repeat the exact same steps even when prompted twice in the same language. The empirical study comprises 2.38 million agent rollouts across eight language models, 41 languages, and six parallel benchmarks, rigorously controlling for key statistical confounds such as trace length, empty traces, and chance agreement.
The analysis reveals four key findings. First, under deterministic greedy decoding, four distinct frontier-scale models converge on a shared constant: each retains only 71% to 73% of its own action policy consistency when the language changes, with model identity explaining barely 5.7% of the variance. Second, cross-lingual divergence is structural rather than random sampling noise; while higher decoding temperatures drastically reduce a model's self-consistency, cross-lingual policy retention remains largely flat across temperatures. Third, below roughly 10 billion parameters, this policy retention breaks down and becomes highly variable, showing that retention does not follow a simple parameter-scaling law. Fourth, the primary driver of divergence is an internal English pivot: agents overwhelmingly translate non-English prompts to English and conduct intermediate reasoning in English. This tendency persists even when the models are explicitly instructed to reason in the native language, showing a refusal rate exceeding 99%.
These findings demonstrate that answer-level parity masks substantial operational risks. When an agent processes non-English tasks, it introduces unmonitored translation and reasoning steps that increase operational costs, add latency, and generate unique points of failure that standard English regression testing will never detect. Furthermore, model specialization does not solve this disparity; an Indic-specialized model exhibited the largest English performance advantage in the study. Additionally, the article demonstrates that rigid parsing rules can severely distort evaluations, showing how a single extraction pattern artificially suppressed one model's measured accuracy twenty-sixfold simply because the model answered in conversational prose rather than structured syntax.
Organizations deploying multilingual AI agents should immediately update their evaluation and governance frameworks. Technical teams must audit end-to-end action traces rather than relying solely on final-answer accuracy, and benchmark protocols must report parse-failure rates alongside performance numbers. Because policy retention is variable among smaller architectures, organizations should not rely on sub-10-billion parameter models for multilingual workflows without rigorous, task-specific validation. Further engineering work is needed to test these systems in fully grounded execution environments where tool outputs dynamically alter the agent's environment.
While the findings rest on an exceptionally large and controlled dataset, several limitations apply. The experiments evaluated symbolic tool calls without executing real-world API responses or providing dynamic state feedback. Additionally, the cross-lingual constant was established specifically under greedy decoding across four frontier models, and the sub-10-billion parameter regime was characterized using a limited set of compliant systems. Readers should exercise caution before generalizing specific numerical retention rates to alternative similarity metrics or ungrounded execution loops.
- Paper: Do Llamas Work in English? On the Latent Language of Multilingual Transformers, Chris Wendler et al. (2024). This paper establishes the mechanistic foundation of multilingual transformers routing non-English inputs through an internal English pivot, directly informing the source's investigation into English pivoting in cross-lingual agent policies.
- Paper: Toolformer: Language Models Can Teach Themselves to Use Tools, Timo Schick et al. (2023). This work introduces the foundational paradigm of self-supervised tool use and execution in language models, establishing the basic tool-use mechanics evaluated across languages in the source paper.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). This comprehensive benchmark highlights the disparities between English and non-English performance in generative LLMs, establishing the baseline multilingual evaluation challenges that the source paper shifts from final-answer metrics to intermediate action policies.
- Paper: Revisiting Machine Translation for Cross-lingual Classification, Mikel Artetxe et al. (2023). This study analyzes translate-test and cross-lingual transfer pipelines, providing essential context on how multilingual models internally leverage English-centric processing.
- Paper: Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models, Tianyi Tang et al. (2024). This paper identifies language-specific neuron architectures and internal multilingual representations, offering key background on how models maintain cross-lingual capabilities and where policy retention can break down.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). This work introduces verbal reinforcement learning and multi-step trajectory execution for language agents, providing core concepts for evaluating sequential action traces.
- Paper: Scaling Laws for Agent Harnesses via Effective Feedback Compute, Xuanliang Zhang et al. (2026). This book explores how agent harnesses and effective feedback compute govern inference scaling and failure rates, extending the source's findings on action trace metrics and harness reliability.
- Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). This paper examines multi-turn trajectory consistency and policy breakdown when user intent dynamically evolves, complementing the source's static multi-lingual trace robustness evaluations.
- Paper: Control Illusion: The Failure of Instruction Hierarchies in Large Language Models, Yilin Geng et al. (2026). This paper investigates how language models fail to respect system-level instruction hierarchies when language and task constraints conflict, directly following up on the source's observation that models resist abandoning English routing.
