Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

Sourabrata MukherjeeKalika BaliSunayana Sitaram

article2026arXiv1 citations

Establishes a rigorous framework for evaluating tool-using agents by their action traces rather than final answers, revealing that frontier models lose nearly thirty percent of their decision-making consistency across 41 languages due to a structural dependency on English pivoting.

Listen

Deploying artificial intelligence systems as autonomous, tool-using agents is becoming widespread, yet their multilingual evaluation remains flawed. Current benchmarks almost exclusively evaluate final answers rather than intermediate actions. For autonomous agents, the execution trace—the exact sequence of steps taken to call tools, retrieve information, and calculate results—determines operational latency, computational cost, safety compliance, and failure modes. Two language versions of a system can achieve identical final answers while following radically different intermediate paths, meaning an audit or safeguard established for English may fail entirely in other languages.

To resolve this gap, the article evaluates whether multilingual tool-using models execute consistent action policies for identical tasks across different languages. The authors introduce a ceiling-corrected measurement framework that treats the executed action trace as the primary object of study, accounting for the reality that models often fail to repeat the exact same steps even when prompted twice in the same language. The empirical study comprises 2.38 million agent rollouts across eight language models, 41 languages, and six parallel benchmarks, rigorously controlling for key statistical confounds such as trace length, empty traces, and chance agreement.

The analysis reveals four key findings. First, under deterministic greedy decoding, four distinct frontier-scale models converge on a shared constant: each retains only 71% to 73% of its own action policy consistency when the language changes, with model identity explaining barely 5.7% of the variance. Second, cross-lingual divergence is structural rather than random sampling noise; while higher decoding temperatures drastically reduce a model's self-consistency, cross-lingual policy retention remains largely flat across temperatures. Third, below roughly 10 billion parameters, this policy retention breaks down and becomes highly variable, showing that retention does not follow a simple parameter-scaling law. Fourth, the primary driver of divergence is an internal English pivot: agents overwhelmingly translate non-English prompts to English and conduct intermediate reasoning in English. This tendency persists even when the models are explicitly instructed to reason in the native language, showing a refusal rate exceeding 99%.

These findings demonstrate that answer-level parity masks substantial operational risks. When an agent processes non-English tasks, it introduces unmonitored translation and reasoning steps that increase operational costs, add latency, and generate unique points of failure that standard English regression testing will never detect. Furthermore, model specialization does not solve this disparity; an Indic-specialized model exhibited the largest English performance advantage in the study. Additionally, the article demonstrates that rigid parsing rules can severely distort evaluations, showing how a single extraction pattern artificially suppressed one model's measured accuracy twenty-sixfold simply because the model answered in conversational prose rather than structured syntax.

Organizations deploying multilingual AI agents should immediately update their evaluation and governance frameworks. Technical teams must audit end-to-end action traces rather than relying solely on final-answer accuracy, and benchmark protocols must report parse-failure rates alongside performance numbers. Because policy retention is variable among smaller architectures, organizations should not rely on sub-10-billion parameter models for multilingual workflows without rigorous, task-specific validation. Further engineering work is needed to test these systems in fully grounded execution environments where tool outputs dynamically alter the agent's environment.

While the findings rest on an exceptionally large and controlled dataset, several limitations apply. The experiments evaluated symbolic tool calls without executing real-world API responses or providing dynamic state feedback. Additionally, the cross-lingual constant was established specifically under greedy decoding across four frontier models, and the sub-10-billion parameter regime was characterized using a limited set of compliant systems. Readers should exercise caution before generalizing specific numerical retention rates to alternative similarity metrics or ungrounded execution loops.

arXiv: 2608.11110
  • Paper: Scaling Laws for Agent Harnesses via Effective Feedback Compute, Xuanliang Zhang et al. (2026). This book explores how agent harnesses and effective feedback compute govern inference scaling and failure rates, extending the source's findings on action trace metrics and harness reliability.
  • Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). This paper examines multi-turn trajectory consistency and policy breakdown when user intent dynamically evolves, complementing the source's static multi-lingual trace robustness evaluations.
  • Paper: Control Illusion: The Failure of Instruction Hierarchies in Large Language Models, Yilin Geng et al. (2026). This paper investigates how language models fail to respect system-level instruction hierarchies when language and task constraints conflict, directly following up on the source's observation that models resist abandoning English routing.
Cover for Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

Abstract

When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system fails, and are the only auditable part of its behaviour. We make the action policy the measured object across 8 models, 6 parallel benchmarks and 41 languages (2.38M rollouts). The naive measurement fails: five confounds sit between raw trace similarity and any defensible claim, each able to flip a conclusion. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, the gap is capped by each model's reproducibility, and a model asked the same question twice in one language answers differently, leaving no baseline. We remove all five, and every correction makes the effect larger. Divergence proves structural, not sampling noise: it survives greedy decoding in every cell and stays flat as temperature rises, even as models grow less self-consistent. Normalised by their own reproducibility, four very different frontier models converge under greedy decoding, each keeping 71-73% of its action policy across languages, with model identity explaining only 5.7% of the variance. Below roughly 10B parameters it breaks down, and the ordering among smaller models is largely an artifact of a chance floor we measure by permutation rather than assume. Agents route non-English tasks through English; this pivot is causally load-bearing, confirmed by a pre-registered prediction across four models, and models will not abandon it when told to. Finally, a single trace-extraction regex, not the model, manufactured a multilingual failure: two worked examples raise one model's measured accuracy twenty-sixfold while its accuracy on readable outputs barely moves.

Table of Contents

  • 1 Introduction
  • 2 Experimental Setup
  • 3 Cross-Lingual Policy Divergence
  • 4 Policy Retention and Its Boundary
  • 5 The English Pivot
  • 6 Outcomes and Measurement Validity
  • 7 Conclusion
  • References
  • A Key Terms

Knowls

  1. Knowl 1 — Normalised Cross-Lingual Policy Retention Estimand

    model/method

    Cross-lingual evaluation of tool-using language model agents requires evaluating the executed sequence of tool actions rather than only the final output. Given a canonical task zz presented in language ℓ\ell as input x(ℓ)(z)x^{(\ell)}(z), an agent induces an action trace A(ℓ)(z)=⟨τ1,…,τk⟩A^{(\ell)}(z) = \langle \tau_1, \dots, \tau_k \rangle over a fixed symbolic tool alphabet T={Search,Calc,Translate,Summarize,Finish}\mathcal{T} = \{\text{Search}, \text{Calc}, \text{Translate}, \text{Summarize}, \text{Finish}\}.

    To make cross-lingual policy divergence identifiable, the measurement must account for the model's stochastic self-inconsistency (since within-language trace reproducibility Iwithin<1I_{\text{within}} < 1). The protocol generates two independent replicates r1,r2r_1, r_2 per task-language instance under identical decoding configurations, varying only the serving seed. Trace similarity S(A,B)∈[0,1]S(A, B) \in [0, 1] is computed as the normalised matching-block sequence similarity:

    S(A,B)=2M∣A∣+∣B∣S(A, B) = \frac{2M}{|A| + |B|}

    where MM is the sum of lengths of matched contiguous blocks under longest-contiguous-match decomposition.

    Within-language reproducibility IwithinI_{\text{within}} and cross-language agreement IcrossI_{\text{cross}} are both computed across the two replicates across all tasks zz and language pairs ℓi≠ℓj\ell_i \neq \ell_j:

    Iwithin=Ez,ℓ[S(Ar1(ℓ)(z),Ar2(ℓ)(z))],Icross=Ez,ℓi≠ℓj[S(Ar1(ℓi)(z),Ar2(ℓj)(z))]I_{\text{within}} = \mathbb{E}_{z, \ell} \left[ S\left(A_{r_1}^{(\ell)}(z), A_{r_2}^{(\ell)}(z)\right) \right], \quad I_{\text{cross}} = \mathbb{E}_{z, \ell_i \neq \ell_j} \left[ S\left(A_{r_1}^{(\ell_i)}(z), A_{r_2}^{(\ell_j)}(z)\right) \right]

    The ceiling-corrected metric, termed normalised policy retention I~\tilde{I}, is the ratio:

    I~=IcrossIwithin\tilde{I} = \frac{I_{\text{cross}}}{I_{\text{within}}}

    representing the fraction of an agent's self-reproducibility preserved across languages. To eliminate length and empty-trace artifacts, pairs where either trace is empty are dropped, and all comparisons are reweighted to match trace-length distributions over max⁡(∣A∣,∣B∣)\max(|A|, |B|) bins in both directions.

  2. Knowl 2 — Convergence of Cross-Lingual Policy Retention in Frontier Models

    empirical result

    Under deterministic greedy decoding (T=0T = 0), four structurally diverse frontier language models converge to nearly the same normalised cross-lingual policy retention I~\tilde{I} across 6 multilingual parallel benchmarks (FLORES-200, XQuAD, XNLI, Belebele, XCOPA, and a synthetic benchmark):

    Model Architecture Scale IwithinI_{\text{within}} Pooled I~\tilde{I}
    Gemma-3-27B Dense 27B 0.9060 0.7276
    Sarvam-M Dense (Indic-specialised) 24B 0.8632 0.7076
    Qwen3-235B-A22B Mixture-of-Experts 235B (22B active) 0.6787 0.7331
    Llama-4-Maverick Mixture-of-Experts 17B active (128 exp) 0.9005 0.7092

    Key empirical findings:

    1. The pooled policy retention of the four models falls within a tight 2.6 percentage point band: I~∈[0.708,0.733]\tilde{I} \in [0.708, 0.733] (mean I~=0.722\tilde{I} = 0.722).
    2. In a variance decomposition across all n=24n = 24 (model ×\times benchmark) cells, model identity explains only 5.7%5.7\% of the variance in I~\tilde{I}, while benchmark identity explains 26.9%26.9\%, and residual variance is 67.4%67.4\% (overall standard deviation is 0.0520.052).
    3. This cross-model invariance collapses under stochastic decoding (T=0.5T = 0.5), where model identity explains 74.8%74.8\% of variance due to differential model self-inconsistency, widening the cross-model spread to 13.613.6 percentage points.
  3. Knowl 3 — Temperature Invariance of Cross-Lingual Policy Agreement

    empirical result

    Testing tool-using agents across a temperature ladder T∈{0.0,0.3,0.5,0.7,1.0}T \in \{0.0, 0.3, 0.5, 0.7, 1.0\} reveals that cross-lingual policy agreement (IcrossI_{\text{cross}}) is invariant to decoding temperature, changing by at most Δ=0.009\Delta = 0.009 across the entire range (e.g., for Gemma-3-27B, raw IcrossI_{\text{cross}} remains between 0.65460.6546 and 0.66350.6635).

    In contrast, same-language reproducibility (IwithinI_{\text{within}}) falls steeply as temperature increases (moving 21×21\times faster than IcrossI_{\text{cross}}). Consequently, the raw gap Δ=Iwithin−Icross\Delta = I_{\text{within}} - I_{\text{cross}} narrows at higher temperatures purely due to the collapse of the within-language reproducibility ceiling IwithinI_{\text{within}}, rather than any genuine cross-lingual policy convergence. Cross-lingual divergence is thus an intrinsic property of the model's structural mapping from language to actions, rather than sampling noise.

  4. Knowl 4 — Methodological Confounds in Agent Trace Similarity Measurement

    model/method

    Five distinct measurement confounds distort raw sequence similarity comparisons of agent tool traces, several of which can reverse published model rankings or intervention effects:

    1. C1 (Missing Baseline / Identification Failure): Models are not fully self-consistent (Iwithin∈[0.63,0.80]I_{\text{within}} \in [0.63, 0.80] under sampling). Comparing cross-language runs without a matched same-language baseline confounds language divergence with decoding noise. Matching token budgets alone shifts the pooled gap from +0.0625+0.0625 to +0.0878+0.0878.
    2. C2 (Trace-Length Sensitivity): Normalized matching-block sequence similarity scores shorter traces systematically higher (R2=0.67R^2 = 0.67 with irreproducibility). Fixed by bidirectional reweighting over length bins.
    3. C3 (Empty Traces): Completely unparseable or blank traces match each other with a similarity of 1.01.0, artificially rewarding protocol failure. Strict pair-dropping lowers headline invariance by up to −0.119-0.119 on individual cells.
    4. C4 (Reproducibility Ceiling): The raw gap Iwithin−IcrossI_{\text{within}} - I_{\text{cross}} is tightly bounded by IwithinI_{\text{within}} (r=+0.97r = +0.97 at T=0T = 0), so uncorrected rankings measure model determinism rather than multilinguality, flipping rank orderings between T=0.5T = 0.5 and T=0T = 0 (Spearman ρ=−0.80\rho = -0.80).
    5. C5 (Chance Floor): Unrelated traces over a 5-tool alphabet agree well above zero (c≈0.56c \approx 0.56, ranging from 0.350.35 on long traces to 0.950.95 on short traces). Measured via task-permutation null distributions.
  5. Knowl 5 — Causal Role and Head-Room Ordering of the English Pivot

    empirical result

    Multilingual agents predominantly execute non-English tasks by pivoting through English: Translate is the most frequently called tool across all adapted benchmarks (e.g., 50.5%50.5\% of all calls in Qwen3-235B vs 18.0%18.0\% in Sarvam-M), and reasoning thoughts remain ≈99%\approx 99\% ASCII even on Indic script inputs.

    Causal interventions on the tool alphabet demonstrate:

    1. Tool Removal: Deleting Translate leads to tool substitution (shifting to Summarize in Qwen3 and Llama-4, and Search in Sarvam). Length-matched cross-lingual agreement drops on 4/64/6 benchmarks for models that pivot heavily, while having negligible effect (1/61/6 benchmarks negative) on Sarvam-M, confirming a dose-response mechanism.
    2. Tool Mandate: Mandating Translate as the first action on non-English inputs improves length-matched agreement in a pre-registered order strictly determined by a model's baseline "head-room" (100−baseline first-action Translate rate100 - \text{baseline first-action Translate rate}):
      • Sarvam-M (57.257.2 headroom): +38.6%+38.6\% first-action rate move, positive in 5/65/6 benchmarks.
      • Gemma-3-27B (30.430.4 headroom): +30.0%+30.0\% first-action rate move, positive in 5/65/6 benchmarks.
      • Llama-4-Maverick (18.718.7 headroom): +17.4%+17.4\% first-action rate move, positive in 3/63/6 benchmarks.
      • Qwen3-235B (15.715.7 headroom): +14.7%+14.7\% first-action rate move, positive in 2/62/6 benchmarks.
  6. Knowl 6 — Prompt Resistance of Internal English Reasoning

    empirical result

    When explicitly instructed via system prompts to write intermediate thoughts (Thought: ...) in the language and script of the task rather than English, multilingual agents overwhelmingly refuse compliance:

    • Gemma-3-27B generated non-Latin script thoughts on only 0.79%0.79\% of non-Latin script rollouts.
    • Sarvam-M generated non-Latin script thoughts on only 0.08%0.08\% of non-Latin script rollouts.
    • The mean ASCII character fraction in generated reasoning text remained 0.9860.986 to 0.9930.993 across all conditions.

    This confirms that latent internal English reasoning is an ingrained architectural/representational feature of multilingual LLMs that cannot be overridden by prompting constraints.

  7. Knowl 7 — Breakdown of Policy Retention Regularity Below 10B Parameters

    empirical result

    The 2.62.6 percentage point retention band observed across 17B–235B frontier models breaks down on models under 10B parameters. Across three evaluated smaller systems at T=0T = 0:

    Model Scale Pooled Raw I~\tilde{I} Macro-Ratio Raw I~\tilde{I} Chance-Corrected I~\tilde{I}
    Frontier Band ≥17B\ge 17\text{B} [0.708,0.733][0.708, 0.733] [0.713,0.742][0.713, 0.742] [0.146,0.176][0.146, 0.176]
    Gemma-3-4B 4B 0.6259 0.654 0.155
    Qwen3-8B 8B 0.8016 0.735 0.109
    Aya-Expanse-8B 8B 0.4738 0.446 0.047

    Key observations:

    1. Variance Expansion: The spread across the clean small models (Gemma-3-4B and Qwen3-8B) is 17.617.6 points (standard deviation 0.0880.088), an 8.0×8.0\times increase in standard deviation over the frontier models (SD 0.0110.011).
    2. Chance-Correction Reversal: Qwen3-8B produces short traces with a massive chance collision floor (c=0.669c = 0.669 overall, up to 0.9470.947 on Belebele). Chance correction κ=(S−c)/(1−c)\kappa = (S - c)/(1 - c) reverses the apparent raw ordering: Gemma-3-4B moves into the frontier band (0.1550.155) while Qwen3-8B drops below it (0.1090.109), contradicting a naive scaling hypothesis.
  8. Knowl 8 — Causal Trace-Length Manipulation and Policy Retention

    empirical result

    A 3-level prompt manipulation constraining trace lengths (prompts specifying ≤2\le 2 actions, ≥5\ge 5 actions, and ≥8\ge 8 actions) was evaluated at T=0T = 0 on 72 cells (271,200271,200 rollouts) across Gemma-3-27B and Qwen3-235B.

    The intervention achieved a monotonic ≈3×\approx 3\times expansion in mean executed actions per rollout (Gemma: 2.86→6.32→9.092.86 \to 6.32 \to 9.09; Qwen3: 3.45→5.67→8.683.45 \to 5.67 \to 8.68). This length change caused substantial shifts in policy retention I~\tilde{I} with mutually disjoint 95%95\% bootstrap confidence intervals:

    • Gemma-3-27B: I~\tilde{I} shifted by 6.06.0 points (0.7689→0.8292→0.81260.7689 \to 0.8292 \to 0.8126).
    • Qwen3-235B: I~\tilde{I} decreased monotonically by 6.96.9 points (0.8078→0.7722→0.73920.8078 \to 0.7722 \to 0.7392).

    Because the shift (6.06.0–6.96.9 points) exceeds the across-model frontier band (2.62.6 points), this proves trace length is a causal driver of policy retention that is not eliminated by computing the ratio I~=Icross/Iwithin\tilde{I} = I_{\text{cross}} / I_{\text{within}}.

  9. Knowl 9 — Parser Sensitivity Artifacts vs. Model Capability (Legibility Illusion)

    empirical result

    Single-regex trace extraction harnesses can create severe measurement failures that mimic model capability collapse. In zero-shot evaluation, GPT-OSS-120B failed regex extraction on 74.4%74.4\% of rollouts (due to outputting valid natural language descriptions such as "We will use Translate" instead of exact syntax), yielding a measured task accuracy of 1.74%1.74\%, while simultaneously scoring artificially high raw invariance due to empty traces.

    Providing 2-shot format exemplars resolved syntax adherence without improving underlying reasoning ability:

    Prompting Condition Empty Trace Rate Parse Failure Measured Accuracy Accuracy ∣| Scorable
    0-shot (baseline) 74.4% 59.5% 0.0174 0.8127
    2-shot 28.7% 21.7% 0.4421 0.7574
    4-shot 27.8% 19.7% 0.4539 0.7418

    Measured accuracy increased 26×26\times (0.0174→0.44210.0174 \to 0.4421), whereas accuracy among successfully parsed rollouts remained flat (0.8127→0.75740.8127 \to 0.7574). The effect saturated at 2 exemplars, proving the intervention repaired syntactic legibility rather than task-solving competence.

  10. Knowl 10 — Decoupling of Agent Action Policy Invariance from Task Correctness

    empirical result

    Evaluation of gold answer correctness across four benchmarks (XQuAD, XNLI, Belebele, XCOPA) demonstrates that action policy invariance does not correlate strongly with task accuracy:

    1. English Advantage: All evaluated models achieve higher accuracy in English than non-English languages. Sarvam-M (an Indic-specialised 24B model) exhibits the largest English accuracy advantage (+0.155+0.155 mean accuracy gap over non-English languages), showing that language specialization does not eliminate agentic disparity.
    2. Accuracy Variation: Task accuracy varies widely across languages within a single model-benchmark cell (e.g., Qwen3 spans 0.3300.330 to 0.9900.990 on XCOPA; Sarvam spans 0.2780.278 to 0.7890.789 on XQuAD).
    3. Decoupling from Invariance: Among protocol-adherent models, the linear correlation between policy invariance and correctness is moderate (r=+0.378,R2=0.14r = +0.378, R^2 = 0.14) and reverses sign in 25%25\% of individual model-benchmark cells (e.g., Gemma on XNLI: r=−0.477r = -0.477; Llama-4 on Belebele: r=−0.669r = -0.669).
  11. Knowl 11 — Effect of Self-Consistency Voting on Policy Retention

    empirical result

    Applying self-consistency majority voting over action sequences using k=5k = 5 replicate votes (at T=0.7T = 0.7 on 300 tasks) increases raw cross-lingual sequence agreement IcrossI_{\text{cross}} by +0.052+0.052 to +0.053+0.053.

    However, majority voting increases same-language self-consistency IwithinI_{\text{within}} by a larger margin (+0.078+0.078 in Gemma-3-27B and +0.071+0.071 in Sarvam-M). Consequently, the ceiling-corrected policy retention I~=Icross/Iwithin\tilde{I} = I_{\text{cross}} / I_{\text{within}} decreases:

    • Gemma-3-27B: I~\tilde{I} falls from 0.84810.8481 (k=1k=1) to 0.83230.8323 (k=5k=5).
    • Sarvam-M: I~\tilde{I} falls from 0.92960.9296 (k=1k=1) to 0.91090.9109 (k=5k=5).

    Both shifts exhibit disjoint 95%95\% bootstrap intervals. Self-consistency voting acts as an effective variance reducer for raw sequence agreement, but does not improve cross-lingual policy retention relative to the model's own reproducibility ceiling.

Coverage note — None was omitted; all key theoretical definitions, measurement protocols, empirical convergence findings, causal mechanism tests, temperature ladders, failure taxonomy, and confound analyses are fully covered.

References

  1. 1.Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. MEGA: Multilingual evaluation of generative AI. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 4232–4267, 2023. doi: 10.18653/v1/2023.emnlp-main.258.
  2. 2.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4623–4637, 2020. doi: 10.18653/v1/2020.acl-main.421.
  3. 3.Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 749–775, 2024. doi: 10.18653/v1/2024.acl-long.44.
  4. 4.Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8:454–470, 2020. doi: 10.1162/tacl a 00317.
  5. 5.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2475–2485, 2018. doi: 10.18653/v1/D18-1269.
  6. 6.John Dang, Shivalika Singh, Daniel D’souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, Tom Kocmi, Florian Strub, Nathan Grinsztajn, Yannis Flet-Berliac, Acyr Locatelli, et al. Aya expanse: Combining research breakthroughs for a new multilingual frontier. arXiv preprint arXiv:2412.04261, 2024.
  7. 7.Bradley Efron and Robert J. Tibshirani. An Introduction to the Bootstrap. Number 57 in Monographs on Statistics and Applied Probability. Chapman & Hall, New York, 1993.
  8. 8.Gemma Team. Gemma 3 technical report. arXiv preprint arXiv:2503.19786, 2025.
  9. 9.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, YongBin Kang, M. Sohel Rahman, and Rifat Shahriyar. XL-Sum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 4693–4703, 2021. doi: 10.18653/v1/2021.findings-acl. 413.
  10. 10.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. XTREME: A massively multilingual multi-task benchmark for evaluating crosslingual generalisation. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp. 4411–4421, 2020.
  11. 11.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles, pp. 611–626, 2023. doi: 10.1145/3600006.3613165.
  12. 12.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. MLQA: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7315–7330, 2020. doi: 10.18653/ v1/2020.acl-main.653.
  13. 13.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Re, et al. Holistic evaluation of language models. ´ Transactions on Machine Learning Research, 2023. arXiv:2211.09110.
  14. 14.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. AgentBench: Evaluating LLMs as agents. In International Conference on Learning Representations, 2024. arXiv:2308.03688.
  15. 15.Gregoire Mialon, Cl ´ ementine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas ´ Scialom. GAIA: A benchmark for general AI assistants. In International Conference on Learning Representations, 2024. arXiv:2311.12983.
  16. 16.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M. Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. Crosslingual generalization through multitask finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15991–16111, 2023. doi: 10.18653/v1/2023.acl-long.891.
  17. 17.NLLB Team, Marta R. Costa-jussa, James Cross, Onur ` C¸ elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, et al. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672, 2022.
  18. 18.OpenAI. gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925, 2025.
  19. 19.Aaron Parisi, Yao Zhao, and Noah Fiedel. TALM: Tool augmented language models. arXiv preprint arXiv:2205.12255, 2022.
  20. 20.Edoardo Maria Ponti, Goran Glavas, Olga Majewska, Qianchu Liu, Ivan Vuli ˇ c, and Anna ´ Korhonen. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pp. 2362–2376, 2020. doi: 10.18653/v1/2020.emnlp-main.185.
  21. 21.Martin L. Puterman. Markov Decision Processes: Discrete Stochastic Dynamic Programming. Wiley Series in Probability and Statistics. John Wiley & Sons, New York, 1994.
  22. 22.Jirui Qi, Raquel Fernandez, and Arianna Bisazza. Cross-lingual consistency of factual ´ knowledge in multilingual language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 10650–10666, 2023. doi: 10.18653/ v1/2023.emnlp-main.658.
  23. 23.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. ToolLLM: Facilitating large language models to master 16000+ real-world APIs. In International Conference on Learning Representations, 2024. arXiv:2307.16789.
  24. 24.Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Shantadevi Khapra. Samanantar: The largest publicly available parallel corpora collection for 11 Indic languages. Transactions of the Association for Computational Linguistics, 10:145–162, 2022. doi: 10.1162/tacl a 00452.
  25. 25.Timo Schick, Jane Dwivedi-Yu, Roberto Dess`ı, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, volume 36, 2023. arXiv:2302.04761.
  26. 26.Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models’ sensitivity to spurious features in prompt design, or: How i learned to start worrying about prompt formatting. In International Conference on Learning Representations, 2024. arXiv:2310.11324.
  27. 27.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. Language models are multilingual chain-of-thought reasoners. In International Conference on Learning Representations, 2023. arXiv:2210.03057.
  28. 28.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations, 2023. arXiv:2203.11171.
  29. 29.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pp. 24824–24837, 2022.
  30. 30.Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West. Do llamas work in English? on the latent language of multilingual transformers. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15366–15394, 2024. doi: 10.18653/v1/2024.acl-long.820.
  31. 31.An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
  32. 32.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations, 2023. arXiv:2210.03629.
  33. 33.Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. Improving massively multilingual neural machine translation and zero-shot translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 1628–1639, 2020. doi: 10.18653/v1/2020.acl-main.148.
  34. 34.Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, and Lidong Bing. How do large language models handle multilingualism? In Advances in Neural Information Processing Systems, volume 37, 2024. arXiv:2402.18815.

Citation

MLA
Mukherjee, S., et al. “Actions Speak Louder Than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents”. arXiv, 2026, http://arxiv.org/abs/2608.11110v2.
APA
Mukherjee, S., Bali, K., & Sitaram, S. (2026). Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents. arXiv. http://arxiv.org/abs/2608.11110v2
Chicago
Mukherjee, S., K. Bali, and S. Sitaram. 2026. “Actions Speak Louder Than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents”. arXiv. http://arxiv.org/abs/2608.11110v2.
Harvard
Mukherjee, S., Bali, K. and Sitaram, S. (2026) “Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2608.11110v2.
Vancouver
1. Mukherjee S, Bali K, Sitaram S (2026) Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents. arXiv

BibTeX

@article{mukherjee2026actions,
  title = {Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents},
  author = {Mukherjee, Sourabrata and Bali, Kalika and Sitaram, Sunayana},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2608.11110v2},
  eprint = {2608.11110}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/