Gemma 3 Technical Report

Gemma Team Aishwarya KamathJohan FerretShreya PathakNino VieillardRamona MerhejSarah PerrinTatiana MatejovicovaAlexandre Ram'eMorgane RivièreLouis Rouil-lard

article2025arXiv1,836 citations

Presents Gemma 3, an open family of lightweight multimodal models ranging from 1B to 27B parameters that combines vision understanding, 128K-token context processing, and memory-efficient attention to achieve performance competitive with much larger systems.

Listen

Deploying state-of-the-art artificial intelligence models typically demands massive computing resources, limiting their accessibility on consumer-grade hardware such as laptops and smartphones. At the same time, practical applications increasingly require models that can process visual inputs, handle long documents, and support multiple languages without suffering from prohibitive memory overhead. Developing open, highly capable models that operate efficiently within constrained hardware environments has therefore become a critical priority for the AI research and developer community.

The article introduces and evaluates Gemma 3, a family of lightweight, multimodal open models ranging from 1 to 27 billion parameters. It aims to demonstrate that architectural innovations and optimized training pipelines can deliver competitive visual understanding, extended context lengths, and strong reasoning capabilities while maintaining a compact memory footprint.

To achieve this, the team trained models across four sizes (1B, 4B, 12B, and 27B parameters) using pre-training datasets up to 14 trillion tokens with knowledge distillation. The architecture incorporates an interleaved local-to-global attention pattern (a five-to-one ratio with a 1,024-token sliding window) to control key-value cache memory growth and extends context windows up to 128,000 tokens. The models integrate a tailored 400-million-parameter SigLIP vision encoder using an adaptive windowing approach to handle varied image aspect ratios. Post-training involved reinforcement learning with varied reward functions covering instruction following, mathematics, and coding, as well as quantization-aware training for efficient deployment.

The evaluations yielded several significant findings. First, the 27B instruction-tuned model achieved a human-preference Elo score of 1338 on the LMSYS Chatbot Arena, ranking among the top ten models globally and outperforming substantially larger open systems such as LLaMA 3.1 405B (1269) and Qwen2.5-72B (1257). Second, the post-training recipe drove steep gains in complex reasoning: on the MATH benchmark, the 27B model reached 89.0% compared to 55.6% for Gemma 2 27B, while the 4B model (75.6%) markedly outperformed the prior 27B generation. Third, the architectural changes reduced inference key-value cache overhead from roughly 60% down to under 15% at 32,000 tokens without sacrificing model perplexity. Fourth, the models demonstrated significantly lower data memorization rates than prior iterations, with zero personal identifiable information detected in memorization audits. Finally, evaluations across chemical, biological, radiological, and nuclear domains confirmed that dangerous domain knowledge remained low.

These results indicate that architectural efficiency and refined post-training can substitute for raw parameter scale. Organizations can run frontier-class reasoning, multimodal, and multilingual capabilities directly on local or lower-cost infrastructure, substantially reducing operational serving costs and reliance on proprietary APIs. Safety evaluations also suggest that releasing these weights introduces minimal incremental risk to the broader AI landscape.

Organizations seeking to deploy cost-effective on-device or edge AI should evaluate Gemma 3 models using the provided 4-bit and 8-bit quantized checkpoints. When processing high-resolution or non-square imagery, teams should enable the adaptive windowing algorithm to maximize text recognition and detail extraction. Future work should focus on tracking long-term usage trends in production environments and refining evaluation suites to mitigate benchmark contamination risks across open-weight models.

Confidence in these findings is strong given the broad range of automated benchmarks, ablations, and blind human side-by-side evaluations. However, users should note that the smallest 1B model supports a shorter 32,000-token context window, does not include native image understanding, and displays weaker multilingual and mathematical performance than the larger variants. Additionally, performance degrades rapidly when extending context lengths beyond the supported 128,000 tokens without further adaptation.

Cover for Gemma 3 Technical Report

Abstract

We introduce Gemma 3, a multimodal addition to the Gemma family of lightweight open models, ranging in scale from 1 to 27 billion parameters. This version introduces vision understanding abilities, a wider coverage of languages and longer context - at least 128K tokens. We also change the architecture of the model to reduce the KV-cache memory that tends to explode with long context. This is achieved by increasing the ratio of local to global attention layers, and keeping the span on local attention short. The Gemma 3 models are trained with distillation and achieve superior performance to Gemma 2 for both pre-trained and instruction finetuned versions. In particular, our novel post-training recipe significantly improves the math, chat, instruction-following and multilingual abilities, making Gemma3-4B-IT competitive with Gemma2-27B-IT and Gemma3-27B-IT comparable to Gemini-1.5-Pro across benchmarks. We release all our models to the community.

Table of Contents

  • 1 Introduction
  • 2 Model Architecture
  • 2.1 Vision modality
  • 2.2 Pre-training
  • 2.3 Quantization Aware Training
  • 2.4 Compute Infrastructure
  • 3 Instruction-Tuning
  • 4 Evaluation of final models
  • 4.1 LMSYS Chatbot Arena
  • 4.2 Standard benchmarks
  • 5 Ablations
  • 5.1 Pre-training ability probing
  • 5.2 Local:Global attention layers
  • 5.3 Enabling long context
  • 5.4 Small versus large teacher
  • 5.5 Vision encoder
  • 6 Memorization and Privacy
  • 7 Responsibility, Safety, Security
  • 7.1 Governance & Assessment
  • 7.2 Safety policies and train-time mitigations
  • 7.3 Assurance Evaluations
  • 7.4 Our approach to responsible open models
  • 8 Discussion and Conclusion
  • References
  • 8.1 Performance of IT models
  • 8.2 Performance of IT models on video understanding

Knowls

  1. Knowl 1 — Gemma 3 Interleaved Local-Global Attention Architecture

    model/method

    Gemma 3 models use a decoder-only transformer architecture with Grouped-Query Attention (GQA), RMSNorm for pre-normalization and post-normalization, and QK-normalization (replacing the attention soft-capping used in Gemma 2). To mitigate KV-cache memory explosion at long context lengths, the architecture alternates between local sliding-window attention and global self-attention in a fixed 5:1 ratio:

    • Interleaving Pattern: Every block consists of 5 local sliding-window attention layers followed by 1 global attention layer, beginning with a local layer at the first layer of the network.
    • Local Window Span: Local attention layers operate on a sliding window span of W=1024W = 1024 tokens and maintain a Rotary Position Embedding (RoPE) base frequency of θlocal=10,000\theta_{\text{local}} = 10{,}000.
    • Global Attention and Positional Scaling: Global attention layers attend across the entire context window. The RoPE base frequency is increased to θglobal=1,000,000\theta_{\text{global}} = 1{,}000{,}000, and positional interpolation with a scaling factor of 8 is applied to support context lengths up to 128K tokens (32K tokens for the 1B variant).

    This 5:1 local-to-global attention structure reduces the KV-cache inference memory overhead to under 15% of total model memory at a 32K context length, compared to ~60% memory overhead in standard global-only architectures, with negligible impact on validation perplexity across tested local-to-global ratios (from 1:1 to 7:1).

  2. Knowl 2 — Gemma 3 Parameter Sizes and Vision Token Compression

    model/method

    Gemma 3 models are released in four parameter scales (1B, 4B, 12B, and 27B) sharing a 262k vocabulary size:

    Model Vision Encoder Parameters Embedding Parameters Non-Embedding Parameters
    1B 0 302M 698M
    4B 417M 675M 3,209M
    12B 417M 1,012M 10,759M
    27B 417M 1,416M 25,600M

    For multimodal variants (4B, 12B, and 27B), visual inputs are processed by a frozen 400M parameter SigLIP Vision Transformer (ViT) encoder operating on 896×896896 \times 896 pixel square inputs. The encoder output is condensed into a fixed sequence of 256 soft tokens via a 4×44 \times 4 spatial average pooling operation before being fed into the autoregressive language decoder.

  3. Knowl 3 — Inference-Time Pan and Scan Adaptive Windowing for Images

    model/method

    Because the frozen SigLIP vision encoder operates at a fixed resolution of 896×896896 \times 896 pixels, directly downsampling high-resolution or non-square images can introduce distortion, blur small text, or omit small visual details. To resolve this without retraining, Gemma 3 uses an inference-time optimization called Pan & Scan (P&S):

    1. Adaptive Segmentation: The input image is divided into a grid of non-overlapping, equally sized crops that cover the entire image while matching native aspect ratios more closely.
    2. Cropping & Resizing: Each crop is individually resized to 896×896896 \times 896 pixels and passed through the vision encoder.
    3. Token Representation: Each crop is encoded into 256 soft tokens via 4×44 \times 4 spatial pooling and presented sequentially to the language model.
    4. Inference Control: The maximum number of crops is bounded to control latency, and P&S can be toggled off when higher throughput is required.
  4. Knowl 4 — Gemma 3 Pre-Training Distillation Objective and Data Mixture

    model/method

    Gemma 3 models are pre-trained with knowledge distillation from larger teacher models across substantial token budgets:

    • Token Allocations: 14T tokens for 27B, 12T tokens for 12B, 4T tokens for 4B, and 2T tokens for 1B.
    • Data Balancing: Multilingual data includes monolingual and parallel corpora balanced via UniMax-style reweighting to ensure broad language coverage. Pre-computed vision embeddings are used for image-text pairs, incurring no vision backpropagation overhead.
    • Sampled Logit Distillation: For each token position tt, the student model samples K=256K = 256 candidate logits from the teacher distribution, weighted by teacher probabilities PteacherP_{\text{teacher}}. The teacher's target distribution P~teacher\tilde{P}_{\text{teacher}} is set to 0 for all non-sampled logits and renormalized over the KK sampled logits:

    P~teacher(w)={Pteacher(w)∑v∈SKPteacher(v),w∈SK0,w∉SK\tilde{P}_{\text{teacher}}(w) = \begin{cases} \frac{P_{\text{teacher}}(w)}{\sum_{v \in \mathcal{S}_K} P_{\text{teacher}}(v)}, & w \in \mathcal{S}_K \\ 0, & w \notin \mathcal{S}_K \end{cases}

    where SK\mathcal{S}_K denotes the set of KK sampled vocabulary items. The student minimizes cross-entropy with respect to P~teacher\tilde{P}_{\text{teacher}}. Distillation scaling experiments demonstrate that while distilling from a smaller teacher is advantageous for short training schedules, distilling from a larger teacher yields superior perplexity as total training tokens exceed 101110^{11} tokens.

  5. Knowl 5 — Gemma 3 Instruction-Tuning and Multi-Reward Alignment Recipe

    model/method

    Gemma 3 pre-trained base checkpoints are aligned into instruction-tuned (IT) models using a multi-stage post-training pipeline:

    1. Distillation from IT Teacher: Cross-entropy distillation on instruction-following datasets using a high-capacity instruction-tuned teacher.
    2. Reinforcement Learning (RL) Fine-Tuning: Policy optimization leveraging algorithms derived from BOND (Best-of-N Distillation), WARM (Weight-Averaged Reward Models), and WARP (Weight-Averaged Rewarded Policies).
    3. Multi-Reward Functions: The RL objectives incorporate:
      • Weight-averaged reward models trained on human preference data to enforce helpfulness and safety.
      • Reinforcement Learning from Execution Feedback (RLEF) using ground-truth unit test execution signals for coding tasks.
      • Ground-truth correctness rewards for mathematical problem-solving steps.
    4. Prompt and Turn Formatting: Text sequences explicitly require a [BOS] token at the start. Turns are formatted with special control tokens <start_of_turn>user, <start_of_turn>model, and <end_of_turn>:
    [BOS]<start_of_turn>user
    {user_prompt}<end_of_turn>
    <start_of_turn>model
    {model_response}<end_of_turn>
    

    While pre-trained base models generate <eos> to terminate text, instruction-tuned models terminate generation with <end_of_turn>.

  6. Knowl 6 — Zero-Shot Benchmark Performance of Gemma 3 Instruction-Tuned Models

    data/table

    Gemma 3 instruction-tuned (IT) models exhibit substantial performance improvements over Gemma 2, with the 4B model achieving performance comparable to the 27B model of the previous generation:

    Benchmark Gemini 1.5 Flash Gemini 1.5 Pro Gemma 2 2B Gemma 2 9B Gemma 2 27B Gemma 3 1B Gemma 3 4B Gemma 3 12B Gemma 3 27B
    MMLU-Pro 67.3 75.8 15.6 46.8 56.9 14.7 43.6 60.6 67.5
    LiveCodeBench 30.7 34.2 1.2 10.8 20.4 1.9 12.6 24.6 29.7
    Bird-SQL (dev) 45.6 54.4 12.2 33.8 46.7 6.4 36.3 47.9 54.4
    GPQA Diamond 51.0 59.1 24.7 28.8 34.3 19.2 30.8 40.9 42.4
    SimpleQA 8.6 24.9 2.8 5.3 9.2 2.2 4.0 6.3 10.0
    FACTS Grounding 82.9 80.0 43.8 62.0 62.4 36.4 70.1 75.8 74.9
    Global MMLU-Lite 73.7 80.8 41.9 64.8 68.6 34.2 54.5 69.5 75.1
    MATH 77.9 86.5 27.2 49.4 55.6 48.0 75.6 83.8 89.0
    HiddenMath 47.2 52.0 1.8 10.4 14.8 15.8 43.0 54.5 60.3
    MMMU (val) 62.3 65.9 - - - - 48.8 59.6 64.9

    In LMSYS Chatbot Arena human preference evaluations, Gemma-3-27B-IT achieved an Elo score of 1338 (95% CI +8/-9, rank 9), outperforming Gemma-2-27B-IT (Elo 1220), Llama-3.3-70B-Instruct (Elo 1257), Qwen2.5-72B-Instruct (Elo 1257), and DeepSeek-V3 (Elo 1318).

  7. Knowl 7 — Quantization-Aware Training and Inference Memory Footprints

    data/table

    Gemma 3 checkpoints were adapted to standard low-bit inference representations by applying Quantization Aware Training (QAT) for approximately 5,000 steps using non-quantized model probabilities as targets. Weight representations include per-channel int4, per-block int4 (block size 32), and switched fp8 (SFP8). Memory footprint comparisons for model weights alone and combined with a 32,768-token 8-bit quantized KV cache (+KV) are reported in Gigabytes (GB):

    Model Raw (bfloat16) Int4 Int4 (blocks=32) SFP8
    1B 2.0 0.5 0.7 1.0
    1B (+KV) 2.9 1.4 1.6 1.9
    4B 8.0 2.6 2.9 4.4
    4B (+KV) 12.7 7.3 7.6 9.1
    12B 24.0 6.6 7.1 12.4
    12B (+KV) 38.9 21.5 22.0 27.3
    27B 54.0 14.1 15.3 27.4
    27B (+KV) 72.7 32.8 34.0 46.1
  8. Knowl 8 — Impact of Pan and Scan and Encoder Resolution on Multimodal Benchmarks

    empirical result

    Ablation evaluations demonstrate that both the native image encoder resolution and the inference-time Pan and Scan (P&S) adaptive windowing algorithm are critical for high-resolution and text-centric visual tasks:

    1. Encoder Input Resolution: When evaluating a 2B baseline schedule on visual QA benchmarks with encoder outputs pooled to 256 tokens, increasing input resolution from 256×256256 \times 256 to 896×896896 \times 896 increased DocVQA performance from 31.9 to 59.8, InfoVQA from 23.1 to 33.7, and TextVQA from 44.1 to 58.0.
    2. Pan & Scan Performance: Applying inference-time P&S cropping to pre-trained base checkpoints yielded substantial accuracy gains on 4-shot evaluation sets containing diverse aspect ratios and embedded text:
    Configuration DocVQA InfoVQA TextVQA
    4B Base 72.8 44.1 58.9
    4B w/ PS 81.0 57.0 60.8
    Δ\Delta +8.2 +12.9 +1.9
    27B Base 85.6 59.4 68.6
    27B w/ PS 90.4 76.4 70.2
    Δ\Delta +4.8 +17.0 +1.6
  9. Knowl 9 — Gemma 3 Multimodal Fine-Tuning Transfer vs PaliGemma 2

    data/table

    When fine-tuned on multimodal benchmark datasets following the PaliGemma 2 transfer protocol, Gemma 3 pre-trained checkpoints surpass PaliGemma 2 on document understanding tasks while requiring significantly lower compute due to token compression (256 tokens per image crop):

    Benchmark PaliGemma 2 2B PaliGemma 2 9B PaliGemma 2 27B Gemma 3 4B Gemma 3 12B Gemma 3 27B
    DocVQA 81.6 86.3 85.1 86.1 89.0 89.5
    InfoVQA 41.4 53.1 50.2 55.6 61.6 64.6
    TextVQA 76.3 76.3 75.1 79.1 81.6 83.2
    ChartQA 70.7 79.1 71.3 79.8 83.5 83.4
    AI2D 76.0 84.4 84.6 80.9 85.6 86.5
    OKVQA 64.1 68.6 70.6 65.2 69.3 71.1
    CountBenchQA 82.0 85.3 87.4 79.4 83.5 87.8
    COCO caption 143.0 145.0 145.0 143.0 143.0 144.0
    VQAv2 84.8 85.8 85.8 84.1 84.9 85.1
    Tally QA 80.6 82.4 82.1 79.0 81.3 81.7

    Due to spatial average pooling of vision features to 256 tokens, fine-tuning transfer for Gemma 3 4B and 12B models is approximately 10×10\times cheaper computationally compared to PaliGemma 2 9B and 27B at 896×896896 \times 896 resolution.

  10. Knowl 10 — Training Data Memorization and Privacy Audit Results

    empirical result

    Training data memorization was audited using discoverable extraction tests with 50-token prefixes and 50-token suffixes sampled uniformly across the pre-training corpus. Generations were classified as exactly memorized (100% token continuation match) or approximately memorized (match within an edit distance of 10%):

    • Memorization Reduction: Gemma 3 models memorize long-form training sequences at substantially lower rates than Gemma 1, Gemma 2, and PaLM models (e.g., total memorization rate orders of magnitude below older baselines on logarithmic scale).
    • Exact vs. Approximate: Across Gemma 3 sizes, approximate memorization occurs roughly 24×24\times more frequently than exact memorization.
    • Scale Invariance: The 4B, 12B, and 27B models display only marginal differences in memorization rate, while the 1B model memorizes less.
    • PII Leakage Assessment: Auditing memorized outputs with the Google Cloud Sensitive Data Protection (SDP) detection service found zero instances of personal identifiable information at any severity level across all Gemma 3 model generations.

Coverage note — Omitted fine-grained benchmark score breakdowns for individual IndicGenBench sub-tasks (Table 14) and individual video/reasoning benchmarks (Tables 17, 18, and 20) as they are fully summarized by the high-level capability comparisons and benchmark tables.

References

  1. 1.Realworldqa. https://x.ai/news/grok-1.5v.
  2. 2.M. Acharya, K. Kafle, and C. Kanan. Tallyqa: Answering complex counting questions. In AAAI, 2018.
  3. 3.R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. R. Garea, M. Geist, and O. Bachem. On-policy distillation of language models: Learning from self-generated mistakes. In ICLR, 2024.
  4. 4.J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. arXiv preprint arXiv:2305.13245, 2023.
  5. 5.R. Anil, G. Pereyra, A. Passos, R. Ormandi, G. E. Dahl, and G. E. Hinton. Large scale distributed neural network training through online distillation. arXiv preprint arXiv:1804.03235, 2018.
  6. 6.R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  7. 7.M. Artetxe, S. Ruder, and D. Yogatama. On the cross-lingual transferability of monolingual representations. In ACL, 2020.
  8. 8.A. Asai, J. Kasai, J. H. Clark, K. Lee, E. Choi, and H. Hajishirzi. Xor qa: Cross-lingual open-retrieval question answering. arXiv preprint arXiv:2010.11856, 2020.
  9. 9.J. Austin, A. Odena, M. I. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. J. Cai, M. Terry, Q. V. Le, and C. Sutton. Program synthesis with large language models. CoRR, abs/2108.07732, 2021.
  10. 10.P. Barham, A. Chowdhery, J. Dean, S. Ghemawat, S. Hand, D. Hurt, M. Isard, H. Lim, R. Pang, S. Roy, B. Saeta, P. Schuh, R. Sepassi, L. E. Shafey, C. A. Thekkath, and Y. Wu. Pathways: Asynchronous distributed dataflow for ml, 2022.
  11. 11.I. Beltagy, M. E. Peters, and A. Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  12. 12.S. Biderman, U. Prashanth, L. Sutawika, H. Schoelkopf, Q. Anthony, S. Purohit, and E. Raff. Emergent and predictable memorization in large language models. NeurIPS, 36: 28072–28090, 2023.
  13. 13.Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi. PIQA: reasoning about physical commonsense in natural language. CoRR, abs/1911.11641, 2019.
  14. 14.N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al. Extracting training data from large language models. In USENIX, 2021.
  15. 15.N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646, 2022.
  16. 16.Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024.
  17. 17.M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba. Evaluating large language models trained on code. CoRR, abs/2107.03374, 2021.
  18. 18.S. Chen, S. Wong, L. Chen, and Y. Tian. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595, 2023.
  19. 19.X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick. Microsoft coco captions: Data collection and evaluation server. ArXiv, abs/1504.00325, 2015.
  20. 20.W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, and I. Stoica. Chatbot arena: An open platform for evaluating llms by human preference, 2024.
  21. 21.F. Chollet. On the measure of intelligence. arXiv preprint arXiv:1911.01547, 2019.
  22. 22.A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel. Palm: Scaling language modeling with pathways, 2022.
  23. 23.H. W. Chung, N. Constant, X. Garcia, A. Roberts, Y. Tay, S. Narang, and O. Firat. Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining, 2023.
  24. 24.C. Clark, K. Lee, M. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. CoRR, abs/1905.10044, 2019.
  25. 25.K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021.
  26. 26.DeepSeek-AI. Deepseek-r1: Incentivizing reasoningt learning, 2025.
  27. 27.M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al. Scaling vision transformers to 22 billion parameters. In ICML, 2023.
  28. 28.D. Deutsch, E. Briakou, I. Caswell, M. Finkelstein, R. Galor, J. Juraska, G. Kovacs, A. Lui, R. Rei, J. Riesa, S. Rijhwani, P. Riley, E. Salesky, F. Trabelsi, S. Winkler, B. Zhang, and M. Freitag. Wmt24++: Expanding the language coverage of wmt24 to 55 languages & dialects, 2025.
  29. 29.A. Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  30. 30.D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In ACL, 2019.
  31. 31.B. Fatemi, M. Kazemi, A. Tsitsulin, K. Malkan, J. Yim, J. Palowitch, S. Seo, J. Halcrow, and B. Perozzi. Test of time: A benchmark for evaluating llms on temporal reasoning. arXiv preprint arXiv:2406.09170, 2024.
  32. 32.X. Fu, Y. Hu, B. Li, Y. Feng, H. Wang, X. Lin, D. Roth, N. A. Smith, W.-C. Ma, and R. Krishna. Blink: Multimodal large language models can see but not perceive. ArXiv, abs/2404.12390, 2024.
  33. 33.J. Gehring, K. Zheng, J. Copet, V. Mella, T. Cohen, and G. Synnaeve. Rlef: Grounding code llms in execution feedback with reinforcement learning. arXiv preprint arXiv:2410.02089, 2024.
  34. 34.Gemini Team. Gemini: A family of highly capable multimodal models, 2023.
  35. 35.Gemini Team. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.
  36. 36.Gemma Team. Gemma: Open models based on gemini research and technology, 2024a.
  37. 37.Gemma Team. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024b.
  38. 38.O. Goldman, U. Shaham, D. Malkin, S. Eiger, A. Hassidim, Y. Matias, J. Maynez, A. M. Gilady, J. Riesa, S. Rijhwani, L. Rimell, I. Szpektor, R. Tsarfaty, and M. Eyal. Eclektic: a novel challenge set for evaluation of cross-lingual knowledge transfer, 2025.
  39. 39.N. Goyal, C. Gao, V. Chaudhary, P.-J. Chen, G. Wenzek, D. Ju, S. Krishnan, M. Ranzato, F. Guzmán, and A. Fan. The flores-101 evaluation benchmark for low-resource and multilingual machine translation. ACL, 2022.
  40. 40.Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering. In CVPR, 2017.
  41. 41.D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. CoRR, abs/2009.03300, 2020.
  42. 42.D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset. NeurIPS, 2021.
  43. 43.J. Hessel, A. Marasović, J. D. Hwang, L. Lee, J. Da, R. Zellers, R. Mankoff, and Y. Choi. Do androids laugh at electric sheep? humor" understanding" benchmarks from the new yorker caption contest. arXiv preprint arXiv:2209.06293, 2022.
  44. 44.G. Hinton, O. Vinyals, and J. Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
  45. 45.C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg. Ruler: What’s the real context size of your long-context language models? arXiv preprint arXiv:2404.06654, 2024.
  46. 46.D. Ippolito, F. Tramèr, M. Nasr, C. Zhang, M. Jagielski, K. Lee, C. A. Choquette-Choo, and N. Carlini. Preventing verbatim memorization in language models gives a false sense of privacy. arXiv preprint arXiv:2210.17546, 2022.
  47. 47.B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In CVPR, 2018.
  48. 48.M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. CoRR, abs/1705.03551, 2017.
  49. 49.M. Kazemi, H. Alvari, A. Anand, J. Wu, X. Chen, and R. Soricut. Geomverse: A systematic evaluation of large models for geometric reasoning. arXiv preprint arXiv:2312.12241, 2023.
  50. 50.M. Kazemi, N. Dikkala, A. Anand, P. Dević, I. Dasgupta, F. Liu, B. Fatemi, P. Awasthi, D. Guo, S. Gollapudi, and A. Qureshi. Remi: A dataset for reasoning with multiple images. ArXiv, abs/2406.09175, 2024a.
  51. 51.M. Kazemi, Q. Yuan, D. Bhatia, N. Kim, X. Xu, V. Imbrasaite, and D. Ramachandran. Boardgameqa: A dataset for natural language reasoning with contradictory information. NeurIPS, 36, 2024b.
  52. 52.M. Kazemi, B. Fatemi, H. Bansal, J. Palowitch, C. Anastasiou, S. V. Mehta, L. K. Jain, V. Aglietti, D. Jindal, P. Chen, et al. Big-bench extra hard. arXiv preprint arXiv:2502.19187, 2025.
  53. 53.A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi. A diagram is worth a dozen images. ArXiv, abs/1603.07396, 2016.
  54. 54.E. Kıcıman, R. Ness, A. Sharma, and C. Tan. Causal reasoning and large language models: Opening a new frontier for causality. arXiv preprint arXiv:2305.00050, 2023.
  55. 55.T. Kudo and J. Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. 2018.
  56. 56.T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov. Natural questions: A benchmark for question answering research. ACL, 2019.
  57. 57.N. Lambert, J. Morrison, V. Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V. Miranda, A. Liu, N. Dziri, S. Lyu, et al. T" ulu 3: Pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124, 2024.
  58. 58.Z. Lin, J. Cui, X. Liao, and X. Wang. Malla: Demystifying real-world large language model integrated malicious services, 2024.
  59. 59.H. Liu, C. Li, Q. Wu, and Y. J. Lee. Visual instruction tuning. NeurIPS, 36, 2024.
  60. 60.LLaMa Team. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  61. 61.M. Luong, H. Pham, and C. D. Manning. Effective approaches to attention-based neural machine translation. 2015.
  62. 62.Macknight, Aung, and Gomes. Personal Communication.
  63. 63.K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In CVPR, 2019.
  64. 64.A. Masry, X. L. Do, J. Q. Tan, S. Joty, and E. Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. ACL, 2022.
  65. 65.M. Mathew, D. Karatzas, R. Manmatha, and C. V. Jawahar. Docvqa: A dataset for vqa on document images. WACV, 2020.
  66. 66.M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar. Infographicvqa. In WACV, 2022.
  67. 67.I. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M. Farajtabar. Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models. arXiv preprint arXiv:2410.05229, 2024.
  68. 68.M. Nasr, N. Carlini, J. Hayase, M. Jagielski, A. F. Cooper, D. Ippolito, C. A. Choquette-Choo, E. Wallace, F. Tramèr, and K. Lee. Scalable extraction of training data from (production) language models. arXiv preprint arXiv:2311.17035, 2023.
  69. 69.A. Nie, Y. Zhang, A. S. Amdekar, C. Piech, T. B. Hashimoto, and T. Gerstenberg. Moca: Measuring human-language model alignment on causal and moral judgment tasks. NeurIPS, 36, 2024.
  70. 70.R. Paiss, A. Ephrat, O. Tov, S. Zada, I. Mosseri, M. Irani, and T. Dekel. Teaching clip to count to ten. ICCV, 2023.
  71. 71.M. Phuong, M. Aitchison, E. Catt, S. Cogan, A. Kaskasoli, V. Krakovna, D. Lindner, M. Rahtz, Y. Assael, S. Hodkinson, H. Howard, T. Lieberum, R. Kumar, M. A. Raad, A. Webson, L. Ho, S. Lin, S. Farquhar, M. Hutter, G. Deletang, A. Ruoss, S. El-Sayed, S. Brown, A. Dragan, R. Shah, A. Dafoe, and T. Shevlane. Evaluating frontier models for dangerous capabilities, 2024.
  72. 72.A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PMLR, 2021.
  73. 73.A. Ramé, J. Ferret, N. Vieillard, R. Dadashi, L. Hussenot, P.-L. Cedoz, P. G. Sessa, S. Girgin, A. Douillard, and O. Bachem. WARP: On the benefits of weight averaged rewarded policies, 2024a.
  74. 74.A. Ramé, N. Vieillard, L. Hussenot, R. Dadashi, G. Cideron, O. Bachem, and J. Ferret. WARM: On the benefits of weight averaged reward models. In ICML, 2024b.
  75. 75.D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark. ArXiv, abs/2311.12022, 2023.
  76. 76.J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He. Zero-offload: Democratizing billion-scale model training. In USENIX, 2021.
  77. 77.A. Roberts, H. W. Chung, G. Mishra, A. Levskaya, J. Bradbury, D. Andor, S. Narang, B. Lester, C. Gaffney, A. Mohiuddin, et al. Scaling up models and data with t5x and seqio. JMLR, 2023.
  78. 78.N. Sachdeva, B. Coleman, W.-C. Kang, J. Ni, L. Hong, E. H. Chi, J. Caverlee, J. McAuley, and D. Z. Cheng. How to train data-efficient llms. arXiv preprint arXiv:2402.09668, 2024.
  79. 79.K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi. WINOGRANDE: an adversarial winograd schema challenge at scale. CoRR, abs/1907.10641, 2019.
  80. 80.E. Sánchez, B. Alastruey, C. Ropers, P. Stenetorp, M. Artetxe, and M. R. Costa-jussà. Linguini: A benchmark for language-agnostic linguistic reasoning. arXiv preprint arXiv:2409.12126, 2024.
  81. 81.M. Sap, H. Rashkin, D. Chen, R. L. Bras, and Y. Choi. Socialiqa: Commonsense reasoning about social interactions. CoRR, abs/1904.09728, 2019.
  82. 82.P. G. Sessa, R. Dadashi, L. Hussenot, J. Ferret, N. Vieillard, A. Ramé, B. Shariari, S. Perrin, A. Friesen, G. Cideron, S. Girgin, P. Stanczyk, A. Michi, D. Sinopalnikov, S. Ramos, A. Héliou, A. Severyn, M. Hoffman, N. Momchev, and O. Bachem. Bond: Aligning llms with best-of-n distillation, 2024.
  83. 83.K. Shah, N. Dikkala, X. Wang, and R. Panigrahy. Causal language modeling can elicit search and reasoning capabilities on logic puzzles. arXiv preprint arXiv:2409.10502, 2024.
  84. 84.T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal, M. Anderljung, N. Kolt, L. Ho, D. Siddarth, S. Avin, W. Hawkins, B. Kim, I. Gabriel, V. Bolina, J. Clark, Y. Bengio, P. Christiano, and A. Dafoe. Model evaluation for extreme risks, 2023.
  85. 85.F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, D. Das, and J. Wei. Language models are multilingual chain-of-thought reasoners. In ICLR, 2023.
  86. 86.A. Singh, V. Natarjan, M. Shah, Y. Jiang, X. Chen, D. Parikh, and M. Rohrbach. Towards vqa models that can read. In CVPR, 2019.
  87. 87.H. Singh, N. Gupta, S. Bharadwaj, D. Tewari, and P. Talukdar. Indicgenbench: a multilingual benchmark to evaluate generation capabilities of llms on indic languages. arXiv preprint arXiv:2404.16816, 2024a.
  88. 88.S. Singh, A. Romanou, C. Fourrier, D. I. Adelani, J. G. Ngui, D. Vila-Suero, P. Limkonchotiwat, K. Marchisio, W. Q. Leong, Y. Susanto, R. Ng, S. Longpre, W.-Y. Ko, M. Smith, A. Bosselut, A. Oh, A. F. T. Martins, L. Choshen, D. Ippolito, E. Ferrante, M. Fadaee, B. Ermis, and S. Hooker. Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation, 2024b.
  89. 89.A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y. Bitton, A. Gritsenko, M. Minderer, A. Sherbondy, S. Long, S. Qin, R. Ingle, E. Bugliarello, S. Kazemzadeh, T. Mesnard, I. Alabdulmohsin, L. Beyer, and X. Zhai. PaliGemma 2: A Family of Versatile VLMs for Transfer. arXiv preprint arXiv:2412.03555, 2024.
  90. 90.M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, and J. Wei. Challenging big-bench tasks and whether chain-of-thought can solve them, 2022.
  91. 91.G. Tyen, H. Mansoor, P. Chen, T. Mak, and V. Cărbune. Llms cannot find reasoning errors, but can correct them! arXiv preprint arXiv:2311.08516, 2023.
  92. 92.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. 2017.
  93. 93.K. Vodrahalli, S. Ontanon, N. Tripuraneni, K. Xu, S. Jain, R. Shivanna, J. Hui, N. Dikkala, M. Kazemi, B. Fatemi, et al. Michelangelo: Long context evaluations beyond haystacks via latent structure queries. arXiv preprint arXiv:2409.12640, 2024.
  94. 94.Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In NeurIPS, 2024.
  95. 95.L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Rimell, L. A. Hendricks, W. Isaac, S. Legassick, G. Irving, and I. Gabriel. Ethical and social risks of harm from language models, 2021.
  96. 96.C. White, S. Dooley, M. Roberts, A. Pal, B. Feuer, S. Jain, R. Shwartz-Ziv, N. Jain, K. Saifullah, S. Naidu, et al. Livebench: A challenging, contamination-free llm benchmark. arXiv preprint arXiv:2406.19314, 2024.
  97. 97.M. Wortsman, P. J. Liu, L. Xiao, K. Everett, A. Alemi, B. Adlam, J. D. Co-Reyes, I. Gur, A. Kumar, R. Novak, et al. Small-scale proxies for large-scale transformer training instabilities. arXiv preprint arXiv:2309.14322, 2023.
  98. 98.XLA. Xla: Optimizing compiler for tensorflow, 2019. URL https://www.tensorflow.org/xla.
  99. 99.Y. Xu, H. Lee, D. Chen, B. A. Hechtman, Y. Huang, R. Joshi, M. Krikun, D. Lepikhin, A. Ly, M. Maggioni, R. Pang, N. Shazeer, S. Wang, T. Wang, Y. Wu, and Z. Chen. GSPMD: general and scalable parallelization for ML computation graphs. 2021.
  100. 100.Y. Yamada, Y. Bao, A. K. Lampinen, J. Kasai, and I. Yildirim. Evaluating spatial understanding of large language models. arXiv preprint arXiv:2310.14540, 2023.
  101. 101.K. Yang, O. Russakovsky, and J. Deng. Spatialsense: An adversarially crowdsourced benchmark for spatial relation recognition. ICCV, 2019.
  102. 102.X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. CVPR, 2023.
  103. 103.R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. HellaSwag: Can a machine really finish your sentence? In ACL, 2019.
  104. 104.X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid loss for language image pre-training. In CVPR, 2023.
  105. 105.B. Zhang and R. Sennrich. Root mean square layer normalization. 2019.
  106. 106.J. Zhang, L. Jain, Y. Guo, J. Chen, K. L. Zhou, S. Suresh, A. Wagenmaker, S. Sievert, T. Rogers, K. Jamieson, et al. Humor in ai: Massive scale crowd-sourced preferences and benchmarks for cartoon captioning. arXiv preprint arXiv:2406.10522, 2024.
  107. 107.W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan. Agieval: A human-centric benchmark for evaluating foundation models, 2023.

Citation

MLA
Team, G., et al. “Gemma 3 Technical Report”. arXiv, 2025, http://arxiv.org/abs/2503.19786v1.
APA
Team, G., Kamath, A., Ferret, J., Pathak, S., Vieillard, N., Merhej, R., Perrin, S., Matejovicova, T., Ramé, A., Rivière, M., Rouillard, L., Mesnard, T., Cideron, G., Grill, J.-. bastien ., Ramos, S., Yvinec, E., Casbon, M., Pot, E., Penchev, I., … Hussenot, L. (2025). Gemma 3 Technical Report. arXiv. http://arxiv.org/abs/2503.19786v1
Chicago
Team, G., A. Kamath, J. Ferret, et al. 2025. “Gemma 3 Technical Report”. arXiv. http://arxiv.org/abs/2503.19786v1.
Harvard
Team, G. et al. (2025) “Gemma 3 Technical Report”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.19786v1.
Vancouver
1. Team G, Kamath A, Ferret J, et al (2025) Gemma 3 Technical Report. arXiv

BibTeX

@article{team2025gemma,
  title = {Gemma 3 Technical Report},
  author = {Team, Gemma and Kamath, Aishwarya and Ferret, Johan and Pathak, Shreya and Vieillard, Nino and Merhej, Ramona and Perrin, Sarah and Matejovicova, Tatiana and Ramé, Alexandre and Rivière, Morgane and Rouillard, Louis and Mesnard, Thomas and Cideron, Geoffrey and Grill, Jean-bastien and Ramos, Sabela and Yvinec, Edouard and Casbon, Michelle and Pot, Etienne and Penchev, Ivo and Liu, Gaël and Visin, Francesco and Kenealy, Kathleen and Beyer, Lucas and Zhai, Xiaohai and Tsitsulin, Anton and Busa-Fekete, Robert and Feng, Alex and Sachdeva, Noveen and Coleman, Benjamin and Gao, Yi and Mustafa, Basil and Barr, Iain and Parisotto, Emilio and Tian, David and Eyal, Matan and Cherry, Colin and Peter, Jan-Thorsten and Sinopalnikov, Danila and Bhupatiraju, Surya and Agarwal, Rishabh and Kazemi, Mehran and Malkin, Dan and Kumar, Ravin and Vilar, David and Brusilovsky, Idan and Luo, Jiaming and Steiner, Andreas and Friesen, Abe and Sharma, Abhanshu and Sharma, Abheesht and Gilady, Adi Mayrav and Goedeckemeyer, Adrian and Saade, Alaa and Feng, Alex and Kolesnikov, Alexander and Bendebury, Alexei and Abdagic, Alvin and Vadi, Amit and György, András and Pinto, André Susano and Das, Anil and Bapna, Ankur and Miech, Antoine and Yang, Antoine and Paterson, Antonia and Shenoy, Ashish and Chakrabarti, Ayan and Piot, Bilal and Wu, Bo and Shahriari, Bobak and Petrini, Bryce and Chen, Charlie and Lan, Charline Le and Choquette-Choo, Christopher A. and Carey, CJ and Brick, Cormac and Deutsch, Daniel and Eisenbud, Danielle and Cattle, Dee and Cheng, Derek and Paparas, Dimitris and Sreepathihalli, Divyashree Shivakumar and Reid, Doug and Tran, Dustin and Zelle, Dustin and Noland, Eric and Huizenga, Erwin and Kharitonov, Eugene and Liu, Frederick and Amirkhanyan, Gagik and Cameron, Glenn and Hashemi, Hadi and Klimczak-Plucińska, Hanna and Singh, Harman and Mehta, Harsh and Lehri, Harshal Tushar and Hazimeh, Hussein and Ballantyne, Ian and Szpektor, Idan and Nardini, Ivan and Pouget-Abadie, Jean and Chan, Jetha and Stanton, Joe and Wieting, John and Lai, Jonathan and Orbay, Jordi and Fernandez, Joseph and Newlan, Josh and Ji, Ju-yeong and Singh, Jyotinder and Black, Kat and Yu, Kathy and Hui, Kevin and Vodrahalli, Kiran and Greff, Klaus and Qiu, Linhai and Valentine, Marcella and Coelho, Marina and Ritter, Marvin and Hoffman, Matt and Watson, Matthew and Chaturvedi, Mayank and Moynihan, Michael and Ma, Min and Babar, Nabila and Noy, Natasha and Byrd, Nathan and Roy, Nick and Momchev, Nikola and Chauhan, Nilay and Sachdeva, Noveen and Bunyan, Oskar and Botarda, Pankil and Caron, Paul and Rubenstein, Paul Kishan and Culliton, Phil and Schmid, Philipp and Sessa, Pier Giuseppe and Xu, Pingmei and Stanczyk, Piotr and Tafti, Pouya and Shivanna, Rakesh and Wu, Renjie and Pan, Renke and Rokni, Reza and Willoughby, Rob and Vallu, Rohith and Mullins, Ryan and Jerome, Sammy and Smoot, Sara and Girgin, Sertan and Iqbal, Shariq and Reddy, Shashir and Sheth, Shruti and Põder, Siim and Bhatnagar, Sijal and Panyam, Sindhu Raghuram and Eiger, Sivan and Zhang, Susan and Liu, Tianqi and Yacovone, Trevor and Liechty, Tyler and Kalra, Uday and Evci, Utku and Misra, Vedant and Roseberry, Vincent and Feinberg, Vlad and Kolesnikov, Vlad and Han, Woohyun and Kwon, Woosuk and Chen, Xi and Chow, Yinlam and Zhu, Yuvein and Wei, Zichuan and Egyed, Zoltan and Cotruta, Victor and Giang, Minh and Kirk, Phoebe and Rao, Anand and Black, Kat and Babar, Nabila and Lo, Jessica and Moreira, Erica and Martins, Luiz Gustavo and Sanseviero, Omar and Gonzalez, Lucas and Gleicher, Zach and Warkentin, Tris and Mirrokni, Vahab and Senter, Evan and Collins, Eli and Barral, Joelle and Ghahramani, Zoubin and Hadsell, Raia and Matias, Yossi and Sculley, D. and Petrov, Slav and Fiedel, Noah and Shazeer, Noam and Vinyals, Oriol and Dean, Jeff and Hassabis, Demis and Kavukcuoglu, Koray and Farabet, Clement and Buchatskaya, Elena and Alayrac, Jean-Baptiste and Anil, Rohan and \{Dmitry\} and \{Lepikhin\} and Borgeaud, Sebastian and Bachem, Olivier and Joulin, Armand and Andreev, Alek and Hardin, Cassidy and Dadashi, Robert and Hussenot, Léonard},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.19786v1},
  eprint = {2503.19786}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/