Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models

Amartya RoyIit SireDelhi RobertIndia Bosch GmbHKripabandhu GhoshIiser KolkataP. KumaraguruAdrian de Wynter

article2025arXiv0 citations

Demonstrates that smaller fine-tuned encoder-based architectures handle complex causal reasoning and distribution shifts far better than in-context learning with decoder-only models, challenging the prevailing reliance on massive autoregressive language models.

Listen

Modern artificial intelligence systems increasingly rely on in-context learning within large language models to perform complex tasks. However, safety-critical decision-making and scientific discovery demand reliable causal reasoning. Causal reasoning requires strict rule-based deduction: chaining multiple logical steps (multi-hop composition) and ensuring all conditions are met before drawing a conclusion (strict conjunctive control). Current systems often rely on superficial word associations rather than true underlying logical structures, raising fundamental questions about whether prevailing neural network designs restrict reasoning capabilities.

This article evaluates how different neural network architectures perform deterministic logical deduction under distribution shifts. Specifically, it tests whether encoder-based models—which process whole inputs simultaneously into a global representation—exhibit greater reasoning stability and generalizability compared to autoregressive, decoder-only models that generate answers sequentially step-by-step.

The authors constructed a controlled synthetic framework based on first-order logic rules, establishing a training dataset of 40,000 samples (reasoning depths 0 to 7) and two distinct out-of-distribution test sets of 3,600 samples each (reasoning depths 0 to 11). The first test set used natural language expressions with deeper reasoning chains, while the second replaced natural words with randomized character strings to eliminate lexical shortcuts. The evaluation benchmarked both zero- and few-shot in-context learning across proprietary and open decoder models (including GPT-4.1, GPT-5, Claude Opus 4.1, and Qwen variants) as well as fine-tuned encoder (BERT), encoder–decoder (BART, Flan-T5), and decoder-only architectures.

The investigation produced four central findings. First, standard in-context prompting proved insufficient for dependable deduction; without task-specific fine-tuning, smaller models exhibited near-random discrimination, and performance barely changed between zero-shot and five-shot prompts. Second, fine-tuned encoder and encoder–decoder architectures generalized more effectively to deeper reasoning paths and abstract symbols than comparable fine-tuned decoders. In the non-natural language test, fine-tuned BERT achieved 61% accuracy, whereas the fine-tuned decoder Qwen3-1.7B dropped to near-random performance (53% accuracy) by depth 3. Third, fine-tuned encoder-based models were exceptionally cost- and compute-efficient; BART-Base achieved its accuracy with an inference efficiency of 640 accuracy points per hour on standard hardware. Fourth, while massive frontier models such as GPT-5 achieved near-perfect accuracy across all depths, they exhibited the lowest operational efficiency (1.1 points per hour), requiring over an hour of API inference per accuracy point.

These findings demonstrate that deploying off-the-shelf decoder models via prompting introduces operational risks and potential reasoning failures when rules extend beyond familiar patterns. Mechanistic analysis indicates that bidirectional encoders maintain geometric consistency across multi-step inferences, whereas sequential decoders accumulate drift unless scaled to massive compute budgets. For organizations implementing rule-based and causal workflows, reliance on large decoder-only models via prompting carries severe cost and latency penalties that may not be justified.

Organizations should adopt targeted, fine-tuned encoder or encoder–decoder models (such as BERT or BART variants) for structured, short-horizon causal deduction to optimize cost, latency, and reliability. Highly expensive frontier reasoning models like GPT-5 should be reserved exclusively for scenarios requiring open-ended generation or deep chains that smaller systems cannot resolve. Where hybrid needs exist, decision-makers should pilot architectures that pair the structured representation capabilities of encoders with the generative interface of decoders.

These conclusions are bounded by the study's reliance on synthetic Horn-clause logic without disjunctions, focusing on formal deductive steps rather than real-world unstructured text. In addition, because proprietary frontier models were accessed via remote application interfaces on undisclosed hardware, precise runtime comparisons represent operational estimates. Decision-makers can place high confidence in the architectural trade-offs demonstrated, while exercising caution when translating synthetic deduction benchmarks directly to noisy, real-world domain data.

arXiv: 2512.10561
Cover for Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models

Abstract

In context learning (ICL) underpins recent advances in large language models (LLMs), although its role and performance in causal reasoning remains unclear. Causal reasoning demands multihop composition and strict conjunctive control, and reliance on spurious lexical relations of the input could provide misleading results. We hypothesize that, due to their ability to project the input into a latent space, encoder and encoder decoder architectures are better suited for said multihop conjunctive reasoning versus decoder only models. To do this, we compare fine-tuned versions of all the aforementioned architectures with zero and few shot ICL in both natural language and non natural language scenarios. We find that ICL alone is insufficient for reliable causal reasoning, often overfocusing on irrelevant input features. In particular, decoder only models are noticeably brittle to distributional shifts, while finetuned encoder and encoder decoder models can generalize more robustly across our tests, including the non natural language split. Both architectures are only matched or surpassed by decoder only architectures at large scales. We conclude by noting that for cost effective, short horizon robust causal reasoning, encoder or encoder decoder architectures with targeted finetuning are preferable.

Citation

MLA
Roy, A., et al. “Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models”. arXiv, 2025, http://arxiv.org/abs/2512.10561v1.
APA
Roy, A., M, E., Ghosh, K., Kumaraguru, P., & Wynter, A. de . (2025). Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models. arXiv. http://arxiv.org/abs/2512.10561v1
Chicago
Roy, A., E. M, K. Ghosh, P. Kumaraguru, and A. de . Wynter. 2025. “Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models”. arXiv. http://arxiv.org/abs/2512.10561v1.
Harvard
Roy, A. et al. (2025) “Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2512.10561v1.
Vancouver
1. Roy A, M E, Ghosh K, Kumaraguru P, Wynter A de (2025) Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models. arXiv

BibTeX

@article{roy2025causal,
  title = {Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models},
  author = {Roy, Amartya and M, Elamparithy and Ghosh, Kripabandhu and Kumaraguru, Ponnurangam and Wynter, Adrian de},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2512.10561v1},
  eprint = {2512.10561}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/