Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
Amartya RoyIit SireDelhi RobertIndia Bosch GmbHKripabandhu GhoshIiser KolkataP. KumaraguruAdrian de Wynter
Demonstrates that smaller fine-tuned encoder-based architectures handle complex causal reasoning and distribution shifts far better than in-context learning with decoder-only models, challenging the prevailing reliance on massive autoregressive language models.
Modern artificial intelligence systems increasingly rely on in-context learning within large language models to perform complex tasks. However, safety-critical decision-making and scientific discovery demand reliable causal reasoning. Causal reasoning requires strict rule-based deduction: chaining multiple logical steps (multi-hop composition) and ensuring all conditions are met before drawing a conclusion (strict conjunctive control). Current systems often rely on superficial word associations rather than true underlying logical structures, raising fundamental questions about whether prevailing neural network designs restrict reasoning capabilities.
This article evaluates how different neural network architectures perform deterministic logical deduction under distribution shifts. Specifically, it tests whether encoder-based models—which process whole inputs simultaneously into a global representation—exhibit greater reasoning stability and generalizability compared to autoregressive, decoder-only models that generate answers sequentially step-by-step.
The authors constructed a controlled synthetic framework based on first-order logic rules, establishing a training dataset of 40,000 samples (reasoning depths 0 to 7) and two distinct out-of-distribution test sets of 3,600 samples each (reasoning depths 0 to 11). The first test set used natural language expressions with deeper reasoning chains, while the second replaced natural words with randomized character strings to eliminate lexical shortcuts. The evaluation benchmarked both zero- and few-shot in-context learning across proprietary and open decoder models (including GPT-4.1, GPT-5, Claude Opus 4.1, and Qwen variants) as well as fine-tuned encoder (BERT), encoder–decoder (BART, Flan-T5), and decoder-only architectures.
The investigation produced four central findings. First, standard in-context prompting proved insufficient for dependable deduction; without task-specific fine-tuning, smaller models exhibited near-random discrimination, and performance barely changed between zero-shot and five-shot prompts. Second, fine-tuned encoder and encoder–decoder architectures generalized more effectively to deeper reasoning paths and abstract symbols than comparable fine-tuned decoders. In the non-natural language test, fine-tuned BERT achieved 61% accuracy, whereas the fine-tuned decoder Qwen3-1.7B dropped to near-random performance (53% accuracy) by depth 3. Third, fine-tuned encoder-based models were exceptionally cost- and compute-efficient; BART-Base achieved its accuracy with an inference efficiency of 640 accuracy points per hour on standard hardware. Fourth, while massive frontier models such as GPT-5 achieved near-perfect accuracy across all depths, they exhibited the lowest operational efficiency (1.1 points per hour), requiring over an hour of API inference per accuracy point.
These findings demonstrate that deploying off-the-shelf decoder models via prompting introduces operational risks and potential reasoning failures when rules extend beyond familiar patterns. Mechanistic analysis indicates that bidirectional encoders maintain geometric consistency across multi-step inferences, whereas sequential decoders accumulate drift unless scaled to massive compute budgets. For organizations implementing rule-based and causal workflows, reliance on large decoder-only models via prompting carries severe cost and latency penalties that may not be justified.
Organizations should adopt targeted, fine-tuned encoder or encoder–decoder models (such as BERT or BART variants) for structured, short-horizon causal deduction to optimize cost, latency, and reliability. Highly expensive frontier reasoning models like GPT-5 should be reserved exclusively for scenarios requiring open-ended generation or deep chains that smaller systems cannot resolve. Where hybrid needs exist, decision-makers should pilot architectures that pair the structured representation capabilities of encoders with the generative interface of decoders.
These conclusions are bounded by the study's reliance on synthetic Horn-clause logic without disjunctions, focusing on formal deductive steps rather than real-world unstructured text. In addition, because proprietary frontier models were accessed via remote application interfaces on undisclosed hardware, precise runtime comparisons represent operational estimates. Decision-makers can place high confidence in the architectural trade-offs demonstrated, while exercising caution when translating synthetic deduction benchmarks directly to noisy, real-world domain data.
- Paper: How Transformers Learn Causal Structure with Gradient Descent, Eshaan Nichani et al. (2024). Provides the theoretical foundation for how transformer attention heads learn latent causal graphs via gradient descent, directly informing the source's study on architecture suitability for causal reasoning.
- Paper: BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information, Mehran Kazemi et al. (2023). Establishes a systematic comparison of encoder, encoder-decoder, and decoder-only architectures across fine-tuning and few-shot in-context learning in multi-hop reasoning tasks.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Analyzes what mechanisms drive in-context learning and why few-shot prompts often rely on surface heuristics rather than true task mappings.
- Paper: Training Large Language Models to Reason in a Continuous Latent Space, Shibo Hao et al. (2024). Investigates multi-hop reasoning inside continuous latent spaces, motivating the source's hypothesis that encoder-based latent projections enhance conjunctive reasoning.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). Exposes how decoder-only autoregressive models falter during multi-step reasoning by greedily following spurious branches rather than maintaining global planning control.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). Demonstrates how NLP models exploit superficial lexical overlap heuristics instead of learning valid logical inference, which the source addresses in causal reasoning contexts.
- Paper: Do Large Language Models Latently Perform Multi-Hop Reasoning?, Sohee Yang et al. (2024). Probes the mechanics and vulnerabilities of internal multi-hop reasoning steps within autoregressive foundation models.
- Paper: MetaICL: Learning to Learn In Context, Sewon Min et al. (2022). Introduces meta-tuning frameworks that contrast fine-tuning versus purely in-context learning behavior across out-of-distribution shifts.
- Paper: Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task, Alicia Curth et al. (2026). Extends the investigation into architectural limits on multi-hop relational reasoning by analyzing whether deep transformer layers process multi-hop dependencies adaptively.
- Paper: Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?, Zhiyuan Zeng 0004 et al. (2025). Critically evaluates whether sequential test-time scaling in decoder reasoning architectures overcomes the brittleness identified in the source.
- Paper: Understanding Reasoning from Pretraining to Post-Training, Jingyan Shen et al. (2026). Analyzes how pretraining versus post-training reinforcement learning impacts multi-hop problem-solving capabilities across different model scales.
- Paper: The Surprising Effectiveness of Test-Time Training for Few-Shot Learning, Ekin Akyrek et al. (2025). Applies targeted test-time parameter updates to mitigate the distribution-shift brittleness of standard in-context reasoning observed in the source.
