Position: Avoid Overstretching LLMs for every Enterprise Task

Kuldeep SinghAnson BastosIsaiah Onando Mulang

article2026arXiv0 citations

Demonstrates the theoretical limitations of monolithic language models in enterprise workflows and advocates for modular architectures that restrict models to interface roles while offloading knowledge and computation to symbolic systems.

Listen

Enterprise artificial intelligence deployments face a significant gap between high experimentation rates and limited measurable financial returns. While organizations widely test frontier large language models, attempts to scale them across high-volume, mission-critical operations frequently encounter steep inference costs, high latency, persistent hallucinations, and compliance challenges. The article evaluates why relying on monolithic language models—or distilling them directly into smaller standalone models—is fundamentally mismatched with structured business tasks. Its main objective is to demonstrate that enterprise workflows achieve superior reliability, scalability, and cost efficiency by decoupling language processing from knowledge storage and algorithmic computation.

The authors analyze enterprise workload characteristics through an information-theoretic and computational lens, synthesizing industry survey data, formal mathematical proofs, and architectural case studies. This approach models how tasks depend on evolving external knowledge and deterministic business rules, evaluating the theoretical limits of parametric language models versus modular, tool-integrated architectures.

The article establishes several key findings. First, finite-parameter language models have an inherent capacity ceiling, meaning they cannot fully encode dynamic, proprietary enterprise knowledge and will suffer an irreducible error floor on knowledge-dependent tasks. Second, standard model distillation techniques drop intermediate reasoning steps, creating an information bottleneck that leaves smaller models prone to brittle shortcuts and generalization failures on multi-step logic. Third, modular systems that restrict small language models strictly to structured extraction—while routing computation and data retrieval to external knowledge bases and deterministic engines—strictly outperform purely parametric models in both accuracy and provable reliability. Fourth, running monolithic frontier models for routine, structured workflows introduces excess computational capacity, generating linearly higher energy and compute costs without meaningful gains in accuracy.

These findings indicate that treating generative models as central reasoning engines introduces avoidable operational and financial risk into production systems. In contrast, modular architectures substantially lower per-transaction execution costs and localize system failures to specific extraction or data layers, making auditing, compliance monitoring, and performance telemetry straightforward. By converting generative tasks into structured interface queries, organizations avoid the severe risks associated with granting unconstrained generative agents end-to-end execution autonomy.

To move successfully from pilots to production, enterprise leaders should adopt a staged deployment strategy. Organizations should prioritize high-frequency, repeatable workflows—such as ticket routing, invoice parsing, and policy lookups—and use specialized small language models solely to convert unstructured inputs into structured schemas. Substantive business logic, database queries, and rule validations should execute in deterministic symbolic tools. Large frontier models should be shifted offline to generate extraction schemas, rule templates, and verification test suites. Monolithic models should remain strictly reserved for low-volume, open-ended, and highly conversational or exploratory tasks.

Leaders should note that this architectural transition relies on the existence of well-defined data schemas and structured workflows. In domains with ambiguous decision criteria or unconstrained creative tasks, externalization provides fewer efficiency benefits and may add interface complexity. Nonetheless, for high-volume and constraint-heavy enterprise operations, the theoretical and practical evidence strongly indicates that modular architectures provide a reliable and economically sustainable path forward.

arXiv: 2605.09365
Cover for Position: Avoid Overstretching LLMs for every Enterprise Task

Abstract

Enterprise workloads are dominated by deterministic, structured, and knowledge-dependent tasks operating under strict cost, latency, and reliability constraints. While these are often addressed through large language model (LLM) deployment or distillation into smaller models, we argue this is inefficient, unreliable, and misaligned with enterprise task structures. Instead, AI systems should treat language models as interfaces rather than monolithic engines, externalizing knowledge and computation into dedicated components for greater reliability, scalability, and transparency. Our theoretical evidences show that finite-capacity models cannot fully capture the breadth of knowledge required for enterprise tasks, creating inherent limits to efficiency and interpretability. Building on this, we take the position that language models should primarily be used for structured extraction in deterministic enterprise workflows, while computation and storage are delegated to knowledge bases and symbolic procedures. We formally demonstrate that such modular architectures are more reliable and maintainable than monolithic frameworks, offering a sustainable foundation for enterprise tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Theoretical Formalization of the Enterprise AI Design Choice
  • 3.1 A Theoretical Separation: Symbolic Reasoning with SLM Extraction and External Knowledge
  • 3.1.1 Problem Setup
  • 3.1.2 Systems Under Comparison
  • 3.1.3 Strict Advantage from Access to External Information
  • 3.1.4 Strict Advantage from Computational (Expressivity) Separation
  • 3.1.5 Discussion: Why Rationale-Free Distillation Is Especially Limited
  • 3.2 Why monolithic LLMs (even with with tools and RAGs) are wasteful vis a vis specialized SLMs: A Theory of Modular Decomposition, Capacity Allocation, and Efficiency
  • 3.3 Entropic decomposition of enterprise tasks
  • 3.4 Separation of concerns and complexity reduction
  • 3.5 Excess capacity and inefficiency of monolithic LLMs
  • 3.5.1 Efficiency gain from modular architectures
  • 3.6 When to prioritize task specific SLMs over LLMs: A Decision-Theoretic Framework for System Selection
  • 3.6.1 Total cost formulation
  • 3.6.2 Break-even analysis
  • 3.7 Selection regimes
  • 4 Position Statements
  • 5 Roadmap: Externalizing Knowledge and Computation for Enterprise AI
  • 5.1 Architectural decomposition: interface, knowledge, and computation
  • 5.2 Deployment pathway: from pilots to production
  • 5.3 Evaluation and telemetry: making failure modes actionable
  • 5.4 Offline use of frontier LLMs to synthesize symbolic assets
  • 5.5 Comparison with standard architecture of LLMs with tool usage and RAG/memory
  • 5.6 Formal Foundations of the Modular Roadmap
  • 5.7 Interpretation and practitioner guidance
  • 6 Alternative Views and Our Responses
  • 7 Conclusion
  • References
  • A Technical Appendix
  • B An Information-Theoretic Theory of Language Model Computational Ability
  • B.1 Task Information Complexity
  • B.2 Model Information Capacity
  • B.3 Mutual Information Between Model and Task
  • B.4 Limits on Generalization and Emergence
  • B.5 Limits on Computation
  • B.6 Unified Characterization
  • C Interpretation: Empirical Limitations of SLMs Through the Lens of Information Theory
  • C.1 Limitations in Reasoning, Compositionality, and Emergence
  • C.2 Limitations in Out-of-Domain Generalization
  • C.3 Limits on Computation Itself
  • C.4 Compression, Pruning, and Quantization
  • C.5 Data-Optimal Training as a Partial Remedy
  • C.6 Summary
  • C.7 Agentic Decomposition and Specialized SLMs
  • D Knowledge Distillation for Language Models: Rationale-Free Distillation and Its Limits
  • D.1 What We Mean by Distillation
  • D.2 Distillation Without Rationale
  • D.3 Why We Use Distillation Without Rationale
  • D.4 An Information-Theoretic View: What Is Lost Without Rationale
  • D.5 Limitations of Distillation Without Rationale
  • D.6 When Rationale-Free Distillation Is Still Appropriate
  • D.7 Enterprise Summary and Practical Implication

Knowls

  1. Knowl 1 — Strict Advantage of Hybrid Symbolic-SLM Architectures Over Purely Parametric Models

    theoretical result

    Let (X,Y)∼D(X, Y) \sim \mathcal{D} denote an enterprise task input-output distribution where the ground-truth target satisfies Y=h(X,Z)Y = h(X, Z), with ZZ representing minimal task-relevant external knowledge retrieved from source K\mathcal{K}.

    1. External Information Separation: If the task depends nontrivially on external knowledge such that conditional mutual information satisfies: I(Y;Z∣X)>0I(Y; Z \mid X) > 0 then any rationale-free distilled small language model (SLM) SWS_W parameterized by WW producing Y^S∼pW(⋅∣X)\hat{Y}_S \sim p_W(\cdot \mid X) without access to K\mathcal{K} incurs an irreducible strictly positive Bayes risk: RX∗:=inf⁡y^:X→YPr⁡[y^(X)≠Y]>0R_X^* := \inf_{\hat{y}: \mathcal{X} \to \mathcal{Y}} \Pr[\hat{y}(X) \ne Y] > 0 For a hybrid system Y^H=R(ϕ^(X),K)\hat{Y}_H = R(\hat{\phi}(X), \mathcal{K}), where ϕ^\hat{\phi} is an SLM extraction interface approximating ideal extraction ϕ∗(X)\phi^*(X) and RR is a sound deterministic symbolic reasoner, system risk is bounded by extraction failure Pr⁡[Y^H≠Y]≤Pr⁡[ϕ^(X)≠ϕ∗(X)]\Pr[\hat{Y}_H \ne Y] \le \Pr[\hat{\phi}(X) \ne \phi^*(X)]. If extraction failure satisfies Pr⁡[ϕ^(X)≠ϕ∗(X)]<RX∗\Pr[\hat{\phi}(X) \ne \phi^*(X)] < R_X^*, the hybrid architecture strictly dominates any parametric distilled model: Pr⁡[Y^H≠Y]<inf⁡WPr⁡[Y^S≠Y]\Pr[\hat{Y}_H \ne Y] < \inf_W \Pr[\hat{Y}_S \ne Y]

    2. Computational Expressivity Separation: Let CSLM(B)\mathcal{C}_{SLM}(B) be the class of functions representable by SLMs of capacity BB, and let CR\mathcal{C}_R be functions computable by symbolic reasoner RR with access to K\mathcal{K}. If there exists a target function f∗∈CRf^* \in \mathcal{C}_R such that f∗∉CSLM(B)f^* \notin \mathcal{C}_{SLM}(B), then there exists a distribution D\mathcal{D} and constant ϵ>0\epsilon > 0 such that any rationale-free SLM of capacity BB incurs error at least ϵ\epsilon: inf⁡W:SW∈CSLM(B)Pr⁡[Y^S≠f∗(X,K)]≥ϵ\inf_{W: S_W \in \mathcal{C}_{SLM}(B)} \Pr[\hat{Y}_S \ne f^*(X, \mathcal{K})] \ge \epsilon When extraction failure satisfies Pr⁡[ϕ^(X)≠ϕ∗(X)]<ϵ\Pr[\hat{\phi}(X) \ne \phi^*(X)] < \epsilon, the hybrid system achieves strictly superior performance over any bounded-capacity parametric SLM on structured computational tasks.

  2. Knowl 2 — Topological and Metric Obstruction to Internalizing World Knowledge in LLMs

    theoretical result

    Let the parameter space of a language model be modeled as a finite-dimensional Euclidean space P:=(Rd,∥⋅∥2)\mathcal{P} := (\mathbb{R}^d, \|\cdot\|_2) of dimension d<∞d < \infty. Let the space of enterprise world knowledge be modeled as the infinite product space Kworld:={0,1}N\mathcal{K}_{world} := \{0, 1\}^\mathbb{N} equipped with the standard ultrametric: ρ(s,t)=∑n=1∞2−n1{sn≠tn},for s=(sn)n≥1,t=(tn)n≥1\rho(s, t) = \sum_{n=1}^\infty 2^{-n} \mathbf{1}\{s_n \ne t_n\}, \quad \text{for } s=(s_n)_{n\ge 1}, t=(t_n)_{n\ge 1} A metric space (M,dM)(M, d_M) is doubling if every ball of radius r>0r > 0 can be covered by at most NN balls of radius r/2r/2, with doubling dimension log⁡2N\log_2 N.

    1. (Rd,∥⋅∥2)(\mathbb{R}^d, \|\cdot\|_2) is doubling with doubling constant at most 2O(d)2^{O(d)}.
    2. The metric space (Kworld,ρ)(\mathcal{K}_{world}, \rho) has infinite doubling dimension and is not doubling.
    3. For any fixed probing extractor with Lipschitz continuous forward maps, the induced parameter-to-knowledge encoding map Φ:P→Kworld\Phi: \mathcal{P} \to \mathcal{K}_{world} is globally LL-Lipschitz continuous, and the image KLLM:=Φ(Rd)\mathcal{K}_{LLM} := \Phi(\mathbb{R}^d) is doubling.

    Because no doubling space can cover a non-doubling metric space, the knowledge retrievable from model parameters is a strict subset of world knowledge: KLLM=Φ(Rd)⊊Kworld\mathcal{K}_{LLM} = \Phi(\mathbb{R}^d) \subsetneq \mathcal{K}_{world} This establishes a metric-topological obstruction proving that finite-dimensional parametric models cannot fully internalize unconstrained world knowledge.

  3. Knowl 3 — Information-Theoretic Capacity and Generalization Bounds for Language Models

    theoretical result

    For a task distribution TT mapping input XX to ground-truth output Y=f∗(X)Y = f^*(X):

    • Task Information Complexity: The intrinsic complexity of task TT is defined as: Itask:=I(Y;f∗∣X)I_{task} := I(Y; f^* \mid X) measuring the minimal bits an agent must encode about the input-output mapping to guarantee bounded error.
    • Model Information Capacity: For a model parameterized by W∈RdW \in \mathbb{R}^d stored with finite precision of BB bits (with Shannon entropy H(W)≤BH(W) \le B), its capacity is: Imodel:=I(W;HW)≤BI_{model} := I(W; \mathcal{H}_W) \le B where HW\mathcal{H}_W is the hypothesis class induced by WW.

    The model's computational and generalization abilities obey the following bounds:

    1. Solvability: Task TT is solvable by the model if and only if Imodel≥ItaskI_{model} \ge I_{task}.
    2. Out-of-Distribution (OOD) Generalization Limit: For in-distribution function fin∗f^*_{in} and out-of-distribution function fOOD∗f^*_{OOD}, the extra information demand is ΔIOOD=Itask,OOD−Itask,in\Delta I_{OOD} = I_{task, OOD} - I_{task, in}. If Itask,OOD>ImodelI_{task, OOD} > I_{model}, OOD generalization and emergent reasoning outside training distribution are provably impossible since I(W;fOOD∗)≈I(W;fin∗)≤ImodelI(W; f^*_{OOD}) \approx I(W; f^*_{in}) \le I_{model}.
    3. Output Mutual Information Bound: For model outputs generated via distribution pW(y∣x)p_W(y \mid x), the data-processing inequality implies: I(outputs;Y)≤min⁡(Imodel,Itask)I(\text{outputs}; Y) \le \min(I_{model}, I_{task}) If Imodel<ItaskI_{model} < I_{task}, the model cannot faithfully represent the required computation even on-distribution.
  4. Knowl 4 — Information Efficiency of Specialized SLMs in Decomposable Agentic Workflows

    theoretical result

    Consider an agentic enterprise workflow composed of mm stages j∈{1,…,m}j \in \{1, \dots, m\}, where stage jj processes local input XjX_j and local external tool state ZjZ_j to compute output Yj=hj(Xj,Zj)Y_j = h_j(X_j, Z_j), with stage information requirement: Ij:=I(Yj;Zj∣Xj)I_j := I(Y_j; Z_j \mid X_j) The total information required across the workflow is Iagent=∑j=1mIjI_{agent} = \sum_{j=1}^m I_j. The workflow is decomposable if its total loss factorizes as: L(W1,…,Wm)=∑j=1mλjLj(Wj)\mathcal{L}(W_1, \dots, W_m) = \sum_{j=1}^m \lambda_j \mathcal{L}_j(W_j) where each stage jj can be evaluated using local state and external memory without requiring the parameter vector of any other stage (with weights λj>0\lambda_j > 0).

    1. Information Allocation Advantage: If a monolithic controller has parameter capacity ImodelsharedI_{model}^{shared} satisfying: ∑j=1mIj>Imodelshared\sum_{j=1}^m I_j > I_{model}^{shared} no single shared model can execute the full workflow without irreducible error on at least one stage. In contrast, modular specialized SLMs SjS_j, each possessing capacity Imodel(j)≥IjI_{model}^{(j)} \ge I_j and connected through external tools, can represent every stage without error.

    2. Optimization Advantage: Independent optimization of specialized stage models yields: inf⁡W1,…,Wm∑j=1mλjLj(Wj)≤inf⁡W∑j=1mλjLj(W)\inf_{W_1, \dots, W_m} \sum_{j=1}^m \lambda_j \mathcal{L}_j(W_j) \le \inf_W \sum_{j=1}^m \lambda_j \mathcal{L}_j(W) with strict inequality whenever stage-wise optimal parameters are mutually incompatible.

  5. Knowl 5 — Projection Bottleneck and Double Compression in Rationale-Free Distillation

    theoretical result

    In rationale-free knowledge distillation, a student model SS with parameters θS\theta_S is trained on pairs (x,yT)(x, y_T) to match the predictive distribution of a teacher TT via Kullback-Leibler divergence Ex∼D[KL(pT(⋅∣x)∥pS(⋅∣x))]\mathbb{E}_{x \sim \mathcal{D}}[\text{KL}(p_T(\cdot \mid x) \parallel p_S(\cdot \mid x))] without observing intermediate reasoning traces RR.

    Assuming the teacher's internal computation follows the causal graph X→R→YX \to R \to Y with predictive distribution pT(y∣x)=∑rpT(y∣r,x)pT(r∣x)p_T(y \mid x) = \sum_r p_T(y \mid r, x) p_T(r \mid x), omitting RR induces fundamental limitations:

    1. Projection Bottleneck: Unless target output YY is a sufficient statistic for reasoning trace RR, residual uncertainty remains: H(R∣X,Y)>0andI(R;f∗∣X,Y)>0H(R \mid X, Y) > 0 \quad \text{and} \quad I(R; f^* \mid X, Y) > 0 where f∗f^* is the task rule. The transferable information about the teacher's underlying reasoning algorithm is bounded by the mutual information of the output channel: I(student parameters;f∗)≲I(YT;f∗∣X)I(\text{student parameters}; f^*) \lesssim I(Y_T; f^* \mid X)

    2. Double Compression Penalty: The student experiences two consecutive information compressions: (i) the teacher compresses algorithmic state RR into final response YY, and (ii) the student compresses the resulting mapping into a reduced parameter budget WW where Imodel<ItaskI_{model} < I_{task}.

    This double compression incentivizes student SLMs to learn superficial, brittle shortcuts rather than faithful algorithmic state transitions, leading to severe failure under distribution shifts, compositional queries, and out-of-distribution reasoning.

  6. Knowl 6 — Entropic Uncertainty Decomposition of Enterprise Workloads

    model/method

    For an enterprise task where input XX and external knowledge ZZ produce output Y=h(X,Z)Y = h(X, Z) with I(Y;Z∣X)>0I(Y; Z \mid X) > 0, the conditional entropy of the output given input decomposes into four distinct uncertainty terms: H(Y∣X)=HKNOW+HCOMP+HLANG+ΔH(Y \mid X) = H_{KNOW} + H_{COMP} + H_{LANG} + \Delta where:

    • HKNOWH_{KNOW} is uncertainty due to missing, private, or dynamically evolving enterprise knowledge.
    • HCOMPH_{COMP} is uncertainty arising from algorithmic, procedural, or rule-based multi-step computation.
    • HLANGH_{LANG} is linguistic uncertainty in mapping unstructured natural language inputs to formal structured schemas.
    • Δ\Delta represents residual interaction effects among knowledge, computation, and language.

    In monolithic LLMs, all four terms must be resolved implicitly inside the network weights. Introducing external retrieval and symbolic tools provides evidence EE, dramatically reducing uncertainty: H(Y∣X,E)≪H(Y∣X)H(Y \mid X, E) \ll H(Y \mid X) Because retrieval eliminates HKNOWH_{KNOW} and deterministic tools eliminate HCOMPH_{COMP}, the system reaches an interface-dominated regime: H(Y∣X,E)≈HLANGH(Y \mid X, E) \approx H_{LANG} Here, task complexity drops from full end-to-end complexity ITASK(END-TO-END)I_{TASK}^{(END\text{-}TO\text{-}END)} to extraction complexity ITASK(INTERFACE)≪ITASK(END-TO-END)I_{TASK}^{(INTERFACE)} \ll I_{TASK}^{(END\text{-}TO\text{-}END)}, allowing lightweight SLMs to perform the mapping without loss of accuracy.

  7. Knowl 7 — Decision-Theoretic Cost Framework and Break-Even Selection Criterion

    model/method

    To decide between deploying a monolithic LLM or a modular SLM-tool architecture, the total expected operational cost over NN requests across a deployment horizon is modeled as: J=CBUILD+NCRUN+NλR+COPSJ = C_{BUILD} + N C_{RUN} + N \lambda R + C_{OPS} where:

    • CBUILDC_{BUILD} is initial system development and integration cost.
    • CRUNC_{RUN} is execution cost per request (CRUN(M)=mTc(M)C_{RUN}(M) = m T c(M) for mm model invocations, TT tokens per invocation, and token cost c(M)c(M)).
    • R∈[0,1]R \in [0, 1] is the expected task failure rate.
    • λ\lambda is the financial or operational penalty per failure.
    • COPSC_{OPS} is ongoing operational maintenance overhead.

    Letting ΔCBUILD=CBUILDmodular−CBUILDmonolithic\Delta C_{BUILD} = C_{BUILD}^{modular} - C_{BUILD}^{monolithic}, ΔCRUN=CRUNmonolithic−CRUNmodular\Delta C_{RUN} = C_{RUN}^{monolithic} - C_{RUN}^{modular}, and ΔR=Rmonolithic−Rmodular\Delta R = R^{monolithic} - R^{modular}, a modular SLM-tool architecture is economically and operationally superior when: ΔCBUILD<N(ΔCRUN+λΔR)\Delta C_{BUILD} < N (\Delta C_{RUN} + \lambda \Delta R) The break-even request volume N∗N^* is given by: N∗=ΔCBUILDΔCRUN+λΔRN^* = \frac{\Delta C_{BUILD}}{\Delta C_{RUN} + \lambda \Delta R} Modular SLM frameworks strictly dominate when workloads are high-volume (N≫N∗N \gg N^*), error costs λ\lambda are high, external knowledge dependence I(Y;Z∣X)>0I(Y; Z \mid X) > 0 is significant, or underlying business logic evolves frequently.

  8. Knowl 8 — Three-Layer Architectural Decomposition for Governed Enterprise AI

    model/method

    To resolve the reliability and scaling failure modes of monolithic LLM orchestrators, enterprise AI systems are architected into three decoupled layers:

    1. Interface Layer (SLM Extraction): Task-specialized small language models (SLMs) map raw, unstructured enterprise inputs (emails, tickets, policy documents, logs) into well-defined structured formats (normalized slots, entity-relation graphs, database query templates, and API payloads). Their scope is constrained to deterministic structured extraction, eliminating generative hallucination in decision logic.
    2. Knowledge Layer (External Knowledge Sources): Enterprise facts, business policies, and domain data reside in external, auditable structures (knowledge graphs, document repositories, relational databases). This ensures access control, live data updates, and enterprise security without requiring model retraining or weight updates.
    3. Computation Layer (Discoverable Formal Methods): Deterministic execution, symbolic reasoning, constraint satisfaction, and formal verification are delegated to specialized symbolic code engines and solvers. These tools are versioned, indexed, and invoked through explicit interfaces.

    Frontier LLMs operate strictly offline as schema and rule synthesizers (drafting candidate ontologies, verification checklists, and canonical query templates for validation by domain experts), keeping online runtime loops fast, auditable, and low-cost.

  9. Knowl 9 — Four-Stage Deployment Pathway for Scaling Enterprise AI to Production

    model/method

    To bridge the enterprise AI "pilot-to-production" gap, systems should transition through four sequential deployment stages rather than defaulting to monolithic end-to-end LLMs:

    • Stage 1: High-Frequency Repeatable Workflows. Implement SLMs on routine, high-volume enterprise tasks characterized by rigid schemas and unambiguous success metrics (e.g., ticket triage, compliance verification, invoice parsing, runbook step selection).
    • Stage 2: Knowledge Grounding and Constraint Enforcement. Externalize domain memory into structured knowledge bases and retrieval indices early, eliminating parametric hallucinations and enabling access control.
    • Stage 3: Verification Modules and Human-in-the-Loop Protocols. Position formal rule verifiers, schema validators, and human review checkpoints at critical operational boundaries where failure penalties are severe (e.g., automated transaction approval, security actions).
    • Stage 4: Governed Agentic Workflows with Explicit Risk Controls. Construct multi-step agentic workflows by routing subtasks across narrow, specialized SLMs (e.g., Extraction SLM, Routing SLM, Tool-Aggregator SLM, Verification SLM) coordinated by an explicit controller over bounded action spaces and audited execution paths.
  10. Knowl 10 — Task Suitability Regimes for Modular SLM versus Monolithic LLM Frameworks

    data/table

    Enterprise workloads exhibit varying suitability for modular SLM-tool frameworks versus monolithic LLMs depending on whether task uncertainty is reducible via external knowledge and symbolic computation:

    Task Category Suitable for Modular SLM Error-Prone / Less Suitable Reasoning
    Structured data extraction Invoice parsing, log extraction, ticket field normalisation Free-form narrative summarisation across domains Well-defined schemas allow small extraction models to suffice; unstructured tasks increase interface entropy.
    Knowledge-grounded QA Enterprise KB querying, policy lookup, compliance checks Open-ended advisory answers without grounding Tasks with explicit sources satisfy I(Y;Z∣X)>0I(Y; Z \mid X) > 0, favoring retrieval and deterministic resolution.
    Workflow automation Ticket routing, runbook selection, approval flows Complex decision-making with evolving ambiguous criteria Stable workflows with explicit rules reduce failures strictly to extraction errors.
    Algorithmic computation Pricing rules, validation checks, constraint enforcement Multi-step reasoning with unclear decomposition Deterministic computation benefits from external execution engines exceeding bounded model capacity.
    Agentic orchestration Tool routing, parameter generation, API planning Long-horizon planning with emergent subgoals Decomposable workflows allow stage-specific SLMs; implicit long-term state abstraction challenges modularization.
    High-volume production tasks Repetitive queries, monitoring, alert triage Low-frequency exploratory queries Modular systems amortise upfront build cost and minimize per-query inference expense at high scale.
    Governance-sensitive tasks Compliance auditing, regulated decision pipelines Informal or creative reasoning tasks Externalisation provides auditability, traceability, and verifiable control surfaces.

    The fundamental design principle is that modular SLM-tool pipelines are optimal when post-retrieval task uncertainty collapses to a low-entropy structured interface mapping (H(Y∣X,E)≈HLANGH(Y \mid X, E) \approx H_{LANG}), whereas monolithic LLMs remain necessary for open-ended, unformalized linguistic synthesis.

Coverage note — The Kolmogorov-complexity and Shannon-capacity bounds from Appendix A (Theorems 1 and 2) were omitted as their substantive conclusions are subsumed by the topological doubling-dimension theorem and the information-theoretic capacity bounds.

References

  1. 1.Patrice Béchard and Orlando Marquez Ayala. Reducing hallucination in structured outputs via retrieval-augmented generation. In NAACL (Industry Track), 2024.
  2. 2.Satwik Bhattamishra, Michael Hahn, Phil Blunsom, and Varun Kanade. Separations in the representational capabilities of transformers and recurrent architectures. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  3. 3.BizzDesign. Enterprise ai adoption: Balancing innovation and roi in 2026. BizzDesign blog, 2026. Published Jan 27, 2026.
  4. 4.Deloitte. Ai trends 2025: Adoption barriers and updated predictions. Deloitte US AI Pulse Check Series, 2025.
  5. 5.Deloitte. State of ai in the enterprise 2026. Deloitte Global, 2026. Survey of 3,235 senior leaders across 24 countries, conducted Aug–Sep 2025.
  6. 6.Aladin Djuhera, Swanand Ravindra Kadhe, Syed Zawad, Farhan Ahmed, Heiko Ludwig, and Holger Boche. Fixing it in post: A comparative study of llm post-training data quality and model performance. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025.
  7. 7.Ernst & Young. Ai adoption outpaces governance: Responsible ai pulse survey. EY Global Survey, 2025.
  8. 8.Gartner. Gartner predicts over 40% of agentic ai projects will be canceled by end of 2027. Gartner Press Release, 2025.
  9. 9.Š. Hána and B. Lameijer. Ai-based systems adoption in business operations: barriers and performance effects. Operations Management Research, 2025.
  10. 10.Harvard Business Review. Overcoming the organizational barriers to ai adoption. Harvard Business Review article, 2025. Published Nov 11, 2025.
  11. 11.Information Services Group (ISG). State of enterprise ai adoption report 2025. ISG report, 2025. Analysis of 1,200 generative, agentic, and traditional AI use cases.
  12. 12.E. Karakurt and A. Akbulut. Retrieval-augmented generation and large language models for enterprise knowledge management: A systematic literature review. Applied Sciences, 2025.
  13. 13.Tassilo Klein and Johannes Hoffart. Foundation models for tabular data within systemic contexts need grounding, 2025.
  14. 14.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021.
  15. 15.Andy Lo, Albert Q Jiang, Wenda Li, and Mateja Jamnik. End-to-end ontology learning with large language models. Advances in Neural Information Processing Systems, 37:87184–87225, 2024.
  16. 16.Zhongyuan Lyu, Shuoyu Hu, Lujie Liu, Hongxia Yang, and Ming LI. Canonical intermediate representation for llm-based optimization problem formulation and code generation. arXiv preprint arXiv:2602.02029, 2026.
  17. 17.McKinsey & Company. The state of ai: Global survey 2025, 2025.
  18. 18.Menlo Ventures. 2025: The state of ai in healthcare. Menlo Ventures report, 2025. Survey of 700+ healthcare executives, conducted Aug–Sep 2025.
  19. 19.OpenAI. The state of enterprise ai. OpenAI report, 2025. Based on usage telemetry from 9,000 workers across nearly 100 enterprises.
  20. 20.OpenText, Capgemini, and Sogeti. merging worlds. Industry survey report, OpenText Corporation, 2025. Survey of enterprise leaders on AI adoption in quality engineering practices. Reports that while nearly 90% of organizations pursue generative AI in quality engineering, only 15% have achieved enterprise-scale deployment. Top barriers include data privacy risks (67%), integration complexity (64%), and hallucination/reliability concerns (60%).
  21. 21.Pepper Foster. The artificial intelligence (ai) roi report. Pepper Foster report, 2025. Cites MIT study “The GenAI Divide: State of AI in Business 2025”.
  22. 22.E. Romeo and J. Lacko. Adoption and integration of ai in organizations: a systematic review of challenges and drivers. Kybernetes, 2025.
  23. 23.S&P Global Sustainable1. Ai adoption is soaring, but few companies are measuring its impact. S&P Global Insights, 2025.
  24. 24.Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin. What formal languages can transformers express? a survey. Transactions of the Association for Computational Linguistics, 2024.
  25. 25.The Guardian. Claude ai agent deletes firm database in seconds. The Guardian, April 2026. Accessed: 2026-05-02.
  26. 26.Vistage and Wharton Human-AI Research / GBK Collective. Enterprise ai adoption and roi: Three-year executive study. Wharton/Vistage report, 2025. Survey of 800 US executives, June 26–July 11, 2025.
  27. 27.WalkMe. The state of digital adoption 2025 (special ai edition). WalkMe Research Report, 2025.
  28. 28.Samuel Yeh, Sharon Li, and Tanwi Mallick. Lumina: Detecting hallucinations in rag system with context–knowledge signals. In Socially Responsible and Trustworthy Foundation Models at NeurIPS 2025, 2025.
  29. 29.Yedi Zhang, Yufan Cai, Xinyue Zuo, Xiaokun Luan, Kailong Wang, Zhe Hou, Yifan Zhang, Zhiyuan Wei, Meng Sun, Jun Sun, et al. Position: Trustworthy ai agents require the integration of large language models and formal methods. In Forty-second International Conference on Machine Learning Position Paper Track, 2025.

Citation

MLA
Singh, K., et al. “Position: Avoid Overstretching LLMs for Every Enterprise Task”. arXiv, 2026, http://arxiv.org/abs/2605.09365v1.
APA
Singh, K., Bastos, A., & Mulang', I. O. (2026). Position: Avoid Overstretching LLMs for every Enterprise Task. arXiv. http://arxiv.org/abs/2605.09365v1
Chicago
Singh, K., A. Bastos, and I. O. Mulang'. 2026. “Position: Avoid Overstretching LLMs for Every Enterprise Task”. arXiv. http://arxiv.org/abs/2605.09365v1.
Harvard
Singh, K., Bastos, A. and Mulang', I.O. (2026) “Position: Avoid Overstretching LLMs for every Enterprise Task”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.09365v1.
Vancouver
1. Singh K, Bastos A, Mulang' IO (2026) Position: Avoid Overstretching LLMs for every Enterprise Task. arXiv

BibTeX

@article{singh2026position,
  title = {Position: Avoid Overstretching LLMs for every Enterprise Task},
  author = {Singh, Kuldeep and Bastos, Anson and Mulang', Isaiah Onando},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.09365v1},
  eprint = {2605.09365}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/