Position: Avoid Overstretching LLMs for every Enterprise Task
Kuldeep SinghAnson BastosIsaiah Onando Mulang
Demonstrates the theoretical limitations of monolithic language models in enterprise workflows and advocates for modular architectures that restrict models to interface roles while offloading knowledge and computation to symbolic systems.
Enterprise artificial intelligence deployments face a significant gap between high experimentation rates and limited measurable financial returns. While organizations widely test frontier large language models, attempts to scale them across high-volume, mission-critical operations frequently encounter steep inference costs, high latency, persistent hallucinations, and compliance challenges. The article evaluates why relying on monolithic language models—or distilling them directly into smaller standalone models—is fundamentally mismatched with structured business tasks. Its main objective is to demonstrate that enterprise workflows achieve superior reliability, scalability, and cost efficiency by decoupling language processing from knowledge storage and algorithmic computation.
The authors analyze enterprise workload characteristics through an information-theoretic and computational lens, synthesizing industry survey data, formal mathematical proofs, and architectural case studies. This approach models how tasks depend on evolving external knowledge and deterministic business rules, evaluating the theoretical limits of parametric language models versus modular, tool-integrated architectures.
The article establishes several key findings. First, finite-parameter language models have an inherent capacity ceiling, meaning they cannot fully encode dynamic, proprietary enterprise knowledge and will suffer an irreducible error floor on knowledge-dependent tasks. Second, standard model distillation techniques drop intermediate reasoning steps, creating an information bottleneck that leaves smaller models prone to brittle shortcuts and generalization failures on multi-step logic. Third, modular systems that restrict small language models strictly to structured extraction—while routing computation and data retrieval to external knowledge bases and deterministic engines—strictly outperform purely parametric models in both accuracy and provable reliability. Fourth, running monolithic frontier models for routine, structured workflows introduces excess computational capacity, generating linearly higher energy and compute costs without meaningful gains in accuracy.
These findings indicate that treating generative models as central reasoning engines introduces avoidable operational and financial risk into production systems. In contrast, modular architectures substantially lower per-transaction execution costs and localize system failures to specific extraction or data layers, making auditing, compliance monitoring, and performance telemetry straightforward. By converting generative tasks into structured interface queries, organizations avoid the severe risks associated with granting unconstrained generative agents end-to-end execution autonomy.
To move successfully from pilots to production, enterprise leaders should adopt a staged deployment strategy. Organizations should prioritize high-frequency, repeatable workflows—such as ticket routing, invoice parsing, and policy lookups—and use specialized small language models solely to convert unstructured inputs into structured schemas. Substantive business logic, database queries, and rule validations should execute in deterministic symbolic tools. Large frontier models should be shifted offline to generate extraction schemas, rule templates, and verification test suites. Monolithic models should remain strictly reserved for low-volume, open-ended, and highly conversational or exploratory tasks.
Leaders should note that this architectural transition relies on the existence of well-defined data schemas and structured workflows. In domains with ambiguous decision criteria or unconstrained creative tasks, externalization provides fewer efficiency benefits and may add interface complexity. Nonetheless, for high-volume and constraint-heavy enterprise operations, the theoretical and practical evidence strongly indicates that modular architectures provide a reliable and economically sustainable path forward.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). StructGPT introduces an iterative reading-then-reasoning framework that accesses external databases and knowledge graphs via structured interfaces, establishing the foundational architecture for externalizing computation and storage away from monolithic models.
- Paper: Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning, Saibo Geng et al. (2023). This paper establishes how grammar-constrained decoding enforces strict adherence to predefined syntax for structured extraction tasks, providing essential technical grounding for treating language models as specialized structured extraction interfaces.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This study demonstrates that scaling parametric memory in language models fails on long-tail factual queries compared to external retrieval, motivating the position that enterprise knowledge must be externalized into dedicated databases.
- Paper: Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs, Jinyang Li et al. (2023). The BIRD benchmark documents the severe failure modes and inefficiencies of monolithic LLMs acting as direct database interfaces without external domain grounding, underpinning the source's enterprise critiques.
- Paper: SGLang: Efficient Execution of Structured Language Model Programs, Lianmin Zheng et al. (2023). SGLang provides the programmatic primitives and runtime caching infrastructure necessary to execute deterministic, structured language model programs alongside external tools.
- Paper: Rethinking Tabular Data Understanding with Large Language Models, Tianyang Liu et al. (2024). This work reveals how LLMs struggle with structural tabular variations and benefit from external code-based normalization, supporting the need for symbolic enterprise workflows.
- Paper: Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks, Minki Kang et al. (2023). This paper demonstrates that augmenting compact models with external knowledge bases overcomes finite memorization bottlenecks in knowledge-intensive enterprise tasks.
- Paper: DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction, Mohammadreza Pourreza et al. (2023). DIN-SQL shows that modular task decomposition and intermediate symbolic generation outperform monolithic prompt strategies in enterprise database environments.
- Paper: LLMs Corrupt Your Documents When You Delegate, Philippe Laban et al. (2026). This empirical study validates the risks of over-delegating complex tasks to monolithic LLMs by quantifying widespread silent document corruption across long-horizon enterprise interactions.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). LEGOMem implements the source's modular philosophy by externalizing procedural memory and workflow execution across specialized multi-agent systems for office automation.
- Paper: Human-Inspired Memory Architecture for LLM Agents, Doga Kerestecioglu et al. (2026). This paper realizes the architectural separation of memory from model inference by deploying an external, tiered biological memory system with semantic knowledge graphs for enterprise agents.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). This survey extends the source's efficiency arguments by cataloging concrete methods to modularize agent memory, tool learning, and planning to reduce runtime overhead.
- Paper: Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills, Agamdeep Singh et al.. This research provides a practical method for amortizing recurring reasoning costs into compact prompt-distilled skills, advancing cost-effective enterprise agent architectures.
- Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). This study tests the operational limits of monolithic LLMs in multi-turn interactive workflows, reinforcing the need for externalized state and intent tracking.
- Paper: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models, Qizheng Zhang et al. (2026). Agentic Context Engineering translates modular context management into self-improving, externalized structured playbooks for enterprise and domain-specific workflows.
