LLM Augmented LLMs: Expanding Capabilities through Composition
Rachit BansalBidisha SamantaSiddharth DalmiaNitish GuptaSriram GanapathyAbhishek BapnaPrateek JainPartha Talukdar
Introduces CALM, an efficient framework that combines frozen large language models with specialized smaller models via cross-attention to acquire new capabilities in coding and low-resource translation without retraining the base model.
Large language models excel at broad reasoning, world knowledge, and fluent communication, but expanding their capabilities to new specialized domains remains difficult and expensive. Retraining or fine-tuning massive models requires vast computing budgets, risks catastrophic forgetting of existing skills, and often faces data privacy constraints across organizational boundaries. Simpler alternatives like basic model merging or output routing fall short when neither model can solve a composite task independently.
The article demonstrates an efficient framework called Composition to Augment Language Models (CALM). The objective is to combine a large, general-purpose anchor model with a smaller, domain-specialized augmenting model to enable new composite capabilities while keeping both original models completely frozen.
The approach introduces a lightweight set of learnable projection and cross-attention parameters between selected intermediate layers of the two models. This structure allows the anchor model to dynamically attend to internal representations from the specialized model during text generation. To evaluate this method, the authors conducted experiments across three diverse domains: synthetic key-value arithmetic, low-resource language translation and mathematical reasoning, and software code generation and explanation. Composition training required only a small fraction—approximately 5% to 7%—of domain examples, avoiding the need for extensive retraining.
The findings show that model composition significantly enhances capabilities without degrading baseline performance. In synthetic tests requiring both key lookup and arithmetic, the combined system achieved an 84.3% success rate, whereas both standalone base models failed completely. For multilingual expansion across low-resource languages, composition improved translation to English across 175 of 192 languages and raised mathematical reasoning accuracy by up to 13% in absolute terms. In coding evaluations, the composed system delivered a relative improvement of about 40% over the base model on code generation benchmarks like HumanEval and MBPP, performing on par with fully fine-tuned alternatives.
These results demonstrate that organizations can systematically reuse existing specialized models rather than incurring the massive expense of training monolithic models from scratch. The method adds only about 1.5% in parameter overhead and operates at roughly 2% to 12% of the compute cost of full model retraining. Furthermore, it protects established capabilities, such as high-resource language understanding and natural language generation, which often degrade during standard fine-tuning.
Organizations seeking to extend language models to proprietary or highly specialized domains should pilot representation-level composition frameworks before investing in costly full-scale retraining. Teams should focus next on testing composition setups across multiple augmenting models simultaneously and measuring runtime inference latency in production environments.
Confidence in these findings is high for the tested domains of translation, arithmetic, and coding using standard transformer architectures. However, decision-makers should note that the approach assumes full access to the internal layer representations of both models, making it unsuitable for proprietary black-box systems accessed strictly through external text interfaces.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). Provides the foundational conceptual framing and benchmarks for diagnosing the compositionality gap when combining distinct model competencies.
- Paper: AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning, Yaqing Wang et al. (2022). Introduces modular adaptation layers that enable efficient domain-specific parameter extensions while keeping base model representations intact.
- Paper: Efficient Multimodal Fusion via Interactive Prompting, Yaowei Li et al. (2023). Demonstrates cross-network layer-level interactive prompting mechanisms for combining frozen representations across separate pre-trained systems.
- Paper: Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication, Zhangyue Yin et al. (2023). Explores the alternative baseline paradigm of multi-model collaboration and message exchange at the text-generation level.
- Paper: Mixture-of-Agents Enhances Large Language Model Capabilities, Junlin Wang et al. (2024). Extends multi-model capability aggregation from internal representation-level composition to layered architectural agent collaboration without internal cross-attention access.
- Paper: Localizing Task Information for Improved Model Merging and Compression, Ke Wang et al. (2024). Investigates weight-space merging and task localization techniques as a parameter-merging alternative to cross-attention representation composition.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). Generalizes the composition of modular mechanisms to tackle continuous multi-task adaptation over extended long-horizon sequences.
- Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). Explores automated natural-language coordination to dynamically orchestrate diverse specialized models without retraining or manual routing.
