Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use
Yuhan ChenAng LvTing-En LinChangyu ChenYuchuan WuFei HuangYongbin LiRui Yan
Proposes Attention Buckets, a training-free inference method that eliminates blind spots in LLM context retrieval caused by rotary position embedding attention waveforms by ensembling parallel processes with complementary angle bases, boosting 7B models to GPT-4-level tool-use accuracy.
Large language models frequently experience position-dependent blind spots when processing long contexts, leading them to overlook crucial information such as application programming interface (API) documentation in tool-use tasks or retrieved evidence in question answering. This inconsistency undermines the reliability of autonomous agents and retrieval-augmented systems. The article demonstrates that these performance dips stem directly from an inherent waveform pattern in the model's rotary position embeddings, where tokens located at attention troughs receive significantly less focus than those at peaks.
The main objective of the article is to identify how attention waveforms affect context awareness and to evaluate Attention Buckets, a training-free inference method that stabilizes model focus across all context positions. To assess this, the authors tested open-source models using controlled synthetic retrieval tasks, the large-scale ToolBench benchmark covering over 16,000 real-world APIs, the ToolAlpaca simulation framework, and open-domain question-answering benchmarks including Natural Questions and WebQA.
The findings show that placing target data at an attention peak consistently yields higher retrieval accuracy than placing it at a trough across various context lengths. Applying the Attention Buckets method—which processes multiple parallel context copies across complementary rotary base angles and merges their weighted predictions—elevated an open-source 7-billion-parameter model to state-of-the-art results on ToolBench, achieving an average pass rate of 71.3% and a win rate of 71.5%, which matches or exceeds proprietary systems like GPT-4. Furthermore, the method improved overall accuracy across tool simulation benchmarks and general document question-answering tasks without requiring retraining.
These results demonstrate that significant performance gains can be achieved during inference alone by resolving positional bias, offering organizations a viable path to match proprietary model quality with cost-effective, smaller, open-source models. Organizations deploying tool-augmented agents or retrieval-based workflows should consider adopting attention-interleaving strategies during generation. However, decision-makers must weigh the trade-off between execution speed and hardware memory, as processing parallel context streams increases GPU memory consumption. Further work is recommended to validate the approach across non-rotary positional encodings and to optimize memory-efficient decoding techniques.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). Its controlled evidence that relevant information is often missed in the middle of long contexts sets up this paper’s deeper account of position-dependent retrieval failures.
- Paper: Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation, Zichong Li et al. (2026). It follows the diagnosis of RoPE-linked positional brittleness with a training-based method for making long-context retrieval more position-robust.
- Paper: ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning, Yanjun Zhao et al. (2026). It carries the context-awareness problem into long-context reasoning with an inference-time evidence-replay method that helps models use information they otherwise overlook.
- Paper: Context-Aware RL for Agentic and Multimodal LLMs, Peiyang Xu Bangzheng Li Sijia Liu Xingyu Fu (2026). It extends context-awareness work to coding agents and multimodal models, using reinforcement learning to teach models to identify decisive supporting evidence.
