Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use

Yuhan ChenAng LvTing-En LinChangyu ChenYuchuan WuFei HuangYongbin LiRui Yan

article2024ACL48 citations

Proposes Attention Buckets, a training-free inference method that eliminates blind spots in LLM context retrieval caused by rotary position embedding attention waveforms by ensembling parallel processes with complementary angle bases, boosting 7B models to GPT-4-level tool-use accuracy.

Listen

Large language models frequently experience position-dependent blind spots when processing long contexts, leading them to overlook crucial information such as application programming interface (API) documentation in tool-use tasks or retrieved evidence in question answering. This inconsistency undermines the reliability of autonomous agents and retrieval-augmented systems. The article demonstrates that these performance dips stem directly from an inherent waveform pattern in the model's rotary position embeddings, where tokens located at attention troughs receive significantly less focus than those at peaks.

The main objective of the article is to identify how attention waveforms affect context awareness and to evaluate Attention Buckets, a training-free inference method that stabilizes model focus across all context positions. To assess this, the authors tested open-source models using controlled synthetic retrieval tasks, the large-scale ToolBench benchmark covering over 16,000 real-world APIs, the ToolAlpaca simulation framework, and open-domain question-answering benchmarks including Natural Questions and WebQA.

The findings show that placing target data at an attention peak consistently yields higher retrieval accuracy than placing it at a trough across various context lengths. Applying the Attention Buckets method—which processes multiple parallel context copies across complementary rotary base angles and merges their weighted predictions—elevated an open-source 7-billion-parameter model to state-of-the-art results on ToolBench, achieving an average pass rate of 71.3% and a win rate of 71.5%, which matches or exceeds proprietary systems like GPT-4. Furthermore, the method improved overall accuracy across tool simulation benchmarks and general document question-answering tasks without requiring retraining.

These results demonstrate that significant performance gains can be achieved during inference alone by resolving positional bias, offering organizations a viable path to match proprietary model quality with cost-effective, smaller, open-source models. Organizations deploying tool-augmented agents or retrieval-based workflows should consider adopting attention-interleaving strategies during generation. However, decision-makers must weigh the trade-off between execution speed and hardware memory, as processing parallel context streams increases GPU memory consumption. Further work is recommended to validate the approach across non-rotary positional encodings and to optimize memory-efficient decoding techniques.

Cover for Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use

Abstract

In this paper, we demonstrate that an inherent waveform pattern in the attention allocation of large language models (LLMs) significantly affects their performance in tasks demanding a high degree of context awareness, such as utilizing LLMs for tool-use. Specifically, the crucial information in the context will be potentially overlooked by model when it is positioned in the trough zone of the attention waveform, leading to decreased performance. To address this issue, we propose a novel inference method named Attention Buckets. It allows LLMs to process their input through multiple parallel processes. Each process utilizes a distinct base angle for the rotary position embedding, thereby creating a unique attention waveform. By compensating an attention trough of a particular process with an attention peak of another process, our approach enhances LLM’s awareness to various contextual positions, thus mitigating their risk of overlooking crucial information. In the largest tool-use benchmark, our method elevates a 7B model to achieve state-of-the-art performance comparable to that of GPT-4. On other benchmarks and some RAG tasks, which also demand a thorough understanding of contextual content, Attention Buckets also exhibited notable enhancements in performance.

Table of Contents

  • 1 Introduction
  • 2 Attention Waves Impact on Context Awareness
  • 2.1 Preliminaries
  • 2.2 Hypothesis Verification
  • 2.3 Results and Analysis
  • 3 Enhancing Context Awareness via Interleaving Attention Waveform
  • 3.1 Preliminaries
  • 3.2 Method
  • 3.3 The Searching of ℬc\mathcal{B}_{c}
  • 4 Experiments
  • 4.1 Experiment Setups
  • 4.2 Results and Analysis
  • 4.3 Discussion on Efficacy
  • 5 Exploring Applications for Retrieval-Augmented Generation
  • 6 Ablation Study
  • 7 Related Work: LLM-Based Tool-Use
  • 8 Conclusion
  • References
  • A The Waveform of Attention Score Before Softmax
  • B Locating Peaks and Troughs in an Attention Waveform
  • C Supplement to The In-Context Retrieval Experiment
  • D The Impact of Base Value Smaller Than That Used In Training
  • E Results on ToolAlpaca

Knowls

  1. Knowl 1 — RoPE induces position-dependent attention waveforms

    model/method

    For a Transformer using rotary position embeddings (RoPE), let qm,kn∈Rdq_m,k_n\in\mathbb{R}^d be a query and key at token positions mm and nn, where dd is even. RoPE rotates each adjacent pair of vector coordinates by an angle determined by the relative token position. With rotary base B>0B>0, the frequency for coordinate pair ii is θi=B−2i/d\theta_i=B^{-2i/d}, and the pre-softmax attention score can be written as

    (Rθ,mqm)⊤(Rθ,nkn)=qm⊤Rθ,n−mkn.(R_{\theta,m}q_m)^\top(R_{\theta,n}k_n)=q_m^\top R_{\theta,n-m}k_n.

    The score is therefore a sum of sinusoidal terms whose phases depend on the relative token distance n−mn-m. This produces oscillating attention scores, alongside a broad tendency for their values to decay at larger relative distances. The paper uses an illustrative waveform proxy, not the actual score for every trained query-key pair: assuming paired query and key coordinates are all ones gives W(m−n)=∑i=0d/2−12cos⁡((m−n)θi)W(m-n)=\sum_{i=0}^{d/2-1}2\cos((m-n)\theta_i). The authors hypothesize that important context tokens at waveform troughs may receive less attention and be missed; this is a proposed explanation for position-sensitive performance, not a claim that the proxy exactly predicts every model's attention.

  2. Knowl 2 — Retrieval accuracy is higher when the target lies at an attention peak

    empirical result

    The authors tested the relationship between RoPE attention-waveform positions and context awareness using Llama-2-chat-7B. Each prompt contained either K=40K=40 or K=50K=50 synthetic JSON key-value pairs, with distinct UUID strings, and asked the model to return the value for a specified key. For each RoPE base, the target pair's final token was placed at the waveform peak nearest the context midpoint in one evaluation and at the nearest trough in another; accuracy was measured by exact value retrieval. The tested bases ranged from 10,000 to 30,000 in steps of 5,000.

    Could not parse LaTeX table

    Peak placement yielded higher accuracy than trough placement for every tested base and context length. Performance also varied substantially with base, and the best base differed by context length: 15,000 for K=40K=40 and 20,000 for K=50K=50. Thus, the experiment supports a position-related effect while also showing why a single base cannot be selected reliably when the location of crucial context information is unknown.

  3. Knowl 3 — Attention Buckets aggregates parallel RoPE-base predictions

    model/method

    Attention Buckets is an inference-time method for Transformer language models using RoPE; it requires no additional training. Given context CC, the method runs NN copies of the model in parallel, assigning copy jj a distinct RoPE base BjB_j from a selected set BcB_c. At autoregressive step kk, copy jj produces a next-token distribution pj(v)=p(Rk=v∣C,Bj,R1:k−1)p_j(v)=p(R_k=v\mid C,B_j,R_{1:k-1}) over vocabulary tokens vv, where R1:k−1R_{1:k-1} is the generated response prefix.

    The method uses the largest token probability in each copy's distribution as a confidence proxy, normalizes these values across copies, and takes their weighted mixture:

    αj′=max⁡vpj(v),αj=exp⁡(αj′)∑i=1Nexp⁡(αi′),p^(v)=∑j=1Nαjpj(v).\alpha'_j=\max_v p_j(v),\qquad \alpha_j=\frac{\exp(\alpha'_j)}{\sum_{i=1}^N\exp(\alpha'_i)},\qquad \hat p(v)=\sum_{j=1}^N\alpha_jp_j(v).

    The next token is decoded from p^\hat p, appended to the response, and the process repeats until the turn ends. The design aims for the different bases to produce complementary attention waveforms, so that a context position receiving low attention in one run can receive higher attention in another. The confidence weighting is the method's proposed way to favor runs that better use currently relevant context.

  4. Knowl 4 — Greedy search constructs a complementary set of RoPE bases

    algorithm

    The base-search procedure selects a set BcB_c whose attention-waveform peaks and troughs are distributed across context positions. It searches candidate bases from Bs={Bmin⁡+iS:i=0,…,(Bmax⁡−Bmin⁡)/S}B_s=\{B_{\min}+iS: i=0,\ldots,(B_{\max}-B_{\min})/S\}, with Bmin⁡B_{\min} set to the model's pretrained RoPE base to avoid using a smaller base that could introduce out-of-distribution positional frequencies. For the main experiments, Bmin⁡=Btrain=10,000B_{\min}=B_{\mathrm{train}}=10{,}000, Bmax⁡=30,000B_{\max}=30{,}000, stride S=500S=500, target set size N=6N=6, and maximum context length 4,096 tokens. The resulting set was {10,000,17,500,18,000,19,000,20,000,25,000}\{10{,}000,17{,}500,18{,}000,19{,}000,20{,}000,25{,}000\}.

    Input: Pretrained base B_train; minimum and maximum bases B_min and B_max; stride S; desired set size N; maximum context length L
    Output: Selected base set B_c
    Set candidate space B_s to bases from B_min to B_max in increments of S
    Initialize B_c to {B_train}
    Find peak positions P_c and trough positions T_c for B_train up to L
    While the size of B_c is less than N:
        For each candidate base B_j in B_s not already selected:
            Find its peak positions P_j and trough positions T_j up to L
            Compute its mismatch score as the aggregate absolute distances
                from candidate peaks to selected troughs, and from candidate
                troughs to selected peaks
        Select the candidate B_j with the smallest mismatch score
        Add B_j to B_c and add its peaks and troughs to P_c and T_c
    Return B_c

    The search favors candidates whose extrema complement those already selected; extrema are considered only within the maximum context length. In the paper's extrema-finding procedure, a local search window is expanded by a factor of 1.5 after each successive peak or trough. The paper does not state the initial window length.

  5. Knowl 5 — Attention Buckets reaches GPT-4-level average scores on ToolBench

    empirical result

    ToolBench evaluates tool use across six combinations of task level and scenario: I1-Inst., I1-Tool., I1-Cat., I2-Inst., I2-Cat., and I3-Inst. Each combination has 200 queries except I3-Inst., which has 100. The reported metrics are pass rate (the fraction of user queries fulfilled) and win rate (judged by ChatGPT against ChatGPT-ReACT). The benchmark contains 3,451 tools, 16,464 APIs, 126,486 instances, and 469,585 API calls. The authors applied Attention Buckets to the 7B ToolLlama model and used greedy decoding; the strongest configuration also used DFSDT-Retriever, an API-retrieval augmentation.

    The entries below are pass rate / win rate. Attention Buckets with DFSDT-Retriever was compared with the corresponding unaugmented ToolLlama configuration and with GPT-4 using DFSDT.

    Could not parse LaTeX table

    The augmented 7B model obtained the highest reported average pass rate and win rate in these comparisons, narrowly exceeding GPT-4's average pass rate and win rate in the specified configurations. Attention Buckets also improved both average metrics over ToolLlama with DFSDT-Retriever.

  6. Knowl 6 — ToolBench comparisons favor parallel base-specific inference

    empirical result

    Using the ToolLlama-DFSDT-Retriever configuration on ToolBench, the authors compared Attention Buckets with three alternatives. Attention Bucketsonce averages the waveforms from multiple RoPE bases before a single model run; Attention Sorting (ASort) rearranges context segments according to attention scores; Universal Self-Consistency (USC) generates multiple responses and uses the model to select a consistent response. The reported values are average pass rate / average win rate.

    Could not parse LaTeX table

    Attention Buckets performed best on both averages. The authors attribute Attention Bucketsonce's lower scores to the out-of-distribution positional information that can result from averaging waveforms before model computation; separate runs with distinct bases avoid this pre-averaging. ASort and USC delivered smaller improvements over the original configuration. The comparison indicates that the benefit is not explained simply by running multiple generations or by reordering context.

  7. Knowl 7 — Attention Buckets improves open-domain question answering

    empirical result

    The authors evaluated retrieval-augmented open-domain question answering with Llama-2-7B on Natural Questions (NQ; 3,610 test examples) and WebQA (2,032 examples). A supervised DPR retriever supplied 10 documents as context, and accuracy counted an answer as correct if it contained an accepted answer. The comparison included FiD-XL, a 3B model trained on ODQA benchmarks, and two inference-time alternatives.

    Could not parse LaTeX table

    Attention Buckets raised Llama-2-7B accuracy by 1.8 points on NQ and 1.4 points on WebQA relative to the unaugmented model, and it exceeded FiD-XL on both reported datasets.

  8. Knowl 8 — ToolAlpaca results support gains across model sizes

    empirical result

    On the simulated-tools subset of ToolAlpaca's evaluation, GPT-4 scored model outputs on Procedure (appropriate actions and parameters without redundant steps), Response (alignment with the user query), and Overall (the full action-response cycle). The authors restricted evaluation to simulated tools because the open-source evaluation code did not permit complete reproduction of the real-tools results. Scores are reported below in the order Procedure / Response / Overall.

    Could not parse LaTeX table

    Adding Attention Buckets improved all three metrics for both ToolAlpaca sizes. The 13B augmented model's Overall score of 74.0 was close to the reported GPT-3.5 score of 75.0.

  9. Knowl 9 — Searched base sets perform robustly on NQ

    empirical result

    The authors tested the base-search procedure on NQ with 10 retrieved documents, varying the number of selected bases NN and search stride SS. Accuracy is reported as a percentage. The four automatically searched sets used (N,S)(N,S) values of (7,100)(7,100), (7,1000)(7,1000), (6,500)(6,500), and (7,500)(7,500). They were compared with individual bases and two manually constructed arithmetic-sequence sets whose successive bases differ by 3,000 or 4,000.

    Could not parse LaTeX table

    All four searched sets scored above either arithmetic-sequence set, with similar accuracy across the different search settings. The strongest single base, 2.00×1042.00\times10^4, scored 50.28, close to the searched sets. Adding 2.25×1042.25\times10^4 to Bc3B_{c3} raised accuracy from 50.22 to 50.31 even though that base alone scored 49.58, consistent with the paper's claim that a base can contribute by covering positions that other bases handle less well.

  10. Knowl 10 — Scope and resource limitations

    limitation

    The reported method is developed for Transformer language models using RoPE; whether the same context-awareness approach works with other positional embeddings remains uninvestigated. Attention Buckets processes multiple copies of the context with different bases, increasing memory demands. The authors report that it did not reduce inference speed when sufficient memory was available and ran their experiments on a single NVIDIA A100-80G GPU, but they had not identified an efficient way to balance memory cost against inference speed.

Coverage note — Omitted per-task ToolBench scores for the ReAct and plain DFSDT configurations because the main benchmark comparison and the ToolBench alternative-method averages capture the central performance findings; no other substantial contributed material was omitted.

References

  1. 1.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. 2023. Qwen technical report.
  2. 2.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on freebase from question-answer pairs. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1533–1544.
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  4. 4.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading wikipedia to answer open-domain questions. arXiv preprint arXiv:1704.00051.
  5. 5.Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023a. Extending context window of large language models via positional interpolation.
  6. 6.Xinyun Chen, Renat Aksitov, Uri Alon, Jie Ren, Kefan Xiao, Pengcheng Yin, Sushant Prakash, Charles Sutton, Xuezhi Wang, and Denny Zhou. 2023b. Universal self-consistency for large language model generation. arXiv preprint arXiv:2311.17311.
  7. 7.Yifei Gao, Lei Wang, Jun Fang, Longhua Hu, and Jun Cheng. 2023. Empower your model with longer and better context comprehension.
  8. 8.Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. 2023. Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.
  9. 9.Jixiang Hong, Quan Tu, Changyu Chen, Xing Gao, Ji Zhang, and Rui Yan. 2023. Cyclealign: Iterative distillation from black-box llm to white-box models for better human alignment.
  10. 10.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering.
  11. 11.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick ˘ Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906.
  12. 12.Omar Khattab, Christopher Potts, and Matei Zaharia. 2021. Relevance-guided supervision for openqa with colbert. Transactions of the association for computational linguistics, 9:929–944.
  13. 13.Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, pages 39–48.
  14. 14.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  15. 15.Paul J. Leach, Rich Salz, and Michael H. Mealling. 2005. A Universally Unique IDentifier (UUID) URN Namespace. RFC 4122.
  16. 16.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  17. 17.Dacheng Li, Rulin Shao, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang. 2023. How long can open-source llms truly promise on context length?
  18. 18.Jia-Nan Li, Quan Tu, Cunli Mao, Zhengtao Yu, Ji-Rong Wen, and Rui Yan. 2024. Streamingdialogue: Prolonged dialogue learning via long context compression with minimal losses.
  19. 19.Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023a. Lost in the middle: How language models use long contexts.
  20. 20.Xiaoran Liu, Hang Yan, Shuo Zhang, Chenxin An, Xipeng Qiu, and Dahua Lin. 2023b. Scaling laws of rope-based extrapolation.
  21. 21.Yuhan Liu, Xiuying Chen, Xiaoqing Zhang, Xing Gao, Ji Zhang, and Rui Yan. 2024. From skepticism to acceptance: Simulating the attitude dynamics toward fake news.
  22. 22.Ang Lv, Yuhan Chen, Kaiyi Zhang, Yulong Wang, Lifeng Liu, Ji-Rong Wen, Jian Xie, and Rui Yan. 2024. Interpreting key mechanisms of factual recall in transformer-based language models.
  23. 23.Ang Lv, Kaiyi Zhang, Shufang Xie, Quan Tu, Yuhan Chen, Ji-Rong Wen, and Rui Yan. 2023. Are we falling in a middle-intelligence trap? an analysis and mitigation of the reversal curse.
  24. 24.Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom. 2023. Augmented language models: a survey.
  25. 25.OpenAI. 2022. OpenAI: Introducing ChatGPT.
  26. 26.OpenAI. 2023. Gpt-4 technical report.
  27. 27.Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. 2023. Gorilla: Large language model connected with massive apis.
  28. 28.Alexander Peysakhovich and Adam Lerer. 2023. Attention sorting combats recency bias in long context language models.
  29. 29.Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Yufei Huang, Chaojun Xiao, Chi Han, Yi Ren Fung, Yusheng Su, Huadong Wang, Cheng Qian, Runchu Tian, Kunlun Zhu, Shihao Liang, Xingyu Shen, Bokai Xu, Zhen Zhang, Yining Ye, Bowen Li, Ziwei Tang, Jing Yi, Yuzhang Zhu, Zhenning Dai, Lan Yan, Xin Cong, Yaxi Lu, Weilin Zhao, Yuxiang Huang, Junxi Yan, Xu Han, Xian Sun, Dahai Li, Jason Phang, Cheng Yang, Tongshuang Wu, Heng Ji, Zhiyuan Liu, and Maosong Sun. 2023a. Tool learning with foundation models.
  30. 30.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023b. Toolllm: Facilitating large language models to master 16000+ real-world apis.
  31. 31.Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2020. Rocketqa: An optimized training approach to dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2010.08191.
  32. 32.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
  33. 33.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools.
  34. 34.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.
  35. 35.Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning, june 2023. arXiv preprint arXiv:2303.11366.
  36. 36.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2022. Roformer: Enhanced transformer with rotary position embedding.
  37. 37.Dídac Surís, Sachit Menon, and Carl Vondrick. 2023. Vipergpt: Visual inference via python execution for reasoning.
  38. 38.Qiaoyu Tang, Ziliang Deng, Hongyu Lin, Xianpei Han, Qiao Liang, and Le Sun. 2023. Toolalpaca: Generalized tool learning for language models with 3000 simulated cases. arXiv preprint arXiv:2306.05301.
  39. 39.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models.
  40. 40.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models.
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  42. 42.Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023a. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In The Eleventh International Conference on Learning Representations.
  43. 43.Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun. 2023b. Label words are anchors: An information flow perspective for understanding in-context learning.
  44. 44.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023c. Self-consistency improves chain of thought reasoning in language models.
  45. 45.Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, and Hao Ma. 2023. Effective long-context scaling of foundation models.
  46. 46.Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong, Haizhou Zhao, Hang Xu, Haoze Sun, Hongda Zhang, Hui Liu, Jiaming Ji, Jian Xie, JunTao Dai, Kun Fang, Lei Su, Liang Song, Lifeng Liu, Liyun Ru, Luyao Ma, Mang Wang, Mickel Liu, MingAn Lin, Nuolan Nie, Peidong Guo, Ruiyang Sun, Tao Zhang, Tianpeng Li, Tianyu Li, Wei Cheng, Weipeng Chen, Xiangrong Zeng, Xiaochuan Wang, Xiaoxi Chen, Xin Men, Xin Yu, Xuehai Pan, Yanjun Shen, Yiding Wang, Yiyu Li, Youxin Jiang, Yuchen Gao, Yupeng Zhang, Zenan Zhou, and Zhiying Wu. 2023. Baichuan 2: Open large-scale language models.
  47. 47.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models.
  48. 48.Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2022. Generate rather than retrieve: Large language models are strong context generators. arXiv preprint arXiv:2209.10063.
  49. 49.Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2023. Webarena: A realistic web environment for building autonomous agents.
  50. 50.Dawei Zhu, Nan Yang, Liang Wang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. 2023. Pose: Efficient context window extension of llms via positional skip-wise training.

Citation

MLA
Chen, Y., et al. “Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 11160–74, https://doi.org/10.18653/v1/2024.acl-long.601.
APA
Chen, Y., Lv, A., Lin, T.-E., Chen, C., Wu, Y., Huang, F., Li, Y., & Yan, R. (2024). Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11160–11174. https://doi.org/10.18653/v1/2024.acl-long.601
Chicago
Chen, Y., A. Lv, T.-E. Lin, et al. 2024. “Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11160–74. https://doi.org/10.18653/v1/2024.acl-long.601.
Harvard
Chen, Y. et al. (2024) “Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11160–11174. Available at: https://doi.org/10.18653/v1/2024.acl-long.601.
Vancouver
1. Chen Y, Lv A, Lin T-E, Chen C, Wu Y, Huang F, Li Y, Yan R (2024) Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 11160–11174

BibTeX

@inproceedings{chen-etal-2024-fortify,
    title = "Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use",
    author = "Chen, Yuhan  and
      Lv, Ang  and
      Lin, Ting-En  and
      Chen, Changyu  and
      Wu, Yuchuan  and
      Huang, Fei  and
      Li, Yongbin  and
      Yan, Rui",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.601/",
    doi = "10.18653/v1/2024.acl-long.601",
    pages = "11160--11174"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/