Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search

Dongge HanMenglin XiaDaniel MadrigalSamuel KesslerAnkur MallickXuchao ZhangMirian Hipolito GarciaJin XuVictor RuehleSaravan Rajmohan

article2025arXiv2 citations

Presents a training-free framework that improves the reasoning accuracy and stability of small language models on math, coding, and logic benchmarks by combining structured blueprints with automated prompt template search.

Listen

Deploying large artificial intelligence models requires massive computational power and high operational costs. While smaller language models offer a lightweight, cost-effective, and privacy-friendly alternative for edge and on-device use, their smaller capacity severely limits their ability to solve complex, multi-step problems. Furthermore, smaller models are highly fragile when prompt layouts change, meaning small structural adjustments in input text can significantly degrade their output accuracy.

The article introduces and evaluates a framework designed to improve the reasoning accuracy and consistency of small models without expanding their size or performing expensive retraining. The objective is to demonstrate that structured, reusable reasoning guides—termed blueprints—combined with systematic prompt template optimization can enable small models to perform complex tasks effectively.

To achieve this, the authors used a larger frontier model to extract reusable, step-by-step problem-solving instructions across twelve distinct stylistic formats. These blueprints were evaluated and iteratively refined for specific small models using automated error analysis. Additionally, the approach employed an efficient search method to identify the optimal structural arrangement of prompt components, including the placement of instructions, reasoning steps, and example problems. The evaluation tested three distinct small language models (Phi3-mini, Mistral-7B, and GPT-4o-mini) across 28 diverse task categories spanning mathematical reasoning, Python code generation, and complex logic puzzles, totaling around 5,600 evaluation points.

The findings show substantial performance improvements across all tested domains. Introducing blueprints alone consistently outperformed traditional baseline techniques, such as standard chain-of-thought and conventional prompt optimization. For instance, blueprint guidance improved Mistral-7B's code generation accuracy by roughly 20 percentage points over standard three-shot baseline prompting. Combining error-refined blueprints with optimal template search delivered the strongest overall results, achieving top performance in five out of nine model-task configurations and near-optimal results across the remainder. Furthermore, the experiments revealed that model architectures exhibit distinct preferences for blueprint styles; smaller models showed up to an 11% to 12% performance variance depending on whether instructions were framed as concrete steps, decision criteria, or bullet points.

These results demonstrate that a significant portion of the performance deficit in small models stems from input framing rather than absolute capacity limits. By providing explicit procedural roadmaps and tailoring prompt structures to specific models, organizations can achieve high-grade reasoning on resource-constrained hardware. This drastically reduces computing costs, lowers latency, and mitigates the security and compliance risks associated with transmitting proprietary data to third-party cloud models.

Organizations seeking to deploy language models on local hardware or at lower operational costs should adopt structured blueprint guidance rather than relying on unstructured prompting. Engineering teams should also avoid universal prompt templates, instead tailoring prompt layouts to the specific target model. Before making large-scale deployment decisions, teams should conduct focused pilots to establish task-specific blueprints and confirm template configurations.

The study's main limitation lies in its validation scope, which relies on synthetic and benchmark datasets with structured outputs rather than open-ended operational workflows. Additionally, because the template selection used small training samples, the search occasionally settled on near-optimal rather than absolute best configurations. Confidence in the underlying conclusion is high for structured reasoning tasks, though readers should test and validate blueprints within their own domain-specific environments.

arXiv: 2506.08669
Cover for Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search

Abstract

Small language models (SLMs) offer promising and efficient alternatives to large language models (LLMs). However, SLMs' limited capacity restricts their reasoning capabilities and makes them sensitive to prompt variations. To address these challenges, we propose a novel framework that enhances SLM reasoning capabilities through LLM generated blueprints. The blueprints provide structured, high-level reasoning guides that help SLMs systematically tackle related problems. Furthermore, our framework integrates a prompt template search mechanism to mitigate the SLMs' sensitivity to prompt variations. Our framework demonstrates improved SLM performance across various tasks, including math (GSM8K), coding (MBPP), and logic reasoning (BBH). Our approach improves the reasoning capabilities of SLMs without increasing model size or requiring additional training, offering a lightweight and deployment-friendly solution for on-device or resource-constrained environments.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Problem Formulation
  • 3.2 Blueprint Generation and Optimization
  • 3.3 Prompt Template Search
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Main Results
  • 4.3 Example Blueprint and SLM Response following Blueprint Steps
  • 4.4 SLM performance with different blueprint styles
  • 5 Conclusion
  • References
  • A Appendix: Blueprint Generation and Optimization
  • A.1 Blueprint Generation Styles
  • A.2 More details on Blueprint APO Refinement
  • B More Details on Template Search
  • C Detailed Per-category Results on BBH
  • D More Probing Experiment Results on Prompt Sensitivity
  • E More Details on Experiment Setup

Knowls

  1. Knowl 1 — Blueprint-Guided Reasoning Framework for Small Language Models

    model/method

    The blueprint-guided reasoning framework enhances the problem-solving capabilities of small language models (SLMs) without modifying model parameters or requiring fine-tuning. A blueprint is defined as a reusable, high-level, step-by-step procedural instruction guide tailored to solve a broad class of related tasks T\mathcal{T} (e.g., arithmetic reasoning, code generation, or symbolic logic).

    Unlike standard Chain-of-Thought (CoT) prompting or In-Context Learning (ICL)—where SLMs often fail to induce or transfer abstract reasoning strategies from few-shot demonstrations—the framework uses a larger teacher model (MllmM_{\text{llm}}, such as GPT-4o) to synthesize an explicit, abstract reasoning plan from exemplar problems.

    The framework consists of three offline optimization stages for each task category T\mathcal{T} and target model MslmM_{\text{slm}}:

    1. Blueprint Generation and Style Selection: MllmM_{\text{llm}} generates candidate blueprints spanning diverse structural and stylistic formats, from which the highest-performing style on validation instances is selected.
    2. Automatic Prompt Optimization (APO) Refinement: The selected blueprint is iteratively updated by generating textual error gradients based on MslmM_{\text{slm}} failure cases.
    3. Prompt Template Search: A systematic search over prompt structural configurations (component ordering, demonstration counts, CoT trigger inclusion) selects the optimal template.

    During inference, the optimized blueprint and template are fixed and reused across all incoming queries belonging to task category T\mathcal{T}.

  2. Knowl 2 — Blueprint Refinement via Automatic Prompt Optimization

    algorithm

    To adapt an initial blueprint BB to the idiosyncratic error patterns of a target small language model MslmM_{\text{slm}} on a task category T\mathcal{T}, the framework applies Automatic Prompt Optimization (APO) using a teacher language model MllmM_{\text{llm}}.

    Input: Target model MslmM_{\text{slm}}, teacher model MllmM_{\text{llm}}, task category T\mathcal{T}, initial blueprint BB, training dataset DtrainTD_{\text{train}}^{\mathcal{T}}, number of gradients Ngrad=2N_{\text{grad}} = 2, evaluation samples neval=25n_{\text{eval}} = 25, selection samples nsel=20n_{\text{sel}} = 20
    Output: Optimized blueprint B∗B^*
    1. Sample nevaln_{\text{eval}} training examples from DtrainTD_{\text{train}}^{\mathcal{T}}
    2. Evaluate MslmM_{\text{slm}} on the sample using blueprint BB
    3. Identify incorrect predictions and sample up to 5 failures to compile an error message ee
    4. Prompt MllmM_{\text{llm}} with task description, blueprint BB, and error message ee to generate NgradN_{\text{grad}} textual gradients {g1,g2}\{g_1, g_2\}, where each gjg_j is a causal error analysis
    5. For each textual gradient gjg_j (j∈{1,2}j \in \{1, 2\}):
         Prompt MllmM_{\text{llm}} with BB, ee, and gjg_j to synthesize edited blueprint BjeditB_j^{\text{edit}}
         Prompt MllmM_{\text{llm}} to generate a simplified paraphrased blueprint BjparaB_j^{\text{para}} from BjeditB_j^{\text{edit}}
    6. Construct candidate pool C={B,B1edit,B2edit,B1para,B2para}\mathcal{C} = \{B, B_1^{\text{edit}}, B_2^{\text{edit}}, B_1^{\text{para}}, B_2^{\text{para}}\}
    7. Sample nseln_{\text{sel}} validation examples from DtrainTD_{\text{train}}^{\mathcal{T}}
    8. For each candidate c∈Cc \in \mathcal{C}:
         Evaluate accuracy of MslmM_{\text{slm}} on the nseln_{\text{sel}} examples using candidate cc
    9. Return B∗=arg⁡max⁡c∈CAccuracy(c)B^* = \arg\max_{c \in \mathcal{C}} \text{Accuracy}(c)

    Executing one round of APO incurs 6 teacher LLM queries (2 gradient generations, 2 blueprint edits, 2 paraphrases) and 125 SLM queries (2525 initial +5×20+ 5 \times 20 candidate selection evaluations).

  3. Knowl 3 — Successive Halving Prompt Template Search

    algorithm

    Because small language models exhibit severe sensitivity to prompt formatting, component ordering, and demonstration counts, an optimal prompt template is selected from a discrete 32-configuration search space using a successive halving search algorithm.

    The search space Ω\Omega is defined by 4 categorical parameters:

    • Number of in-context demonstration examples: kshot∈{0,1,2,3}k_{\text{shot}} \in \{0, 1, 2, 3\}
    • Ordering of task description relative to demonstrations: TaskFirst∈{True,False}\text{TaskFirst} \in \{\text{True}, \text{False}\}
    • Inclusion of blueprint: UseBP∈{True,False}\text{UseBP} \in \{\text{True}, \text{False}\}
    • Inclusion of Chain-of-Thought trigger sentence ("Let's think step by step"): UseCoT∈{True,False}\text{UseCoT} \in \{\text{True}, \text{False}\}
    Input: Initial candidate template set Ω\Omega (∣Ω∣=32|\Omega| = 32), reduction factor f=2f = 2, sample size per round k=5k = 5, training set DtrainTD_{\text{train}}^{\mathcal{T}}, target model MslmM_{\text{slm}}
    Output: Best prompt template τ∗\tau^*
    1. Set S←Ω\mathcal{S} \leftarrow \Omega
    2. while ∣S∣>1|\mathcal{S}| > 1 do:
         Sample kk training examples Dbatch⊂DtrainTD_{\text{batch}} \subset D_{\text{train}}^{\mathcal{T}}
         For each template τ∈S\tau \in \mathcal{S}:
           Format DbatchD_{\text{batch}} using template τ\tau
           Evaluate accuracy of MslmM_{\text{slm}} on DbatchD_{\text{batch}}
         Rank templates in S\mathcal{S} in descending order of performance
         Keep the top ⌊∣S∣/f⌋\lfloor |\mathcal{S}| / f \rfloor candidates: S←TopCandidates(S,⌊∣S∣/f⌋)\mathcal{S} \leftarrow \text{TopCandidates}(\mathcal{S}, \lfloor |\mathcal{S}| / f \rfloor)
    3. Return the single remaining template τ∗∈S\tau^* \in \mathcal{S}

    Running successive halving across the 32 templates evaluates candidate sets of size 32, 16, 8, 4, 2, and 1 across rounds, consuming a total of (32+16+8+4+2)×5=310(32 + 16 + 8 + 4 + 2) \times 5 = 310 SLM queries per task category.

  4. Knowl 4 — Blueprint Generation Styles and Categorization

    model/method

    To generate candidate blueprints, a teacher LLM (MllmM_{\text{llm}}) is prompted with MM exemplar problem-solution pairs alongside one of 12 distinct style instructions S={S1,…,S12}S = \{S_1, \dots, S_{12}\}:

    1. concrete_example: Details each reasoning step accompanied by one or more specific example problems.
    2. abstract_example: Uses a synthesized, domain-abstracted example to illustrate reasoning steps.
    3. detailed_pattern: Outlines a detailed, comprehensive step sequence with explicit rationales.
    4. plain_pattern: Provides a concise general reasoning pattern within 300 words without Markdown formatting.
    5. concise_highlevel: Focuses on high-level key phases without detailed solution steps.
    6. instruction_focus: Emphasizes unambiguous, discrete instructional directives.
    7. contextual_explanation: Provides problem context clarification and background principles.
    8. reflective_refinement: Directs the model to attempt an initial solution, critically evaluate it for errors, and refine it.
    9. workflow: Outlines an end-to-end logically progressing, sequential pipeline.
    10. bullet_points: Lists key considerations and steps in bulleted format.
    11. decision_making: Structures the task into subcomponents and decision criteria for each branch.
    12. plan_and_solve: Instructs the model to define problem goals and explicit decision criteria prior to executing step-by-step resolution.

    During style selection, each of the 12 generated candidate blueprints is evaluated on 10 training instances using MslmM_{\text{slm}} (120 SLM queries total), and the highest-scoring style is selected as the baseline for further optimization.

  5. Knowl 5 — Benchmark Accuracy of Blueprint Prompting and Template Search Across SLMs

    data/table

    The effectiveness of blueprint-based reasoning and prompt template search was evaluated across three SLMs (GPT-4o-mini, Mistral-7B, Phi-3-mini) on mathematical reasoning (GSM8K), code synthesis (MBPP), and logical reasoning (BBH, averaged across subcategories). Baseline comparisons include 1-shot CoT, 3-shot CoT, and standard APO (which optimizes task descriptions only). Our methods include raw blueprint style selection without APO (BP w.o. APO), blueprint with APO refinement (BP w. APO), and the full pipeline with template search (BP w. APO + TS).

    BBH GSM8K MBPP
    Method GPT-4o-m Mistral-7B Phi-3-m GPT-4o-m Mistral-7B Phi-3-m GPT-4o-m Mistral-7B Phi-3-m
    CoT (1-shot) 0.839 0.449 0.678 0.940 0.400 0.807 0.813 0.202 0.642
    CoT (3-shot) 0.850 0.530 0.686 0.923 0.347 0.780 0.821 0.233 0.658
    APO 0.825 0.435 0.651 0.930 0.407 0.797 0.837 0.288 0.658
    BP (w.o. APO) 0.882 0.539 0.743 0.947 0.450 0.817 0.821 0.440 0.689
    BP (w. APO) 0.884 0.559 0.738 0.943 0.480 0.840 0.829 0.401 0.681
    BP (w. APO) + TS 0.884 0.572 0.722 0.953 0.490 0.820 0.833 0.424 0.696

    Blueprint-guided methods consistently outperform CoT and APO baselines across all tasks and models. Notably, BP (w.o. APO) improves Mistral-7B on MBPP by +20.7%+20.7\% over 3-shot CoT (from 0.2330.233 to 0.4400.440) and Phi-3-mini on BBH by +5.7%+5.7\% (from 0.6860.686 to 0.7430.743). The complete pipeline (BP w. APO + TS) achieves the best overall performance in 5 out of 9 benchmark combinations and near-optimal accuracy in the rest.

  6. Knowl 6 — Small Language Model Sensitivity to Prompt Ordering and Few-Shot Counts

    empirical result

    Small language models display marked performance volatility under minor prompt restructuring, demonstrating that optimal prompt configurations are non-universal across model families and tasks:

    1. Component Ordering Sensitivity: Swapping the order of <task-description> and <in-context-example> causes substantial accuracy shifts. For example, on the BBH Sports Understanding subtask, GPT-4o-mini performs significantly better when few-shot examples precede the task description, whereas Mistral-7B and Phi-3-mini achieve substantially higher accuracy when the task description precedes the examples.
    2. Non-monotonic Few-Shot Scaling: Increasing in-context examples from 00 to 33 shots does not universally improve SLM reasoning. On BBH Temporal Sequences, Phi-3-mini accuracy increases from 0-shot to 1-shot but monotonically degrades from 1-shot to 3-shot, whereas Mistral-7B exhibits the inverse behavior (dropping from 0-shot to 1-shot before improving at 3-shot).
  7. Knowl 7 — Architecture-Specific Preferences for Blueprint Prompt Styles

    empirical result

    Evaluating 12 blueprint generation styles across 28 task categories (GSM8K, MBPP, and 26 BBH subcategories; 280 test instances per style) reveals that model families exhibit diverging stylistic preferences and varying degrees of style sensitivity:

    • GPT-4o-mini displays high robustness to stylistic framing, with only a 3%3\% spread between its top-performing styles (bullet_points and workflow at 0.880.88 average accuracy) and lowest-performing styles (contextual_explanation, plain_pattern, and concise_highlevel at 0.850.85).
    • Phi-3-mini exhibits moderate sensitivity (11%11\% performance spread), strongly preferring unambiguous, direct styles (instruction_focus at 0.750.75, plain_pattern at 0.740.74) while failing on reflective multi-step prompts (plan_and_solve at 0.640.64, bullet_points at 0.660.66).
    • Mistral-7B exhibits the highest sensitivity (12%12\% performance spread), favoring structured decision guidelines (decision_making and instruction_focus at 0.500.50) while performing worst with bullet_points (0.380.38).

    A prominent finding is that the bullet_points format is optimal for GPT-4o-mini (0.880.88) but constitutes the single worst style for Mistral-7B (0.380.38) and ranks near the bottom for Phi-3-mini (0.660.66).

Coverage note — Individual per-category subtask bar plots for all 26 BBH categories (Appendix Figure 9) were omitted in favor of the aggregate benchmark table and representative qualitative sensitivity findings.

References

  1. 1.Abdin, M., Jacobs, S. A., Awan, A. A., Aneja, J., Awadallah, A., Awadalla, H., Bach, N., Bahree, A., Bakhtiari, A., Behl, H., et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024.
  2. 2.Agarwal, E., Dani, V., Ganu, T., and Nambi, A. Promptwizard: Task-aware agent-driven prompt optimization framework. arXiv preprint arXiv:2405.18369, 2024.
  3. 3.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  4. 4.Brown, T. B. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  5. 5.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
  6. 6.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  7. 7.Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E. P., and Hu, Z. Rlprompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548, 2022.
  8. 8.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al. Glam: Efficient scaling of language models with mixture-of-experts. In International Conference on Machine Learning, pp. 5547–5569. PMLR, 2022.
  9. 9.Fernando, C., Banarse, D., Michalewski, H., Osindero, S., and Rocktäschel, T. Promptbreeder: Self-referential self-improvement via prompt evolution. arXiv preprint arXiv:2309.16797, 2023.
  10. 10.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  11. 11.Langley, P. Crafting papers on machine learning. In Langley, P. (ed.), Proceedings of the 17th International Conference on Machine Learning (ICML 2000), pp. 1207–1216, Stanford, CA, 2000. Morgan Kaufmann.
  12. 12.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021.
  13. 13.Long, J. Large language model guided tree-of-thought. arXiv preprint arXiv:2305.08291, 2023.
  14. 14.Ma, R., Wang, X., Zhou, X., Li, J., Du, N., Gui, T., Zhang, Q., and Huang, X. Are large language models good prompt optimizers? arXiv preprint arXiv:2402.02101, 2024.
  15. 15.Magister, L. C., Mallinson, J., Adamek, J., Malmi, E., and Severyn, A. Teaching small language models to reason. arXiv preprint arXiv:2212.08410, 2022.
  16. 16.OpenAI. Hello gpt-4o. https://openai.com/index/hello-gpt-4o/, 2024a.
  17. 17.OpenAI. Gpt-4o mini: Advancing cost-efficient intelligence. https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/, 2024b.
  18. 18.Pryzant, R., Iter, D., Li, J., Lee, Y. T., Zhu, C., and Zeng, M. Automatic prompt optimization with" gradient descent" and beam search. arXiv preprint arXiv:2305.03495, 2023.
  19. 19.Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  20. 20.Wan, Z., Wang, X., Liu, C., Alam, S., Zheng, Y., Qu, Z., Yan, S., Zhu, Y., Zhang, Q., Chowdhury, M., et al. Efficient large language models: A survey. arXiv preprint arXiv:2312.03863, 1, 2023.
  21. 21.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  22. 22.Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., et al. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846, 2023.
  23. 23.Yang, S., Zhao, H., Zhu, S., Zhou, G., Xu, H., Jia, Y., and Zan, H. Zhongjing: Enhancing the chinese medical capabilities of large language model through expert feedback and real-world multi-turn dialogue. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 19368–19376, 2024.
  24. 24.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022.
  25. 25.Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. Large language models are humanlevel prompt engineers. arXiv preprint arXiv:2211.01910, 2022.

Citation

MLA
Han, D., et al. “Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search”. arXiv, 2025, http://arxiv.org/abs/2506.08669v1.
APA
Han, D., Xia, M., Diaz, D. M., Kessler, S., Mallick, A., Zhang, X., Garcia, M. D. C. H., Xu, J., Rühle, V., & Rajmohan, S. (2025). Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search. arXiv. http://arxiv.org/abs/2506.08669v1
Chicago
Han, D., M. Xia, D. M. Diaz, et al. 2025. “Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search”. arXiv. http://arxiv.org/abs/2506.08669v1.
Harvard
Han, D. et al. (2025) “Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2506.08669v1.
Vancouver
1. Han D, Xia M, Diaz DM, Kessler S, Mallick A, Zhang X, Garcia MDCH, Xu J, Rühle V, Rajmohan S (2025) Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search. arXiv

BibTeX

@article{han2025enhancing,
  title = {Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search},
  author = {Han, Dongge and Xia, Menglin and Diaz, Daniel Madrigal and Kessler, Samuel and Mallick, Ankur and Zhang, Xuchao and Garcia, Mirian Del Carmen Hipolito and Xu, Jin and Rühle, Victor and Rajmohan, Saravan},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2506.08669v1},
  eprint = {2506.08669}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/