Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models

Mukul SinghAnanya SinghaAishni ParabPronita MehrotraSumit Gulwani

article2025arXiv1 citations

Demonstrates that training language models with reinforcement learning guided by cognitive associative thinking metrics improves their capacity for novel connections, abstraction, and problem-solving across creative writing, programming, and data visualization.

Listen

Large language models often struggle with creative tasks that require original synthesis and connecting conceptually distant ideas. Because standard training optimizes for predicting the most probable next word, these systems frequently default to familiar patterns and high semantic similarity rather than producing novel insights. This limitation presents a challenge for deploying artificial intelligence in tasks where creativity, abstraction, and flexible problem-solving are essential.

The article demonstrates that training language models using reinforcement learning grounded in cognitive principles of associative thinking enhances their performance across both creative and analytical domains. The authors evaluate whether rewarding models for generating conceptually distant and diverse associations improves outputs in creative writing, data visualization, and software coding.

To achieve this, the authors designed an automated evaluation mechanism based on four established divergent thinking metrics: novelty, fluency, flexibility, and elaboration. These metrics measure the rarity of concept combinations, the volume of distinct ideas, the diversity of semantic categories, and the level of explanatory detail. Using policy-gradient reinforcement learning algorithms, the authors fine-tuned several small and large language models on an 8-GPU computing cluster, with training runs averaging about two and a half hours per model. The automated reward function was validated against human evaluators across standard and custom benchmarks covering 5,000 storytelling tasks, over 18,000 code generation problems, and 500 visualization tasks.

The findings show that reinforcement learning focused on associative thinking produces overall performance gains of 8% to 13% across models and tasks. The highest improvements occurred in storytelling, where models saw gains ranging between 9.8% and 13.4%. Data visualization performance also improved consistently by 8.5% to 10.2%. For code generation, performance changes were more moderate, with improvements of 7.2% to 8.3% on larger architectures, though one instruction-tuned model experienced a slight regression of 1.2% to 1.5%. Additionally, the automated creativity reward demonstrated high stability and a strong positive correlation with human creativity scores (a correlation coefficient of 0.78).

These results indicate that embedding associative thinking principles into language model training improves an artificial intelligence system's capacity for abstraction and creative synthesis without requiring large-scale human annotation. Furthermore, the benefits extend beyond traditionally creative fields into analytical workflows such as chart design. However, the slight performance dips seen in certain coding tests indicate a potential trade-off between open-ended associative exploration and the strict syntactic precision required in programming.

Organizations developing or deploying language models for complex problem-solving should consider integrating associative reward mechanisms into their fine-tuning pipelines. Because the approach occasionally impacts precision in highly structured tasks, practitioners should carefully balance creativity rewards with correctness constraints when deploying models for mission-critical technical functions.

Key limitations include the reliance on language-model-based evaluators to score creativity, which carries a risk of evaluation bias and alignment drift. Additionally, the evaluation was restricted to three specific task domains and exhibited occasional degradations in factual grounding and fluency. While the reported gains are strong across the tested benchmarks, broader validation across additional creative and analytical fields is necessary before widespread deployment in safety-critical settings.

arXiv: 2511.17876
Cover for Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models

Abstract

Associative thinking--the ability to connect seemingly unrelated ideas--is a foundational element of human creativity and problem-solving. This paper explores whether reinforcement learning (RL) guided by associative thinking principles can enhance a model's performance across diverse generative tasks, including story writing, code generation, and chart creation. We introduce a reinforcement learning framework that uses a prompt-based evaluation mechanism, incorporating established divergent thinking metrics from creativity research. A base language model is fine-tuned using this framework to reward outputs demonstrating higher novelty through higher degrees of conceptual connectivity. Interestingly, the experimental results suggest that RL-based associative thinking-trained models not only generate more original and coherent stories but also exhibit improved abstraction and flexibility in tasks such as programming and data visualization. Our findings provide initial evidence that modeling cognitive creativity principles through reinforcement learning can yield more adaptive and generative AI.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Associative Thinking and RL
  • 3.1 Measuring Creativity
  • 3.2 Reinforcement Learning Training
  • 3.3 Reward Function
  • 4 Experiment Setup
  • 4.1 Training Harness
  • 4.2 Benchmarks
  • 4.3 Metrics
  • 4.4 Models and Configuration
  • 5 Results
  • 5.1 RQ1: Creativity Improvements with RL
  • 5.2 RQ2: Creativity Reward Model
  • 5.3 RQ3: Transfer to Non-Creative Domains
  • 6 Conclusion
  • 7 Limitations
  • References

Knowls

  1. Knowl 1 — Four-Dimensional Associative Creativity Metric for Text Generation

    model/method

    To evaluate and reinforce associative thinking in language models, generated text outputs are modeled as collections of conceptual associations or "mini-ideas" (such as novel metaphors or juxtapositions of concepts). The overall creativity of a response is quantified along four dimensions adapted from Guilford's divergent thinking framework:

    • Novelty: Evaluates how unusual, rare, or unexpected the conceptual connections are compared to typical corpus patterns, rewarding the pairing of concepts rarely found together in training distributions.
    • Fluency: Quantifies the total count of distinct conceptual associations generated in response to a given prompt.
    • Flexibility: Measures the diversity of distinct semantic categories or conceptual domains spanned by the generated associations, penalizing outputs confined to narrow semantic topics.
    • Elaboration: Evaluates the granularity, depth, and explanatory detail provided to substantiate and develop the identified connections.
  2. Knowl 2 — Automated Checklist-Guided Evaluator for Associative Rewards

    model/method

    The associative creativity reward function is implemented using an automated, checklist-guided evaluator powered by a large language model (LLM):

    1. Novelty Scoring: The evaluator inspects candidate outputs to detect and score unusual or rare concept combinations.
    2. Fluency Scoring: The evaluator enumerates and counts distinct associative connections made in the text.
    3. Flexibility Scoring: The evaluator assigns candidate associations into semantic clusters and computes their categorical breadth.
    4. Elaboration Scoring: The evaluator queries the text for the depth and thoroughness of explanations expanding on each link.

    Component scores are aggregated into a unified scalar reward signal. To ensure reliability and mitigate evaluator bias, the evaluation checklists are calibrated against human annotations (evaluated on a 50-sample set adjusted for annotator bias) to produce a scalable and stable training signal without requiring continuous human grading.

  3. Knowl 3 — Policy-Gradient Reinforcement Learning Framework for Associative Thinking

    model/method

    Base pretrained language models are fine-tuned to acquire associative thinking capabilities using policy-gradient reinforcement learning algorithms, specifically Proximal Policy Optimization (PPO), Group Relative Policy Optimization (GRPO), and REINFORCE.

    The language model serves as the policy parameterizing token generation over prompt rollouts. Training proceeds for 100 iterations with 32 rollouts per iteration on a cluster of 8 NVIDIA H100 GPUs, requiring approximately 2 hours and 35 minutes of compute per model. Model parameters are updated to maximize the scalar creativity reward aggregating novelty, fluency, flexibility, and elaboration. Reward convergence and validation-set creativity metrics are monitored, with early stopping applied to prevent overfitting, policy collapse, and fluency degradation.

  4. Knowl 4 — Evaluation Benchmark Suite for Associative Thinking

    experimental setup

    To evaluate both creative and analytical transfer of associative thinking, models are tested across three task domains:

    1. Storytelling (LitBench): Comprises 5,000 tasks requiring models to generate coherent narratives from given premises, evaluated using LitBench's automated LLM-based rubric.
    2. Code Generation (MultiPL-E): Comprises over 18,000 programming tasks spanning more than 30 high- and low-resource programming languages, evaluated via functional unit test execution pass rates on official test splits.
    3. Data Visualization (ChartEval / Custom): Comprises 500 tasks requiring combined analytical data operations (pivoting, filtering, transformations) and visual design decisions (plot selection, color schemes, visual layout), evaluated using automated chart quality metrics.
  5. Knowl 5 — Downstream Performance Improvements from Associative RL Training

    data/table

    Fine-tuning language models with reinforcement learning guided by the associative thinking reward yields improvements across storytelling, code generation, and data visualization compared to base un-tuned models.

    Model Storytelling (%) Code Generation (%) Data Visualization (%)
    Deepseek-distill-7B 12.1 8.3 9.5
    Phi-4-13B 13.4 7.2 10.2
    Phi-3.5-instruct 9.8 -1.5 8.7

    Storytelling exhibits the largest relative improvements (+9.8% to +13.4%). Data visualization consistently improves (+8.7% to +10.2%), while analytical code generation achieves moderate gains on DeepSeek-distill-7B and Phi-4-13B but regresses by 1.5% on Phi-3.5-instruct. Task accuracy peaks in alignment with the peak associative reward score during training iterations.

  6. Knowl 6 — Human Alignment and Correlation of the Associative Creativity Reward

    data/table

    To assess the validity of the LLM-based associative creativity reward function, reward scores were compared against human-assigned creativity ratings on a sample of story generations.

    Model Correlation (Pearson's rr) Mean Reward Score Mean Human Rating
    Deepseek-distill-7B 0.75 0.68 3.5
    Phi-4-13B 0.78 0.72 3.8
    Phi-3.5-instruct 0.80 0.70 3.7

    The associative reward scores show strong positive correlation with human judgments across model families (r=0.75r = 0.75 to r=0.80r = 0.80, with an overall sample correlation of r=0.78r = 0.78). The reward metric also displays low variance across repeated evaluations on fixed candidate outputs, establishing scoring stability despite evaluator non-determinism.

  7. Knowl 7 — Cross-Domain Transfer of Associative Thinking to Analytical Tasks

    data/table

    Evaluating associative-trained models against base models demonstrates differential transfer between open-ended structural design tasks and strict formal logic tasks.

    Model Code Generation (%) Data Visualization (%)
    Deepseek-distill-7B 7.9 9.5
    Phi-4-13B 6.3 10.0
    Phi-3.5-instruct -1.2 8.5

    Data visualization consistently benefits from associative reinforcement (+8.5% to +10.0%), as chart generation combines analytical data shaping with creative visual hierarchy and element selection. In contrast, code generation exhibits mixed transfer (+6.3% to +7.9% for larger or distilled models, and a 1.2% decline for Phi-3.5-instruct), showing that divergent associative exploration can occasionally conflict with the strict syntactic and semantic precision required for program correctness.

  8. Knowl 8 — Limitations of LLM-Based Reward Grading and Reinforcement-Only Training

    limitation

    The associative reinforcement learning framework is subject to three primary methodological limitations:

    • Evaluator Circularity and Bias: Relying on prompt-based, checklist-guided LLMs to score creativity can introduce model-specific biases, overlook subtle human creative expressions, and cause alignment drift because both the policy generator and evaluator share similar language model architectures.
    • Degradation in Fluency and Factual Correctness: Reinforcement-only training without strict ground-truth constraints can occasionally degrade linguistic fluency or factual grounding, which is particularly evident in correctness-critical code outputs.
    • Restricted Benchmark Scope: Empirical evaluation is confined to three specific domains (storytelling, code generation, and chart visualization), leaving untested other creativity-critical areas such as scientific hypothesis formulation, musical composition, or social interaction design.

Coverage note — None was omitted; all primary contributions—including the Guilford-derived reward metric, the automated checklist grader, RL training mechanics, benchmark evaluations across storytelling/coding/visualization, human alignment validation, analytical domain transfer, and stated limitations—are fully covered.

References

  1. 1.J. Ahn, R. Verma, R. Lou, D. Liu, R. Zhang, and W. Yin. Large language models for mathematical reasoning: Progresses and challenges, 2024.
  2. 2.R. E. Beaty and Y. N. Kenett. Associative thinking at the core of creativity. Trends in Cognitive Sciences, 27(7):671–683, 2023.
  3. 3.S. P. Besemer. Creative product analysis matrix: testing the model structure and a comparison among products–three novel chairs. Creativity Research Journal, 11(4):333–346, 1998.
  4. 4.M. Cascella, J. Montomoli, V. Bellini, and E. Bignami. Evaluating the feasibility of chatgpt in healthcare: an analysis of multiple clinical and research scenarios. Journal of medical systems, 47(1):33, 2023.
  5. 5.F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, et al. Multipl-e: a scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering, 49(7):3675–3691, 2023.
  6. 6.T. Chakrabarty, P. Laban, D. Agarwal, S. Muresan, and C.-S. Wu. Art or artifice? large language models and the false promise of creativity. arXiv preprint arXiv:2309.14556, 2023.
  7. 7.DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.
  8. 8.F. Dell’Acqua, C. Ayoubi, H. Lifshitz, R. Sadun, E. Mollick, L. Mollick, Y. Han, J. Goldman, H. Nair, S. Taub, et al. The cybernetic teammate: A field experiment on generative ai reshaping teamwork and expertise. Technical report, National Bureau of Economic Research, 2025.
  9. 9.M. DeLorenzo, V. Gohil, and J. Rajendran. Creativeval: Evaluating creativity of llm-based hardware code generation. In 2024 IEEE LLM Aided Design Workshop (LAD), pages 1–5. IEEE, 2024.
  10. 10.J. Diedrich, M. Benedek, E. Jauk, and A. C. Neubauer. Are creative ideas novel and useful? Psychology of aesthetics, creativity, and the arts, 9(1):35, 2015.
  11. 11.M. Elgarf, H. Salam, and C. Peters. Fostering children’s creativity through llm-driven storytelling with a social robot. Frontiers in Robotics and AI, 11:1457429, 2024.
  12. 12.D. Fein, S. Russo, V. Xiang, K. Jolly, R. Rafailov, and N. Haber. Litbench: A benchmark and dataset for reliable evaluation of creative writing, 2025.
  13. 13.D. Henriksen, P. Mishra, and R. Mehta. Novel, effective, whole: Toward a new framework for evaluations of creative products. Journal of Technology and Teacher Education, 23(3):455–478, 2015.
  14. 14.S. B. Kaufman, C. G. DeYoung, J. R. Gray, J. Brown, and N. Mackintosh. Associative learning predicts intelligence above and beyond working memory and processing speed. Intelligence, 37(4):374–382, 2009.
  15. 15.H. Le, Y. Wang, A. D. Gotmare, S. Savarese, and S. C. H. Hoi. Coderl: Mastering code generation through pretrained models and deep reinforcement learning. Advances in Neural Information Processing Systems, 35:21314–21328, 2022.
  16. 16.Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, et al. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022.
  17. 17.S. Mednick. The associative basis of the creative process. Psychological review, 69(3):220, 1962.
  18. 18.Y. Mroueh. Reinforcement learning with verifiable rewards: Grpo’s effective loss, dynamics, and success amplification, 2025.
  19. 19.M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  20. 20.OpenAI, :, A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, A. Iftimie, A. Karpenko, A. T. Passos, A. Neitz, A. Prokofiev, A. Wei, A. Tam, A. Bennett, A. Kumar, A. Saraiva, A. Vallone, A. Duberstein, A. Kondrich, A. Mishchenko, A. Applebaum, A. Jiang, A. Nair, B. Zoph, B. Ghorbani, B. Rossen, B. Sokolowsky, B. Barak, B. McGrew, B. Minaiev, B. Hao, B. Baker, B. Houghton, B. McKinzie, B. Eastman, C. Lugaresi, C. Bassin, C. Hudson, C. M. Li, C. de Bourcy, C. Voss, C. Shen, C. Zhang, C. Koch, C. Orsinger, C. Hesse, C. Fischer, C. Chan, D. Roberts, D. Kappler, D. Levy, D. Selsam, D. Dohan, D. Farhi, D. Mely, D. Robinson, D. Tsipras, D. Li, D. Oprica, E. Freeman, E. Zhang, E. Wong, E. Proehl, E. Cheung, E. Mitchell, E. Wallace, E. Ritter, E. Mays, F. Wang, F. P. Such, F. Raso, F. Leoni, F. Tsimpourlas, F. Song, F. von Lohmann, F. Sulit, G. Salmon, G. Parascandolo, G. Chabot, G. Zhao, G. Brockman, G. Leclerc, H. Salman, H. Bao, H. Sheng, H. Andrin, H. Bagherinezhad, H. Ren, H. Lightman, H. W. Chung, I. Kivlichan, I. O’Connell, I. Osband, I. C. Gilaberte, I. Akkaya, I. Kostrikov, I. Sutskever, I. Kofman, J. Pachocki, J. Lennon, J. Wei, J. Harb, J. Twore, J. Feng, J. Yu, J. Weng, J. Tang, J. Yu, J. Q. Candela, J. Palermo, J. Parish, J. Heidecke, J. Hallman, J. Rizzo, J. Gordon, J. Uesato, J. Ward, J. Huizinga, J. Wang, K. Chen, K. Xiao, K. Singhal, K. Nguyen, K. Cobbe, K. Shi, K. Wood, K. Rimbach, K. Gu-Lemberg, K. Liu, K. Lu, K. Stone, K. Yu, L. Ahmad, L. Yang, L. Liu, L. Maksin, L. Ho, L. Fedus, L. Weng, L. Li, L. McCallum, L. Held, L. Kuhn, L. Kondraciuk, L. Kaiser, L. Metz, M. Boyd, M. Trebacz, M. Joglekar, M. Chen, M. Tintor, M. Meyer, M. Jones, M. Kaufer, M. Schwarzer, M. Shah, M. Yatbaz, M. Y. Guan, M. Xu, M. Yan, M. Glaese, M. Chen, M. Lampe, M. Malek, M. Wang, M. Fradin, M. McClay, M. Pavlov, M. Wang, M. Wang, M. Murati, M. Bavarian, M. Rohaninejad, N. McAleese, N. Chowdhury, N. Chowdhury, N. Ryder, N. Tezak, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, P. Chao, P. Ashbourne, P. Izmailov, P. Zhokhov, R. Dias, R. Arora, R. Lin, R. G. Lopes, R. Gaon, R. Miyara, R. Leike, R. Hwang, R. Garg, R. Brown, R. James, R. Shu, R. Cheu, R. Greene, S. Jain, S. Altman, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Hernandez, S. Baker, S. McKinney, S. Yan, S. Zhao, S. Hu, S. Santurkar, S. R. Chaudhuri, S. Zhang, S. Fu, S. Papay, S. Lin, S. Balaji, S. Sanjeev, S. Sidor, T. Broda, A. Clark, T. Wang, T. Gordon, T. Sanders, T. Patwardhan, T. Sottiaux, T. Degry, T. Dimson, T. Zheng, T. Garipov, T. Stasi, T. Bansal, T. Creech, T. Peterson, T. Eloundou, V. Qi, V. Kosaraju, V. Monaco, V. Pong, V. Fomenko, W. Zheng, W. Zhou, W. McCabe, W. Zaremba, Y. Dubois, Y. Lu, Y. Chen, Y. Cha, Y. Bai, Y. He, Y. Zhang, Y. Wang, Z. Shao, and Z. Li. Openai o1 system card, 2024.
  21. 21.J. A. Plucker, M. C. Makel, and M. Qian. Assessment of creativity. The Cambridge handbook of creativity, pages 48–73, 2010.
  22. 22.M. Rhodes. An analysis of creativity. The Phi delta kappan, 42(7):305–310, 1961.
  23. 23.M. A. Runco and G. J. Jaeger. The standard definition of creativity. Creativity research journal, 24(1):92–96, 2012.
  24. 24.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms, 2017.
  25. 25.S. Srivastava, S. Oberoi, and V. K. Gupta. The story and the storyteller: Strategic storytelling that gets human attention for entrepreneurs. Business Horizons, 66(3):347–358, 2023.
  26. 26.R. J. Sternberg and T. I. Lubart. The concept of creativity: Prospects and paradigms. Handbook of creativity, 1(3-15), 1999.
  27. 27.C. Stevenson, I. Smal, M. Baas, R. Grasman, and H. van der Maas. Putting gpt-3’s creativity to the (alternative uses) test. arXiv preprint arXiv:2206.08932, 2022.
  28. 28.S. Vennam, D. Valente, D. Herel, and P. Kumaraguru. Rethinking thinking tokens: Understanding why they underperform in practice, 2024.
  29. 29.F. Vinchon, V. Gironnay, and T. Lubart. Genai creativity in narrative tasks: Exploring new forms of creativity. Journal of Intelligence, 12(12):125, 2024.
  30. 30.X. Wang, L. Caccia, O. Ostapenko, X. Yuan, W. Y. Wang, and A. Sordoni. Guiding language model reasoning with planning tokens, 2024.
  31. 31.J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. 35:24824–24837, 2022.
  32. 32.T. Wilson and J. Schooler. Thinking too much: Introspection can reduce the quality of preferences and decisions. Journal of personality and social psychology, 60:181–92, 03 1991.
  33. 33.C.-L. Wu, S.-Y. Huang, P.-Z. Chen, and H.-C. Chen. A systematic review of creativity-related studies applying the remote associates test from 2000 to 2019. Frontiers in psychology, 11:573432, 2020.
  34. 34.S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601, 2023.

Citation

MLA
Singh, M., et al. “Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models”. arXiv, 2025, http://arxiv.org/abs/2511.17876v1.
APA
Singh, M., Singha, A., Parab, A., Mehrotra, P., & Gulwani, S. (2025). Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models. arXiv. http://arxiv.org/abs/2511.17876v1
Chicago
Singh, M., A. Singha, A. Parab, P. Mehrotra, and S. Gulwani. 2025. “Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models”. arXiv. http://arxiv.org/abs/2511.17876v1.
Harvard
Singh, M. et al. (2025) “Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2511.17876v1.
Vancouver
1. Singh M, Singha A, Parab A, Mehrotra P, Gulwani S (2025) Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models. arXiv

BibTeX

@article{singh2025training,
  title = {Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models},
  author = {Singh, Mukul and Singha, Ananya and Parab, Aishni and Mehrotra, Pronita and Gulwani, Sumit},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2511.17876v1},
  eprint = {2511.17876}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/