Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models
Mukul SinghAnanya SinghaAishni ParabPronita MehrotraSumit Gulwani
Demonstrates that training language models with reinforcement learning guided by cognitive associative thinking metrics improves their capacity for novel connections, abstraction, and problem-solving across creative writing, programming, and data visualization.
Large language models often struggle with creative tasks that require original synthesis and connecting conceptually distant ideas. Because standard training optimizes for predicting the most probable next word, these systems frequently default to familiar patterns and high semantic similarity rather than producing novel insights. This limitation presents a challenge for deploying artificial intelligence in tasks where creativity, abstraction, and flexible problem-solving are essential.
The article demonstrates that training language models using reinforcement learning grounded in cognitive principles of associative thinking enhances their performance across both creative and analytical domains. The authors evaluate whether rewarding models for generating conceptually distant and diverse associations improves outputs in creative writing, data visualization, and software coding.
To achieve this, the authors designed an automated evaluation mechanism based on four established divergent thinking metrics: novelty, fluency, flexibility, and elaboration. These metrics measure the rarity of concept combinations, the volume of distinct ideas, the diversity of semantic categories, and the level of explanatory detail. Using policy-gradient reinforcement learning algorithms, the authors fine-tuned several small and large language models on an 8-GPU computing cluster, with training runs averaging about two and a half hours per model. The automated reward function was validated against human evaluators across standard and custom benchmarks covering 5,000 storytelling tasks, over 18,000 code generation problems, and 500 visualization tasks.
The findings show that reinforcement learning focused on associative thinking produces overall performance gains of 8% to 13% across models and tasks. The highest improvements occurred in storytelling, where models saw gains ranging between 9.8% and 13.4%. Data visualization performance also improved consistently by 8.5% to 10.2%. For code generation, performance changes were more moderate, with improvements of 7.2% to 8.3% on larger architectures, though one instruction-tuned model experienced a slight regression of 1.2% to 1.5%. Additionally, the automated creativity reward demonstrated high stability and a strong positive correlation with human creativity scores (a correlation coefficient of 0.78).
These results indicate that embedding associative thinking principles into language model training improves an artificial intelligence system's capacity for abstraction and creative synthesis without requiring large-scale human annotation. Furthermore, the benefits extend beyond traditionally creative fields into analytical workflows such as chart design. However, the slight performance dips seen in certain coding tests indicate a potential trade-off between open-ended associative exploration and the strict syntactic precision required in programming.
Organizations developing or deploying language models for complex problem-solving should consider integrating associative reward mechanisms into their fine-tuning pipelines. Because the approach occasionally impacts precision in highly structured tasks, practitioners should carefully balance creativity rewards with correctness constraints when deploying models for mission-critical technical functions.
Key limitations include the reliance on language-model-based evaluators to score creativity, which carries a risk of evaluation bias and alignment drift. Additionally, the evaluation was restricted to three specific task domains and exhibited occasional degradations in factual grounding and fluency. While the reported gains are strong across the tested benchmarks, broader validation across additional creative and analytical fields is necessary before widespread deployment in safety-critical settings.
- Paper: DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning, A list of authors and their affiliations appears at the end of the paper (2025). This work establishes the paradigm of using reinforcement learning to elicit emergent, sophisticated reasoning and generative behaviors in base language models.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). It provides the foundational methodology for fine-tuning language models with reinforcement learning using reward signals tailored to complex, qualitative objectives.
- Paper: Deep Reinforcement Learning for Dialogue Generation, Jiwei Li et al. (2016). It introduces reinforcement learning reward modeling designed to promote informativity, novelty, and coherence in neural text generation.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). It outlines deliberate, multi-branch exploration strategies in language models, setting a baseline for evaluating conceptual connectivity and creative problem solving.
- Paper: Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs, Zhiyuan Hu et al. (2026). It extends RL-driven creative generation by designing uniqueness-aware rewards to combat exploration collapse and encourage rare problem-solving strategies.
- Paper: From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement, Qinsi Wang et al. (2026). It advances the reinforcement learning framework for open-ended and creative tasks by creating self-verifiable reward mechanisms without relying on human evaluators.
- Paper: Demystifying Reinforcement Learning Post-Training of Language Models, Donovan Clay et al. (2026). It analyzes the foundational mechanics of post-training RL to explain how reward design and base model priors govern the emergence of novel behaviors.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). It investigates how reinforcement learning cultivates internal multi-perspective and associative dialogue structures within reasoning models.
- Paper: Understanding Reasoning from Pretraining to Post-Training, Jingyan Shen et al. (2026). It investigates how pre-training representations interact with post-training reinforcement learning to promote or constrain novel generative capabilities.
