Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models
Raj JaiswalDhruv JainRishabh DhawanSree Krishna UppalapatiShin'ichi SatohTanuja GanuRajiv Ratn Shah
Proposes a step-level reinforcement learning framework that trains small language models to correct multi-step physics derivation errors using structured feedback, achieving up to 20% accuracy gains across five benchmarks without requiring human preference data or test-time verifiers.
Physics problem-solving requires sequential, multi-step reasoning where a single mistake early in a derivation invalidates all subsequent steps. While large artificial intelligence models mitigate this issue through massive scale and broad knowledge compression, small language models under four billion parameters frequently fail due to error propagation, limited domain knowledge, and arithmetic errors. Addressing these failures is increasingly important as organizations seek to deploy cost-effective, energy-efficient, and low-latency models for scientific and technical tasks without relying on massive computing infrastructure.
The article develops and evaluates a step-level reward training framework designed to improve physics reasoning in small language models. The primary objective is to demonstrate that targeted, step-level feedback provided only during training enables small models to self-correct and improve problem-solving accuracy without exposing them to ground-truth answers or requiring costly human preference annotations.
The researchers evaluated the approach using four open-source models ranging from 1 billion to 3.8 billion parameters across five physics benchmarks of varying difficulty, including high school, undergraduate, and advanced competitive entrance exam levels. The methodology follows three stages: first, models are warm-started using supervised fine-tuning on 2,494 structured physics problems to establish consistent step formatting; second, during training, an external verifier identifies the exact step and type of the first reasoning error to calculate a position-based reward that penalizes earlier failures more heavily; third, structured feedback tailored to the specific error type—problem miscomprehension, conceptual misapplication, or calculation error—is provided to guide a revision, updating model weights via reinforcement learning. At inference time, the trained models operate independently without any external verifier or extra computing overhead.
The evaluation produced several key findings. First, the proposed framework achieved substantial performance improvements, delivering 17% to 20% accuracy gains over standard chain-of-thought prompting and outperforming strong baselines, including retrieval-augmented generation and direct preference optimization, by 10% to 16% across all tested benchmarks. Second, the peak performance gain reached 27.1% on the highly challenging JEEBench dataset with a 3-billion-parameter model. Third, error-specific feedback reduced calculation errors from 56.9% down to 23.5% and problem miscomprehension errors from 22.3% to 12.0% in the best observed cases. Fourth, conceptual misapplication proved to be the most stubborn failure mode; although reduced from 89.7% to 68.7% in the best case, it remained above 42% across all models and conditions, showing that providing correct governing formulas does not ensure their correct application.
These findings indicate that small, lightweight models can achieve significant reasoning gains when trained with targeted error localization rather than full-solution preference signals, lowering operational costs and latency for technical deployments. The results also reveal that standard preference optimization methods can degrade domain-specific reasoning by encouraging surface-level fluency without underlying logical accuracy. However, because conceptual errors remain high, organizations cannot rely on small models for fully autonomous technical decision-making without expert supervision.
Organizations developing or deploying automated reasoning systems should consider adopting step-level reward mechanisms to improve model efficiency during training. For operational deployment, hybrid workflows are recommended: small models can handle structured derivations and calculations, while human experts or larger models should verify core conceptual formulations. Future work should focus on developing improved training techniques specifically targeting conceptual understanding, testing multi-domain generalization beyond physics, and evaluating multi-seed consistency across training runs.
The primary limitations of the study include its focus solely on English-language physics benchmarks, the lack of multi-seed validation, and the framework's reliance on an external advanced model for training-time verification, where any verifier misclassification could misdirect training. Confidence in the reported accuracy gains and arithmetic error reductions is high across the tested benchmarks, but caution is warranted regarding the framework's ability to resolve deep conceptual misconceptions.
- Paper: Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, Peiyi Wang et al. (2024). Math-Shepherd establishes how to derive automated step-level process supervision and rewards without human annotation, which directly underpins the training-time step-verification framework developed in this work.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). This seminal paper introduces process-supervised reward modeling for step-by-step reasoning verification up to the first error, providing the core foundational paradigm that the source adapts.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). This work formulates multi-step reasoning as a sequential decision process and analyzes step-level error impact, providing essential theoretical context for step-level reward modeling.
- Paper: Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement, Weimin Xiong et al. (2024). This paper develops step-level refinement and iterative process feedback, offering foundational mechanisms for correcting intermediate reasoning missteps.
- Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). This research analyzes how small language models struggle with multi-step reasoning chains and explores transferring structured chain-of-thought capabilities into compact models.
- Paper: ReFT: Reasoning with Reinforced Fine-Tuning, Luong Quoc Trung et al. (2024). ReFT establishes policy-gradient reinforcement learning frameworks for reasoning path exploration, directly informing the policy gradient optimization used by the source.
- Paper: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, Charlie Snell et al. (2024). This study examines iterative revision chains and process verifiers for multi-step reasoning, providing key concepts for step-level refinement without human annotations.
- Paper: Experiential Reinforcement Learning, Taiwei Shi et al. (2026). Experiential Reinforcement Learning extends the principle of structured reflection and behavioral refinement during training so that models internalize self-corrections without requiring test-time overhead.
- Paper: When Can LLMs Learn to Reason with Weak Supervision?, Salman Rahman et al. (2026). This work explores the boundaries of training-time reinforcement learning under weak or proxy supervision signals across diverse scientific and mathematical reasoning tasks.
- Paper: A Primer in Post-Training Reasoning Data: What We Know About How It Works, Yaoming Li et al. (2026). This comprehensive post-training primer synthesizes how intermediate feedback channels, verifier taxonomies, and reasoning trajectories function across modern reasoning models.
- Paper: Fractured Chain-of-Thought Reasoning, Baohao Liao et al. (2026). This paper investigates fractured sampling and intermediate depth evaluation to optimize reasoning chains and token efficiency during multi-step inference.
- Paper: SPIRAL: Learning to Search and Aggregate, Jubayer Ibn Hamid et al. (2026). SPIRAL advances beyond single-trace step correction by training models via reinforcement learning to search across parallel reasoning paths and aggregate candidate traces.
