Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models

Raj JaiswalDhruv JainRishabh DhawanSree Krishna UppalapatiShin'ichi SatohTanuja GanuRajiv Ratn Shah

article2026arXiv0 citations

Proposes a step-level reinforcement learning framework that trains small language models to correct multi-step physics derivation errors using structured feedback, achieving up to 20% accuracy gains across five benchmarks without requiring human preference data or test-time verifiers.

Listen

Physics problem-solving requires sequential, multi-step reasoning where a single mistake early in a derivation invalidates all subsequent steps. While large artificial intelligence models mitigate this issue through massive scale and broad knowledge compression, small language models under four billion parameters frequently fail due to error propagation, limited domain knowledge, and arithmetic errors. Addressing these failures is increasingly important as organizations seek to deploy cost-effective, energy-efficient, and low-latency models for scientific and technical tasks without relying on massive computing infrastructure.

The article develops and evaluates a step-level reward training framework designed to improve physics reasoning in small language models. The primary objective is to demonstrate that targeted, step-level feedback provided only during training enables small models to self-correct and improve problem-solving accuracy without exposing them to ground-truth answers or requiring costly human preference annotations.

The researchers evaluated the approach using four open-source models ranging from 1 billion to 3.8 billion parameters across five physics benchmarks of varying difficulty, including high school, undergraduate, and advanced competitive entrance exam levels. The methodology follows three stages: first, models are warm-started using supervised fine-tuning on 2,494 structured physics problems to establish consistent step formatting; second, during training, an external verifier identifies the exact step and type of the first reasoning error to calculate a position-based reward that penalizes earlier failures more heavily; third, structured feedback tailored to the specific error type—problem miscomprehension, conceptual misapplication, or calculation error—is provided to guide a revision, updating model weights via reinforcement learning. At inference time, the trained models operate independently without any external verifier or extra computing overhead.

The evaluation produced several key findings. First, the proposed framework achieved substantial performance improvements, delivering 17% to 20% accuracy gains over standard chain-of-thought prompting and outperforming strong baselines, including retrieval-augmented generation and direct preference optimization, by 10% to 16% across all tested benchmarks. Second, the peak performance gain reached 27.1% on the highly challenging JEEBench dataset with a 3-billion-parameter model. Third, error-specific feedback reduced calculation errors from 56.9% down to 23.5% and problem miscomprehension errors from 22.3% to 12.0% in the best observed cases. Fourth, conceptual misapplication proved to be the most stubborn failure mode; although reduced from 89.7% to 68.7% in the best case, it remained above 42% across all models and conditions, showing that providing correct governing formulas does not ensure their correct application.

These findings indicate that small, lightweight models can achieve significant reasoning gains when trained with targeted error localization rather than full-solution preference signals, lowering operational costs and latency for technical deployments. The results also reveal that standard preference optimization methods can degrade domain-specific reasoning by encouraging surface-level fluency without underlying logical accuracy. However, because conceptual errors remain high, organizations cannot rely on small models for fully autonomous technical decision-making without expert supervision.

Organizations developing or deploying automated reasoning systems should consider adopting step-level reward mechanisms to improve model efficiency during training. For operational deployment, hybrid workflows are recommended: small models can handle structured derivations and calculations, while human experts or larger models should verify core conceptual formulations. Future work should focus on developing improved training techniques specifically targeting conceptual understanding, testing multi-domain generalization beyond physics, and evaluating multi-seed consistency across training runs.

The primary limitations of the study include its focus solely on English-language physics benchmarks, the lack of multi-seed validation, and the framework's reliance on an external advanced model for training-time verification, where any verifier misclassification could misdirect training. Confidence in the reported accuracy gains and arithmetic error reductions is high across the tested benchmarks, but caution is warranted regarding the framework's ability to resolve deep conceptual misconceptions.

arXiv: 2607.05199
  • Paper: Experiential Reinforcement Learning, Taiwei Shi et al. (2026). Experiential Reinforcement Learning extends the principle of structured reflection and behavioral refinement during training so that models internalize self-corrections without requiring test-time overhead.
  • Paper: When Can LLMs Learn to Reason with Weak Supervision?, Salman Rahman et al. (2026). This work explores the boundaries of training-time reinforcement learning under weak or proxy supervision signals across diverse scientific and mathematical reasoning tasks.
  • Paper: A Primer in Post-Training Reasoning Data: What We Know About How It Works, Yaoming Li et al. (2026). This comprehensive post-training primer synthesizes how intermediate feedback channels, verifier taxonomies, and reasoning trajectories function across modern reasoning models.
  • Paper: Fractured Chain-of-Thought Reasoning, Baohao Liao et al. (2026). This paper investigates fractured sampling and intermediate depth evaluation to optimize reasoning chains and token efficiency during multi-step inference.
  • Paper: SPIRAL: Learning to Search and Aggregate, Jubayer Ibn Hamid et al. (2026). SPIRAL advances beyond single-trace step correction by training models via reinforcement learning to search across parallel reasoning paths and aggregate candidate traces.
Cover for Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models

Abstract

Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucination under multi-step derivation, and distributional sensitivity compound this failure. We propose a step-level reward framework that identifies the first reasoning error, generates targeted structured feedback, and trains the model to revise its solution via policy gradient with KL regularization, without exposing it to ground truth solutions as generation targets. Unlike annotation-dependent step-level methods, no preference data construction is required and the external verifier operates exclusively at training time. Across five physics benchmarks, our framework delivers accuracy gains of 17-20% over CoT prompting and 10-16% over the strongest baseline, reduces calculation errors from 56.9% to 23.5%, and reduces miscomprehension errors from 22.3% to 12.0% in the best observed cases. Conceptual errors reduce from 89.7% to 68.7%, yet persist as the hardest failure mode across all conditions.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Stage 1: Supervised Fine-Tuning Warm-up
  • 3.2 Stage 2: Step-Level Reward Mechanism
  • 3.3 Stage 3: Feedback Generation
  • 3.4 Policy Update
  • 4 Experiments
  • 4.1 Benchmark Datasets
  • 4.2 Models
  • 4.3 Baseline Setup
  • 4.4 Evaluation
  • 5 Results
  • 5.1 Accuracy Increases With Model Scale.
  • 5.2 Accuracy Across Conditions.
  • 5.3 Reasoning Error Distribution Across Conditions.
  • 6 Discussion
  • 6.1 Structured Error Feedback Improvement.
  • 6.2 Where Step-Level Reward and Feedback Fall Short.
  • 7 Conclusion
  • 8 Limitations
  • 9 Ethical Considerations
  • References
  • A Verifier Prompt
  • B Step-Level Reward Computation
  • C Structured Feedback Generation
  • D Policy Gradient Training Objective
  • E Training Hyperparameters

Knowls

  1. Knowl 1 — Step-Level Reward and Structured Feedback Framework for Physics Reasoning

    model/method

    The step-level reward framework improves multi-step physics reasoning in small language models (SLMs) without requiring step-by-step human preference datasets or exposing models to ground truth reference solutions as generation targets. The training process operates in three successive stages:

    1. Stage 1: Supervised Fine-Tuning (SFT) Warm-Up: The model policy πθ\pi_\theta is fine-tuned using Low-Rank Adaptation (LoRA) on a domain corpus of physics problems paired with step-indexed chain-of-thought solutions divided into seven distinct XML tags. This establishes the structural consistency needed for step identification. Once Stage 1 converges, the adapter weights are frozen to serve as a fixed reference policy πref\pi_{\text{ref}}, and πθ\pi_\theta is initialized from these weights.
    2. Stage 2: Step-Level Error Identification and Reward Assignment: Given a problem xx, the policy generates an initial solution attempt y1∼πθ(⋅∣x)y_1 \sim \pi_\theta(\cdot \mid x). An external verifier evaluates y1y_1 against the ground truth reference solution y∗y^* to locate the index of the first reasoning error efirst∈{1,…,n}e_{\text{first}} \in \{1, \dots, n\} and categorize the failure into one of three classes: Problem Miscomprehension (MC), Conceptual Misapplication (CM), or Calculation Error (CE). The verifier computes a position-dependent step reward r1=efirst/(n+1)r_1 = e_{\text{first}} / (n + 1), where nn is the total step count of y1y_1. If r1>τr_1 > \tau (with early-stopping threshold τ=0.9\tau = 0.9), the sample is considered correct and skipped.
    3. Stage 3: Error-Conditioned Feedback and Policy Optimization: For flawed initial attempts (r1≤τr_1 \le \tau), a feedback module generates structured corrective feedback p1p_1 conditioned on the specific error class. Conditioned on the problem xx, first attempt y1y_1, and feedback p1p_1, the model generates a revised solution y2∼πθ(⋅∣x,y1,p1)y_2 \sim \pi_\theta(\cdot \mid x, y_1, p_1). The verifier evaluates y2y_2 to determine a second reward r2r_2. The policy πθ\pi_\theta is updated via a policy gradient objective with Kullback-Leibler (KL) divergence regularization anchored to πref\pi_{\text{ref}}.

    The external verifier and feedback generators operate exclusively during training; at test time, the fine-tuned model solves problems autonomously in a standard single-pass or chain-of-thought manner.

  2. Knowl 2 — Step-Level Position-Dependent Reward Function

    equation

    The scalar reward rr assigned to a generated multi-step reasoning trajectory is defined as:

    r={1.0if no reasoning error is detected (efirst=0),efirstn+1if a reasoning error occurs at step index efirst,r = \begin{cases} 1.0 & \text{if no reasoning error is detected } (e_{\text{first}} = 0), \\ \frac{e_{\text{first}}}{n + 1} & \text{if a reasoning error occurs at step index } e_{\text{first}}, \end{cases}

    where:

    • efirst∈{1,2,…,n}e_{\text{first}} \in \{1, 2, \dots, n\} is the positive integer step index of the earliest reasoning failure where the generated solution deviates conceptually, mathematically, or factually from the reference solution y∗y^*.
    • n∈N≥1n \in \mathbb{N}_{\ge 1} is the total number of reasoning steps present in the generated solution.
    • The normalization denominator n+1n + 1 guarantees r∈(0,1)r \in (0, 1) for any erroneous solution, assigning higher scalar rewards to solutions that maintain physical correctness deeper into the derivation chain and penalizing earlier failure points more severely.
  3. Knowl 3 — Policy Gradient Loss with Reference Policy KL Regularization

    equation

    The policy parameters θ\theta are updated by minimizing the loss function L(θ)\mathcal{L}(\theta) computed over revised reasoning trajectories:

    L(θ)=−[(log⁡πθ(y2∣x,y1,p1)−log⁡πθ(y1∣x))⋅r2−βDKL(πθ(⋅∣x)∥πref(⋅∣x))]\mathcal{L}(\theta) = -\left[ \left(\log \pi_\theta(y_2 \mid x, y_1, p_1) - \log \pi_\theta(y_1 \mid x)\right) \cdot r_2 - \beta D_{\text{KL}}\left(\pi_\theta(\cdot \mid x) \parallel \pi_{\text{ref}}(\cdot \mid x)\right) \right]

    where:

    • xx is the input physics problem.
    • y1∼πθ(⋅∣x)y_1 \sim \pi_\theta(\cdot \mid x) is the initial generated solution attempt; its log-probability log⁡πθ(y1∣x)\log \pi_\theta(y_1 \mid x) is treated as a detached constant baseline.
    • p1p_1 is the error-type-conditioned structured feedback generated for attempt y1y_1.
    • y2∼πθ(⋅∣x,y1,p1)y_2 \sim \pi_\theta(\cdot \mid x, y_1, p_1) is the revised solution generated by the policy; gradients are computed with respect to y2y_2.
    • r2∈(0,1]r_2 \in (0, 1] is the step-level reward assigned to the revision y2y_2 based on its first error index e2e_2 (r2=1.0r_2 = 1.0 if y2y_2 contains no reasoning error).
    • πref\pi_{\text{ref}} is the frozen reference policy checkpoint initialized from Stage 1 supervised fine-tuning.
    • DKL(πθ(⋅∣x)∥πref(⋅∣x))D_{\text{KL}}(\pi_\theta(\cdot \mid x) \parallel \pi_{\text{ref}}(\cdot \mid x)) is the forward Kullback-Leibler divergence between the active policy and reference policy conditioned only on input xx, preventing policy collapse independently of feedback conditioning.
    • β>0\beta > 0 is the regularization coefficient controlling the KL penalty weight (set to β=0.1\beta = 0.1).
  4. Knowl 4 — Step-Level Reward Training Algorithm for Physics Reasoning

    algorithm

    The algorithm trains a small language model policy πθ\pi_\theta using step-level failure verification, error-type-conditioned structured feedback, and policy gradient optimization.

    Input: Training set D={(x,y∗)}D = \{(x, y^*)\}, policy πθ\pi_\theta warm-started from Stage 1 SFT, frozen reference policy πref\pi_{\text{ref}}, external verifier VV, feedback generator FF, threshold τ=0.9\tau = 0.9, regularization strength β=0.1\beta = 0.1.
    Output: Optimized policy parameters θ\theta.
    for each training iteration do
        Sample problem-solution pair (x,y∗)∼D(x, y^*) \sim D
        Generate initial attempt y1∼πθ(⋅∣x)y_1 \sim \pi_\theta(\cdot \mid x)
        Extract step count n1n_1 from y1y_1
        (e1,c1,expl1)←V(y1,y∗)(e_1, c_1, \text{expl}_1) \leftarrow V(y_1, y^*)
        if e1=0e_1 = 0 then
            r1←1.0r_1 \leftarrow 1.0
        else
            r1←e1/(n1+1)r_1 \leftarrow e_1 / (n_1 + 1)
        end if
        if r1>τr_1 > \tau then
            continue
        end if
        p1←F(x,y1,e1,c1,expl1)p_1 \leftarrow F(x, y_1, e_1, c_1, \text{expl}_1)
        Generate revised attempt y2∼πθ(⋅∣x,y1,p1)y_2 \sim \pi_\theta(\cdot \mid x, y_1, p_1)
        Extract step count n2n_2 from y2y_2
        (e2,c2,expl2)←V(y2,y∗)(e_2, c_2, \text{expl}_2) \leftarrow V(y_2, y^*)
        if e2=0e_2 = 0 then
            r2←1.0r_2 \leftarrow 1.0
        else
            r2←e2/(n2+1)r_2 \leftarrow e_2 / (n_2 + 1)
        end if
        L(θ)←−[(log⁡πθ(y2∣x,y1,p1)−log⁡πθ(y1∣x))⋅r2−βDKL(πθ(⋅∣x)∥πref(⋅∣x))]\mathcal{L}(\theta) \leftarrow -[(\log \pi_\theta(y_2 \mid x, y_1, p_1) - \log \pi_\theta(y_1 \mid x)) \cdot r_2 - \beta D_{\text{KL}}(\pi_\theta(\cdot \mid x) \parallel \pi_{\text{ref}}(\cdot \mid x))]
        Update θ\theta by taking a gradient descent step on ∇θL(θ)\nabla_\theta \mathcal{L}(\theta)
    end for
    return πθ\pi_\theta

    Gradients are backpropagated solely through y2y_2. The initial log-probability log⁡πθ(y1∣x)\log \pi_\theta(y_1 \mid x) is treated as a baseline and detached from the computational graph.

  5. Knowl 5 — Error Taxonomy and Multi-Channel Feedback Generation

    model/method

    Physics reasoning errors in generated solutions are categorized into three hierarchical, interdependent classes, evaluated in strict priority order (MC→CM→CEMC \rightarrow CM \rightarrow CE):

    1. Problem Miscomprehension (MC): The model misreads given values, target objectives, or physical boundary constraints before applying physical principles. Feedback is produced via direct structured prompting instructing the model to re-examine the given quantities, variables, and goal state before regenerating.
    2. Conceptual Misapplication (CM): The model invokes an incorrect physical law, an invalid formula, or applies a principle outside its domain of validity (e.g., applying monatomic heat capacities to diatomic gases). Feedback is generated using a retrieval-augmented agent (FeedbackRAGAgent) that queries a ChromaDB vector store of high school physics formulas (embedded via all-MiniLM-L6-v2) and supplies the model with the relevant physical theorem and error explanation.
    3. Calculation Error (CE): The model correctly selects the physical principle and sets up governing equations, but makes an arithmetic, algebraic manipulation, or calculus mistake during execution. Feedback is generated by an execution agent (CodeAgent) that writes and executes Python code for the exact computation required at the failure step, returning the verified numerical evaluation.

    Routing feedback through dedicated specialized mechanisms ensures that conceptual corrections provide physical grounding while computational corrections leverage deterministic numerical execution.

  6. Knowl 6 — Seven-Tag Structured XML Format for Physics Reasoning

    model/method

    To prevent error propagation across multi-step derivations and enable precise step-level indexing by external verifiers, solutions are structured into seven sequential XML tags:

    1. <problem_analysis>: Restatement of the given physical quantities, target variables, domain, and explicit assumptions.
    2. <principle>: Identification of the governing physical laws and principles described in plain English.
    3. <governing_equation>: Symbolic mathematical formulation of the governing equations without numerical substitutions.
    4. <value_identification>: Comprehensive itemization of all known numeric quantities and physical constants along with their SI units.
    5. <substitution>: Step-by-step algebraic substitution of identified numeric values into the governing equations.
    6. <calculation>: Step-by-step arithmetic and algebraic computation carrying units throughout.
    7. <final_answer>: The definitive result with appropriate units enclosed inside \boxed{}.

    Enforcing this schema during Stage 1 SFT separates physical reasoning from symbolic and numeric calculation, ensuring consistent regex matching for step counting and error localization during verification.

  7. Knowl 7 — Benchmark Accuracy Across Small Language Models and Baselines

    data/table

    Final answer accuracy (%) evaluated across five benchmarks for four open-source small language models under five training/prompting conditions: Chain-of-Thought (CoT), Retrieval-Augmented Generation (RAG), Supervised Fine-Tuning (SFT), Direct Preference Optimization with IPO loss (DPO), and the step-level reward framework (Ours).

    Benchmark Model CoT RAG SFT DPO Ours
    SciEval-Static Qwen 2.5 1.5B 62.73 68.12 57.93 59.32 79.29
    LLaMA 3.2 1B 44.51 57.37 55.49 56.93 68.09
    LLaMA 3.2 3B 62.26 66.87 61.59 57.93 81.52
    Phi 3.5 Mini 3.8B 62.43 61.01 59.17 56.87 65.19
    MMLU-High Qwen 2.5 1.5B 47.05 59.41 41.18 46.47 55.03
    LLaMA 3.2 1B 30.58 38.02 42.35 39.41 55.29
    LLaMA 3.2 3B 50.00 55.29 48.29 47.65 67.20
    Phi 3.5 Mini 3.8B 55.17 60.53 55.97 56.12 65.83
    MMLU-College Qwen 2.5 1.5B 52.62 62.27 43.22 43.22 53.64
    LLaMA 3.2 1B 33.63 45.82 38.14 36.44 57.89
    LLaMA 3.2 3B 53.63 61.82 45.76 50.85 69.09
    Phi 3.5 Mini 3.8B 65.63 67.27 62.77 64.34 73.12
    JEEBench Qwen 2.5 1.5B 37.53 40.00 39.02 30.08 54.86
    LLaMA 3.2 1B 32.52 35.50 28.46 35.77 49.17
    LLaMA 3.2 3B 30.21 40.83 36.59 38.21 57.34
    Phi 3.5 Mini 3.8B 35.90 41.25 38.04 41.65 50.23
    PhysicsQA Qwen 2.5 1.5B 30.64 36.91 23.51 24.32 49.62
    LLaMA 3.2 1B 21.35 28.64 23.24 18.92 39.02
    LLaMA 3.2 3B 27.67 31.21 27.84 25.95 46.73
    Phi 3.5 Mini 3.8B 33.35 41.59 39.19 41.49 53.22

    The step-level reward framework consistently attains the highest overall performance across models and benchmarks, outperforming standard CoT prompting by 17–20% on average and the strongest baseline (RAG) by 10–16%. Standard SFT and DPO frequently degrade accuracy relative to CoT or show inconsistent gains due to holistic response-level loss formulations.

  8. Knowl 8 — Distribution of Physics Reasoning Errors Across Training Conditions

    data/table

    Proportion (%) of Problem Miscomprehension (MC), Conceptual Misapplication (CM), and Calculation Errors (CE) observed among incorrect model generations on PhysicsQA across experimental conditions.

    Model MC (%) CM (%) CE (%)
    CoT RAG SFT DPO Ours CoT RAG SFT DPO Ours CoT RAG SFT DPO Ours
    Qwen 2.5 1.5B 9.3 15.0 4.6 16.1 4.1 38.6 61.4 49.5 83.9 66.1 52.1 24.5 35.0 38.9 31.4
    LLaMA 3.2 1B 22.3 34.1 38.7 25.3 12.0 89.7 79.5 93.0 97.3 68.7 55.0 45.5 49.6 30.0 27.5
    LLaMA 3.2 3B 7.9 5.9 6.4 9.1 12.0 73.8 47.2 55.8 80.7 42.7 37.8 44.9 40.4 38.7 38.5
    Phi 3.5 Mini 3.8B 8.1 6.0 6.7 8.3 6.4 50.8 57.9 44.6 55.6 56.7 56.9 26.4 62.5 27.8 23.5

    The data shows that Calculation Errors (CE) reduce consistently across most models under the step-level reward framework (e.g., dropping from 56.9% to 23.5% in Phi 3.5 Mini and from 55.0% to 27.5% in LLaMA 3.2 1B). In contrast, Conceptual Misapplication (CM) remains the most prominent and resistant error mode across all models and settings, rising substantially under standard DPO (up to 97.3% in LLaMA 3.2 1B and 83.9% in Qwen 2.5 1.5B) and remaining above 42% even with step-level feedback.

  9. Knowl 9 — Comparative Impact of Error-Conditioned Feedback Channels

    empirical result

    Ablation analysis on PhysicsQA comparing the error reduction (%) of individual feedback channels (Problem Statement / MC, Conceptual / CM, Calculation / CE) against baseline conditions reveals distinct error dynamics:

    1. Calculation Feedback Consistency: Calculation feedback is the most reliably effective channel, driving significant CE reductions across models regardless of baseline (e.g., achieving +33.4% error reduction over CoT and +39.0% over SFT for Phi 3.5 Mini 3.8B, and +27.5% over CoT for LLaMA 3.2 1B).
    2. Concept Feedback vs. Preference Baselines: Conceptual feedback produces the largest reductions where response-level DPO caused severe conceptual degradation, reducing CM errors by +17.8% (Qwen 2.5 1.5B), +28.6% (LLaMA 3.2 1B), and +38.0% (LLaMA 3.2 3B) relative to DPO.
    3. Problem Statement Feedback Sensitivity: Statement-level MC feedback exhibits variable outcomes across models: it reduces MC errors by +10.3% over CoT and +22.1% over RAG for LLaMA 3.2 1B, but slightly increases MC errors for LLaMA 3.2 3B (-4.1% relative to CoT). This indicates that heavily prioritizing conceptual corrections during training can inadvertently disrupt problem parsing and initial value extraction.
  10. Knowl 10 — Limitations of the Step-Level Physics Reasoning Framework

    limitation

    The step-level reward training and structured feedback framework possesses several documented constraints:

    1. Language and Domain Generalization: Evaluation is restricted to English-language physics benchmarks; applicability to multilingual settings or other scientific disciplines (such as chemistry or biology) has not been validated.
    2. External Verifier Dependency: The framework depends entirely on GPT-4o for training-time error identification and feedback generation. Systematic misclassifications by the verifier route incorrect feedback and distort the reward signal without being directly detectable from downstream accuracy metrics.
    3. Formatting Fragility: Reward computation assumes strict compliance with the seven-tag XML format and regex step detection. Any non-standard formatting or JSON decoding failure defaults to a minimal reward score.
    4. Incomplete Conceptual Resolution: Conceptual Misapplication errors remain above 42% across all tested small language models. External retrieval provides correct equations but does not guarantee their correct physical application in novel problem setups.
    5. Evaluation Protocol and Logging: Reported results are based on single training runs over 60 epochs without multi-seed confidence intervals, and intermediate sample skipping rates (r1>τr_1 > \tau) were not logged during training.

Coverage note — All core contributions—including the three-stage framework, mathematical reward formulation, policy gradient loss with KL regularization, training algorithm, seven-tag schema, multi-channel feedback mechanism, complete experimental benchmark evaluations, error distribution analyses, feedback ablations, and stated limitations—are fully covered. Raw verifier prompt strings and domain formula listings from the appendix were omitted as implementation details.

References

  1. 1.Avinash Anand, Kritarth Prasad, Chhavi Kirtani, Ashwin R Nair, Mohit Gupta, Saloni Garg, Anurag Gautam, Snehal Buldeo, and Rajiv Ratn Shah. 2024. Enhancing llms for physics problem-solving using reinforcement learning with human-ai feedback. Preprint, arXiv:2412.06827.
  2. 2.Daman Arora, Himanshu Gaurav Singh, and 1 others. 2023. Have LLMs advanced enough? A challenging problem solving benchmark for large language models. arXiv preprint arXiv:2305.15074.
  3. 3.Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos. 2023. A general theoretical paradigm to understand learning from human preferences. Preprint, arXiv:2310.12036.
  4. 4.Johan Boye and Birger Myrberg. 2025. Large language models for physics reasoning. arXiv preprint arXiv:2502.11537.
  5. 5.Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan. 2024. Step-level value preference optimization for mathematical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2024.
  6. 6.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30.
  7. 7.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, and 1 others. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  8. 8.Jingzhe Ding, Yan Cen, and Xinyuan Wei. 2023. Using large language model to solve and explain physics word problems approaching human level. arXiv preprint arXiv:2309.08182.
  9. 9.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2023. Complexity-based prompting for multi-step reasoning. In The Eleventh International Conference on Learning Representations.
  10. 10.Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, and 1 others. 2024. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  11. 11.David Halliday, Robert Resnick, and Jearl Walker. 2014. Fundamentals of Physics, 10th edition. Wiley.
  12. 12.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  13. 13.Namgyu Ho, Laura Schmid, and Se-Young Yun. 2022. Large language models are reasoning teachers. arXiv preprint arXiv:2212.10071.
  14. 14.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  15. 15.Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023. Large language models cannot self-correct reasoning yet. arXiv preprint arXiv:2310.01848.
  16. 16.Raj Jaiswal, Dhruv Jain, Harsh Parimal Popat, Avinash Anand, Abhishek Dharmadhikari, Atharva Marathe, and Rajiv Ratn Shah. 2024. Improving physics reasoning in large language models using mixture of refinement agents. Preprint, arXiv:2412.00821.
  17. 17.Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, and 1 others. 2024. Training language models to self-correct via reinforcement learning. arXiv preprint arXiv:2409.12917.
  18. 18.Xin Lai, Zhuotao Tian, Yukang Chen, Senqiao Yang, Xiangru Peng, and Jiaya Jia. 2024a. Step-DPO: Stepwise preference optimization for long-chain reasoning of LLMs. arXiv preprint arXiv:2406.18629.
  19. 19.Xin Lai, Zhuotao Tian, Yukang Chen, Senqiao Yang, Xiangru Peng, and Jiaya Jia. 2024b. Step-DPO: Stepwise preference optimization for long-chain reasoning of LLMs. arXiv preprint arXiv:2406.18629.
  20. 20.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, and 1 others. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  21. 21.Xiaoze Li, Feifei Sun, Yixin Zhang, Xinlong Xu, Jielong Jiang, Bhaskar Mitra, and Lucian Popa. 2024. Evaluating the factuality of large language models using large-scale knowledge graphs. arXiv preprint arXiv:2404.00216.
  22. 22.Yen-Ting Lin, Di Jin, Tengyu Xu, Tianhao Wu, Sainbayar Sukhbaatar, Chen Zhu, Yun He, Yun-Nung Chen, Jason Weston, Yuandong Tian, Arash Rahnama, Sinong Wang, Hao Ma, and Han Fang. 2025. Step-kto: Optimizing mathematical reasoning through stepwise binary feedback. Preprint, arXiv:2501.10799.
  23. 23.Elita Lobo, Chirag Agarwal, and Himabindu Lakkaraju. 2024. On the impact of fine-tuning on chain-of-thought reasoning. arXiv preprint arXiv:2411.15382.
  24. 24.Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023. WizardMath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. arXiv preprint arXiv:2308.09583.
  25. 25.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, and 1 others. 2024. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36.
  26. 26.Michael McCloskey and Neal J Cohen. 1989. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of Learning and Motivation, volume 24, pages 109–165. Elsevier.
  27. 27.Microsoft Research. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219.
  28. 28.National Testing Agency. 2023. NEET physics formula sheet. Physics formula reference used for retrieval corpus construction.
  29. 29.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  30. 30.D. C. Pandey. 2020. IIT JEE Physics: 35 Years Chapterwise Solved Papers. Arihant Publications.
  31. 31.Xinyu Pang, Ruixin Hong, Zhanke Zhou, and 1 others. 2024. Physics reasoner: Knowledge-augmented reasoning for solving physics problems with large language models. arXiv preprint arXiv:2412.13791.
  32. 32.A. A. Pinsky. 1989. Problems in Physics. Mir Publishers.
  33. 33.Qwen Team. 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671.
  34. 34.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36:53728–53741.
  35. 35.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, pages 3982–3992.
  36. 36.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36.
  37. 37.Aarohi Srivastava and 1 others. 2025. Reasoning ability of small language models. arXiv preprint arXiv:2502.11462.
  38. 38.Liangtai Sun, Yang Han, Zihan Zhao, Da Ma, Zhennan Shen, Baocai Chen, Lu Chen, and Kai Yu. 2024. SciEval: A multi-level large language model evaluation benchmark for scientific research. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 19053–19061.
  39. 39.Paul A. Tipler. 1999. Physics for Scientists and Engineers, 4th edition. W. H. Freeman and Company.
  40. 40.Gladys Tyen, Hassan Mansoor, Victor Carbune, Peter Chen, and Tony Mak. 2024. LLMs cannot find reasoning errors, but can correct them given the error location. In Findings of the Association for Computational Linguistics: ACL 2024.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  42. 42.Huimin Xu, Xin Mao, Feng-Lin Li, Xiaobao Wu, Wang Chen, Wei Zhang, and Anh Tuan Luu. 2025. Full-step-DPO: Self-supervised preference optimization with step-wise rewards for mathematical reasoning. arXiv preprint arXiv:2502.14356.
  43. 43.Yunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee, and Lu Wang. 2024. Small language models need strong verifiers to self-correct reasoning. arXiv preprint arXiv:2404.17140.
  44. 44.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.

Citation

MLA
Jaiswal, R., et al. “Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models”. arXiv, 2026, https://doi.org/10.48550/arxiv.2607.05199.
APA
Jaiswal, R., Jain, D., Dhawan, R., Uppalapati, S. K., Satoh, S., Ganu, T., & Shah, R. R. (2026). Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models. arXiv. https://doi.org/10.48550/arxiv.2607.05199
Chicago
Jaiswal, R., D. Jain, R. Dhawan, et al. 2026. “Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2607.05199.
Harvard
Jaiswal, R. et al. (2026) “Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models”. arXiv. Available at: https://doi.org/10.48550/arxiv.2607.05199.
Vancouver
1. Jaiswal R, Jain D, Dhawan R, Uppalapati SK, Satoh S, Ganu T, Shah RR (2026) Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models. https://doi.org/10.48550/arxiv.2607.05199

BibTeX

@misc{https://doi.org/10.48550/arxiv.2607.05199,
  doi = {10.48550/ARXIV.2607.05199},
  url = {https://arxiv.org/abs/2607.05199},
  author = {Jaiswal, Raj and Jain, Dhruv and Dhawan, Rishabh and Uppalapati, Sree Krishna and Satoh, Shin'ichi and Ganu, Tanuja and Shah, Rajiv Ratn},
  keywords = {Artificial Intelligence (cs.AI), FOS: Computer and information sciences},
  title = {Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/