A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science
Clayton CohnNicole HutchinsTuan LeGautam Biswas
Develops a human-in-the-loop framework combining GPT-4 with chain-of-thought prompting and active learning to automatically grade open-ended middle school science assessments while generating actionable explanatory feedback.
Assessing open-ended student responses in science education is essential for tracking conceptual understanding, but frequent manual evaluation places a substantial time burden on teachers and often leads to inconsistent scoring and delayed feedback. Existing automated evaluation systems primarily handle formulaic math or computer science tasks, leaving a gap for free-form, reasoning-focused science assessments where datasets are typically small and imbalanced. The article evaluates a human-in-the-loop framework that combines GPT-4 with few-shot in-context learning, chain-of-thought prompting, and active learning to automatically score and generate explanatory feedback for middle school science assessments.
The researchers analyzed responses from 270 public middle school students participating in a three-week Earth Science curriculum focused on water runoff and the conservation of matter. The evaluation examined three open-ended assessment questions designed under evidence-centered principles to evaluate both core scientific concepts and scientific reasoning. Human coders scored an initial subset of responses to achieve inter-rater reliability, incorporating areas of consensus and disagreement into the prompt designs. The methodology then iteratively analyzed errors across an 80% training/validation split to target and correct recurring model reasoning flaws before evaluating final performance on a 20% held-out test set.
The analysis produced several key findings. First, the human-in-the-loop prompting pipeline achieved strong alignment with human graders, reaching Quadratic Weighted Kappa agreement scores of 0.80 or higher across 9 of the 11 evaluated subscores and total scores, with 4 subscores exceeding 0.90. Second, the model reached high overall predictive balance, recording a Macro F1 score of at least 0.90 on 10 out of 11 scoring categories. Third, the framework successfully produced clear, evidence-based explanations connecting student quotes directly to grading rubrics, supporting actionable feedback generation. However, the study also revealed that excessive chain-of-thought granularity and active learning prompts created occasional overfitting, particularly on simpler factual concepts and highly ambiguous reasoning items.
These findings demonstrate that modern language models can significantly reduce educator workload while generating transparent, rubric-aligned feedback that supports classroom learning goals. The results also show that automated evaluation patterns correlate with human grading difficulty; questions requiring multiple human review cycles to reach consensus posed similar classification challenges for the artificial intelligence. Consequently, discrepancy reviews between models and human scorers provide a valuable diagnostic tool for educational leaders to identify ambiguous test prompts and refine assessment rubrics.
Educational organizations considering automated evaluation tools should pilot human-in-the-loop workflows rather than relying on fully autonomous grading. Decision-makers should prioritize using model-generated explanations for formative guidance rather than assigning high-stakes summative scores, while also using error logs to iteratively revise assessment materials. Stakeholders should note that these findings are constrained by a single middle school Earth Science context and a limited sample size, and broad deployments must account for potential model hallucinations, bias, and student privacy risks.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces the foundational chain-of-thought prompting methodology that the source paper applies to formative assessment scoring and explanation generation.
- Paper: Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering, Pan Lu et al. (2022). Demonstrates how chain-of-thought explanations can be leveraged to reason through K-12 science questions, providing direct conceptual grounding for evaluating science responses.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Establishes a framework for using GPT-4 with chain-of-thought reasoning as an automated evaluator aligned with human judgments.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). Analyzes the reliability and alignment of GPT-4 when serving as an automated judge over complex, open-ended textual responses.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Presents methods for improving the robustness and consistency of chain-of-thought reasoning paths in large language models.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). Examines potential unfaithfulness and rationalization issues in chain-of-thought explanations, highlighting key challenges addressed by human-in-the-loop evaluation.
- Paper: T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering, Lei Wang et al. (2024). Extends chain-of-thought evaluation and explanation techniques from text-only science assessments into complex multimodal science question answering.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). Investigates the adversarial vulnerabilities and instruction-following fidelity of LLM evaluators like GPT-4 across diverse judging tasks.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). Provides a comprehensive taxonomy and survey of techniques, rubrics, and challenges for LLM-as-a-judge systems beyond single-domain assessments.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). Develops segmentation and merging strategies to mitigate inherent evaluation biases when employing LLMs as automated judges.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). Proposes evaluation metrics focused on step-by-step validity and redundancy in intermediate reasoning chains rather than just final score outcomes.
