A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science

Clayton CohnNicole HutchinsTuan LeGautam Biswas

article2024AAAI91 citations

Develops a human-in-the-loop framework combining GPT-4 with chain-of-thought prompting and active learning to automatically grade open-ended middle school science assessments while generating actionable explanatory feedback.

Listen

Assessing open-ended student responses in science education is essential for tracking conceptual understanding, but frequent manual evaluation places a substantial time burden on teachers and often leads to inconsistent scoring and delayed feedback. Existing automated evaluation systems primarily handle formulaic math or computer science tasks, leaving a gap for free-form, reasoning-focused science assessments where datasets are typically small and imbalanced. The article evaluates a human-in-the-loop framework that combines GPT-4 with few-shot in-context learning, chain-of-thought prompting, and active learning to automatically score and generate explanatory feedback for middle school science assessments.

The researchers analyzed responses from 270 public middle school students participating in a three-week Earth Science curriculum focused on water runoff and the conservation of matter. The evaluation examined three open-ended assessment questions designed under evidence-centered principles to evaluate both core scientific concepts and scientific reasoning. Human coders scored an initial subset of responses to achieve inter-rater reliability, incorporating areas of consensus and disagreement into the prompt designs. The methodology then iteratively analyzed errors across an 80% training/validation split to target and correct recurring model reasoning flaws before evaluating final performance on a 20% held-out test set.

The analysis produced several key findings. First, the human-in-the-loop prompting pipeline achieved strong alignment with human graders, reaching Quadratic Weighted Kappa agreement scores of 0.80 or higher across 9 of the 11 evaluated subscores and total scores, with 4 subscores exceeding 0.90. Second, the model reached high overall predictive balance, recording a Macro F1 score of at least 0.90 on 10 out of 11 scoring categories. Third, the framework successfully produced clear, evidence-based explanations connecting student quotes directly to grading rubrics, supporting actionable feedback generation. However, the study also revealed that excessive chain-of-thought granularity and active learning prompts created occasional overfitting, particularly on simpler factual concepts and highly ambiguous reasoning items.

These findings demonstrate that modern language models can significantly reduce educator workload while generating transparent, rubric-aligned feedback that supports classroom learning goals. The results also show that automated evaluation patterns correlate with human grading difficulty; questions requiring multiple human review cycles to reach consensus posed similar classification challenges for the artificial intelligence. Consequently, discrepancy reviews between models and human scorers provide a valuable diagnostic tool for educational leaders to identify ambiguous test prompts and refine assessment rubrics.

Educational organizations considering automated evaluation tools should pilot human-in-the-loop workflows rather than relying on fully autonomous grading. Decision-makers should prioritize using model-generated explanations for formative guidance rather than assigning high-stakes summative scores, while also using error logs to iteratively revise assessment materials. Stakeholders should note that these findings are constrained by a single middle school Earth Science context and a limited sample size, and broad deployments must account for potential model hallucinations, bias, and student privacy risks.

Cover for A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science

Abstract

This paper explores the use of large language models (LLMs) to score and explain short-answer assessments in K-12 science. While existing methods can score more structured math and computer science assessments, they often do not provide explanations for the scores. Our study focuses on employing GPT-4 for automated assessment in middle school Earth Science, combining few-shot and active learning with chain-of-thought reasoning. Using a human-in-the-loop approach, we successfully score and provide meaningful explanations for formative assessment responses. A systematic analysis of our method’s pros and cons sheds light on the potential for human-in-the-loop techniques to enhance automated grading for open-ended science assessments.

Table of Contents

  • Introduction
  • Background
  • Methods
  • Curricular context
  • Study Design and Dataset
  • Model
  • Approach
  • Results
  • Comparing Model and Human Performance
  • Conclusion and Future Implications
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Human-in-the-loop GPT-4 assessment pipeline

    model/method

    The paper’s contribution is a human-in-the-loop pipeline that uses GPT-4 to score and explain middle-school Earth Science short answers. The pipeline combines four elements: a rubric specifying conceptual and scientific-reasoning subscores; few-shot in-context examples; chain-of-thought (CoT) explanations linking student evidence to rubric criteria; and active learning in which a researcher examines model errors and adds corrected examples to the prompt. After prompt refinement, GPT-4 produces both rubric-aligned scores and evidence-based explanations intended to support student feedback and teacher analysis. The workflow is illustrated by the process diagram on page 4, which connects formative-assessment and rubric design to response scoring, prompt development, active learning, and eventual assessment scores and feedback.

  2. Knowl 2 — Earth Science assessment context and scoring structure

    experimental setup

    The study uses formative-assessment data from the three-week, NGSS-aligned SPICE middle-school Earth Science curriculum, in which students redesign a schoolyard to reduce post-storm water runoff under cost and accessibility constraints. The data came from two Vanderbilt University-approved studies involving 270 students at a public middle school in the southeastern United States; non-consenting participants were removed, and some responses were unavailable because of absences or incomplete submissions. The assessment diagram shown on page 2 depicts rainfall from clouds, absorption, runoff, and arrows whose direction and size represent properties of water movement.

    Three questions required students to inspect that diagram. Question 1 awarded one conceptual point for recognizing that arrow size represents the quantity of water. Question 2 awarded four points: a conceptual and a reasoning point for Arrow Direction, and a conceptual and a reasoning point for Arrow Size. Correct Arrow Direction examples identified water arriving from the sky as rain, being absorbed, or becoming runoff; correct Arrow Size responses recognized that arrow size represents water amount. Question 3 also awarded four points: a conceptual and a reasoning point for Runoff Direction, based on identifying that the runoff arrow pointed uphill, and a conceptual and a reasoning point for Arrow Size, based on identifying that arrow sizes should obey conservation of matter.

  3. Knowl 3 — Human consensus labels and prompt-development split

    experimental setup

    For each of the three questions, two authors independently scored a randomly selected 20% of responses with the question-specific rubric. They discussed disagreements until reaching consensus, and the consensus labels became the reference labels for model alignment. Inter-rater reliability was repeated until Cohen’s κ>0.7\kappa>0.7 for every subscore of every question; Question 1 required two rounds, Question 2 required three rounds, and Question 3 required one round. One author then manually scored the complete response set, so every response had a human score during evaluation.

    Before prompt development, responses for each question were divided into 80% training and 20% held-out testing data. The training portion supplied both the few-shot prompt examples and a validation set for active learning; the validation-to-training ratio was approximately 43:1. The validation set consisted of training responses not used as initial few-shot examples. The paper therefore evaluated GPT-4 against human-consensus labels on held-out responses while retaining manual grading of the full dataset.

  4. Knowl 4 — Rubric-grounded chain-of-thought prompt construction

    model/method

    GPT-4 was prompted to act as a middle-school teacher evaluating formative-assessment responses. Each question prompt contained the assessment question, its scoring rubric, and an output format designed for readable and programmatically parsable scores. The prompt first included consensus-labeled examples containing both scores and chain-of-thought explanations.

    For each labeled example, the explanation followed the template: evidence from the student response, a reference to the applicable rubric criterion, and the resulting score. In practice, the prompt used quoted student text in a pattern such as: the student says a specified statement; the rubric requires a specified property; therefore the student earns or does not earn a specified point. Examples on which human scorers initially disagreed were also included, with explanations encoding the final consensus decision. The researchers attempted to balance positive and negative examples for every subscore; for the four-subscore questions, perfect balance was sometimes impossible, but each prompt contained at least one positive and one negative example for every subscore.

  5. Knowl 5 — Error-directed active learning for prompt refinement

    algorithm

    The active-learning procedure uses a researcher as the labeling oracle and modifies the GPT-4 prompt rather than updating model parameters. It was applied once for each assessment question.

    Input: Initial few-shot prompt P, validation responses V, human-consensus labels, and question rubric
    Output: Refined prompt P'
    Run GPT-4 on every response in V using P.
    Compare each predicted subscore with its human-consensus label.
    For each subscore, count false positives and false negatives and inspect the model's explanations.
    Identify the recurring reasoning error responsible for the largest number of wrong predictions.
    Select a validation response that exemplifies that error and preserves balance across subscores.
    Write a corrected chain-of-thought explanation linking the response evidence to the rubric and consensus score.
    Add the corrected example to P and rebalance the few-shot examples when possible.
    Repeat until the selected active-learning iteration is complete or a general stopping condition is met.
    Return P'.

    The general stopping conditions considered were convergence with no incorrect validation scores, overfitting indicated by more errors than in the previous iteration, or an insufficient number of remaining validation examples to maintain acceptable prompt balance. In the reported experiment, Question 1 added all of its few incorrectly predicted validation cases, while Questions 2 and 3 selected the minimum number of cases needed to address all observed reasoning errors while preserving balance. The final prompts contained 12 examples for Question 1, 10 for Question 2, and 9 for Question 3, compared with initial few-shot prompts containing 4, 5, and 5 examples respectively.

  6. Knowl 6 — Incremental evaluation of prompting components

    experimental setup

    GPT-4 was compared under four prompt configurations on held-out test responses: (1) Zero-Shot, containing the rubric but no labeled examples; (2) Few-Shot, containing labeled examples with numerical scores but no reasoning; (3) Few-Shot, CoT, containing the same type of labeled examples with rubric-grounded reasoning; and (4) CoT + AL, the complete method with active-learning examples and corrected reasoning. The comparison isolates the effects of demonstrations, explicit reasoning, and error-directed prompt refinement.

    Performance was measured using accuracy, Macro F1, and Cohen’s Quadratic Weighted Kappa (QWK). Macro F1 was emphasized because the subscore labels were imbalanced, especially for scientific reasoning. QWK was emphasized because it accounts for the degree of disagreement between ordinal scores. Accuracy was reported for reference but was not used as the principal comparison metric.

  7. Knowl 7 — Question 1 Arrow Size performance

    data/table

    Question 1 required a binary conceptual judgment: whether a student recognized that arrow size in the water-runoff diagram represents water quantity. The held-out-test results reported on page 5 are:

    Could not parse LaTeX table

    Here, nn is the number of labeled prompt examples. GPT-4 already reached moderate agreement in the Zero-Shot condition, achieved perfect test performance with four ordinary few-shot examples, and declined when initially supplied with chain-of-thought explanations. Active learning corrected most of that decline, producing strong agreement with the human labels, although it did not recover the perfect Few-Shot result.

  8. Knowl 8 — Question 2 performance and ambiguity effects

    data/table

    Question 2 asked students to identify two features that the conceptual model represented well. It contained four binary subscores—Arrow Direction, Arrow Direction Reasoning, Arrow Size, and Arrow Size Reasoning—and a four-point total score. The held-out-test results reported on page 5 are:

    Could not parse LaTeX table

    The complete method was best or tied for best on the two conceptual subscores and yielded the highest total-score accuracy and QWK, with total QWK of 0.87. The scientific-reasoning subscores did not improve consistently and the final Arrow Direction Reasoning score was particularly weak, with QWK of 0.30. Question 2 was also the hardest question for human scorers to agree on. Its open-ended wording allowed multiple correct phenomena, vague responses, and responses mixing correct and incorrect statements. A response saying that an absorption arrow represented water amount illustrates the problem: although it may show understanding of arrow size, the absorption arrow itself was scientifically incorrect because it exceeded the rainfall arrow and violated conservation of matter. Active-learning corrections for such cases caused the model to overfit other cases, leading the researchers to conclude that both the question wording and rubric required revision.

  9. Knowl 9 — Question 3 performance and correction of evidence reuse

    data/table

    Question 3 asked students to identify two erroneous features of the conceptual model and explain what should be changed. It contained four binary subscores—Runoff Direction, Runoff Direction Reasoning, Arrow Size, and Arrow Size Reasoning—and a four-point total score. The held-out-test results reported on page 6 are:

    Could not parse LaTeX table

    Few-shot examples improved every Question 3 subscore on both accuracy and QWK, except that total-score Macro F1 decreased. Adding CoT improved both Runoff Direction measures but initially harmed both Arrow Size measures. Active learning restored or improved the Arrow Size measures and produced the best Runoff Direction Reasoning result, with QWK of 0.94; however, it overfit the Runoff Direction conceptual subscore, reducing its QWK from 0.84 with Few-Shot, CoT to 0.75. The principal corrected error was evidence reuse: GPT-4 sometimes used the same student statement to award separate conceptual and reasoning points. Adding a corrected CoT example largely reduced this overscoring.

  10. Knowl 10 — Overall agreement and observed model errors

    empirical result

    Across the three questions, 9 of the 11 reported subscores and total scores reached strong or better human agreement at some point in the four-condition evaluation, defined by QWK ≥0.8\geq 0.8; four reached almost-perfect agreement with QWK >0.9>0.9. Every score except Question 2 Arrow Direction Reasoning reached Macro F1 of at least 0.90 in at least one condition. The results therefore support the feasibility of rubric-conditioned GPT-4 scoring for these short-answer science assessments, but they do not show that adding CoT or active learning always improves performance.

    Qualitative review on page 7 found that the model and human scorer disagreed on 22 cases, and the lead author judged the model’s score preferable in 3 of those cases—one conceptual and two reasoning disagreements. The model sometimes overfit to surface cues: it treated words such as “because” as evidence of reasoning, treated the phrase “arrow size” as sufficient even when the scientific attribution was wrong, and failed to recognize “run off” and “runoff” as equivalent. It also sometimes cited faulty student logic as justification for a positive score, reused evidence across subscores, or produced a positive score without the rubric-linked evidence expected from the prompt. Human agreement appeared to provide a qualitative indication of task difficulty: questions that were easier for human scorers to reconcile were generally easier for GPT-4 to align with, whereas the ambiguous Question 2 was difficult for both.

  11. Knowl 11 — Limitations and deployment risks

    limitation

    The study’s own limitations include the small, imbalanced educational dataset; token and labeling costs that restrict the number of prompt examples; and overfitting caused by highly targeted CoT and active-learning examples. The authors did not update the rubric during active learning, did not quantitatively establish the relationship between human inter-rater agreement and model performance, and evaluated the approach with manually scored data rather than replacing human grading in classroom deployment. They also note that simpler subproblems may not justify an LLM because rule-based methods can be preferable.

    The authors identify broader deployment risks involving student privacy, bias, hallucination, and high-stakes use with children. They also caution that a generated CoT explanation may not faithfully represent the reasoning that caused GPT-4’s decision, so the explanation should not automatically be treated as a trustworthy account of the model’s internal process. Finally, the qualitative teacher feedback reported by the authors suggests that rubric-linked explanations may be more useful as guidance about students’ next learning steps than as standalone performance scores, but this proposed use was not experimentally evaluated.

Coverage note — No substantial contributed component was omitted; bibliographic background, related work, acknowledgements, and the paper’s proposed future teacher-partnership work were excluded because they do not add independently evaluated method or results.

References

  1. 1.Adair, A.; Pedro, M. S.; Gobert, J.; and Segan, E. 2023. Real-Time AI-Driven Assessment and Scaffolding that Improves Students’ Mathematical Modeling during Science Investigations. In Wang, N.; Rebolledo-Mendez, G.; Matsuda, N.; Santos, O. C.; and Dimitrova, V., eds., Artificial Intelligence in Education, 202–216. Cham: Springer Nature Switzerland. ISBN 978-3-031-36272-9.
  2. 2.Bloom, B.; Madaus, G.; and Hastings, J. 1971. Handbook on Formative and Summative Evaluation of Student Learning. New York: McGraw-Hill.
  3. 3.Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020. Language Models are Few-Shot Learners. arXiv e-prints, arXiv:2005.14165.
  4. 4.Charmaz, K. 2006. Constructing grounded theory: A practical guide through qualitative analysis. Sage.
  5. 5.Chiu, J.; McElhaney, K.; Zhang, N.; Biswas, G.; Fried, R.; Basu, S.; Alozie, N.; and Hong, J. 2019. A Principled Approach to NGSS-aligned Curriculum Development Integrating Science, Engineering, and Computation: A Pilot Study. In NARST Annual International Conference. NARST.
  6. 6.Cizek, G. J.; and Lim, S. N. 2023. Formative assessment: an overview of history, theory and application. In Tierney, R. J.; Rizvi, F.; and Ercikan, K., eds., International Encyclopedia of Education (Fourth Edition), 1–9. Oxford: Elsevier, fourth edition edition. ISBN 978-0-12-818629-9.
  7. 7.Cochran, K.; Cohn, C.; and Hastings, P. 2023. Improving NLP model performance on small educational data sets using self-augmentation. In Proceedings of the 15th International Conference on Computer Supported Education (2023, to appear).
  8. 8.Cochran, K.; Cohn, C.; Hutchins, N.; Biswas, G.; and Hastings, P. 2022. Improving automated evaluation of formative assessments with text data augmentation. In International Conference on Artificial Intelligence in Education, 390–401. Springer.
  9. 9.Cochran, K.; Cohn, C.; Rouet, J. F.; and Hastings, P. 2023. Improving Automated Evaluation of Student Text Responses Using GPT-3.5 for Text Data Augmentation. In International Conference on Artificial Intelligence in Education, 217–228. Springer.
  10. 10.Cohen, J. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement, 20(1): 37–46.
  11. 11.Cohen, J. 1968. Weighted kappa: nominal scale agreement provision for scaled disagreement or partial credit. Psychological bulletin, 70(4): 213.
  12. 12.Cohn, C. 2020. BERT efficacy on scientific and medical datasets: a systematic literature review. DePaul University.
  13. 13.Funayama, H.; Asazuma, Y.; Matsubayashi, Y.; Mizumoto, T.; and Inui, K. 2023. Reducing the Cost: Cross-Prompt Pre-finetuning for Short Answer Scoring. In International Conference on Artificial Intelligence in Education, 78–89. Springer.
  14. 14.Harris, C.; Wiebe, E.; Grover, S.; and Pellegrino, J. 2023. Classroom-based STEM assessment: Contemporary issues and perspectives. Community for Advancing Discovery Research in Education (CADRE). Education Development Center, Inc.
  15. 15.Hatch, J. A. 2002. Doing qualitative research in education settings. SUNY Press.
  16. 16.Haudek, K. C.; Kaplan, J. J.; Knight, J.; Long, T.; Merrill, J.; Munn, A.; Nehm, R.; Smith, M.; and Urban-Lurain, M. 2011. Harnessing technology to improve formative assessment of student conceptions in STEM: forging a national network. CBE—Life Sciences Education, 10(2): 149–155.
  17. 17.Hutchins, N.; and Biswas, G. 2023. Using Teacher Dashboards to Customize Lesson Plans for a Problem-Based, Middle School STEM Curriculum. In LAK23: 13th International Learning Analytics and Knowledge Conference, LAK2023, 324–332. New York, NY, USA: Association for Computing Machinery. ISBN 9781450398657.
  18. 18.Hutchins, N. M.; Basu, S.; McElhaney, K.; Chiu, J.; Fick, S.; Zhang, N.; and Biswas, G. 2021. Coherence across conceptual and computational representations of students’ scientific models. In The International Society of the Learning Sciences Annual Meeting 2021. International Society of the Learning Sciences (ISLS).
  19. 19.Liu, E.; Stephan, M.; Nie, A.; Piech, C.; Brunskill, E.; and Finn, C. 2022. Giving Feedback on Interactive Student Programs with Meta-Exploration. Advances in Neural Information Processing Systems, 35: 36282–36294.
  20. 20.Liu, O. L.; Rios, J. A.; Heilman, M.; Gerard, L.; and Linn, M. C. 2016. Validation of automated scoring of science assessments. Journal of Research in Science Teaching, 53(2): 215–233.
  21. 21.Mao, L.; Liu, O. L.; Roohr, K.; Belur, V.; Mulholland, M.; Lee, H.-S.; and Pallant, A. 2018. Validation of automated scoring for a formative assessment that employs scientific argumentation. Educational Assessment, 23(2): 121–138.
  22. 22.McHugh, M. L. 2012. Interrater reliability: the kappa statistic. Biochemia medica, 22(3): 276–282.
  23. 23.Min, S.; Lyu, X.; Holtzman, A.; Artetxe, M.; Lewis, M.; Hajishirzi, H.; and Zettlemoyer, L. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  24. 24.Mislevy, R. J.; and Haertel, G. D. 2006. Implications of Evidence-Centered Design for Educational Testing. Educational Measurement: Issues and Practice, 25(4): 6–20.
  25. 25.Moore, S.; Nguyen, H. A.; Chen, T.; and Stamper, J. 2023. Assessing the Quality of Multiple-Choice Questions Using GPT-4 and Rule-Based Methods. In European Conference on Technology Enhanced Learning, 229–245. Springer.
  26. 26.Morris, W.; Crossley, S.; Holmes, L.; Ou, C.; McNamara, D.; and Dascalu, M. 2023. Using Large Language Models to Provide Formative Feedback in Intelligent Textbooks. In International Conference on Artificial Intelligence in Education, 484–489. Springer.
  27. 27.Mosbach, M.; Andriushchenko, M.; and Klakow, D. 2020. On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines. arXiv preprint arXiv:2006.04884.
  28. 28.NGSS. 2013. Next Generation Science Standards: For States, By States. The National Academies Press.
  29. 29.OpenAI. 2023. GPT-4 Technical Report. arXiv e-prints, arXiv:2303.08774.
  30. 30.Ren, P.; Xiao, Y.; Chang, X.; Huang, P.-Y.; Li, Z.; Gupta, B. B.; Chen, X.; and Wang, X. 2021. A survey of deep active learning. ACM computing surveys (CSUR), 54(9): 1–40.
  31. 31.Rodrigues, F.; and Oliveira, P. 2014. A system for formative assessment and monitoring of students’ progress. Computers & Education, 76: 30–41.
  32. 32.Singh, S.; Pupneja, A.; Mital, S.; Shah, C.; Bawkar, M.; Gupta, L. P.; Kumar, A.; Kumar, Y.; Gupta, R.; and Shah, R. R. 2023. H-AES: Towards Automated Essay Scoring for Hindi. arXiv preprint arXiv:2302.14635.
  33. 33.Singla, Y. K.; Krishna, S.; Shah, R. R.; and Chen, C. 2022. Using sampling to estimate and improve performance of automated scoring systems with guarantees. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36 (11), 12835–12843.
  34. 34.Tan, W.; Lin, J.; Lang, D.; Chen, G.; Gasevič, D.; Du, L.; and Buntine, W. 2023. Does informativeness matter? Active learning for educational dialogue act classification. In International Conference on Artificial Intelligence in Education, 176–188. Springer.
  35. 35.Turpin, M.; Michael, J.; Perez, E.; and Bowman, S. R. 2023. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. arXiv preprint arXiv:2305.04388.
  36. 36.Walkoe, J.; Wilkerson, M.; and Elby, A. 2017. Technology-Mediated Teacher Noticing: A Goal for Classroom Practice, Tool Design, and Professional Development. In Proceedings of the 12th International Conference on Computer Supported Collaborative Learning (CSCL) 2017. International Society of the Learning Sciences.
  37. 37.Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; and Zhou, D. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv e-prints, arXiv:2201.11903.
  38. 38.White, J.; Fu, Q.; Hays, S.; Sandborn, M.; Olea, C.; Gilbert, H.; Elnashar, A.; Spencer-Smith, J.; and Schmidt, D. C. 2023. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382.
  39. 39.Wiley, K. J.; Dimitriadis, Y.; Bradford, A.; and Linn, M. C. 2020. From Theory to Action: Developing and Evaluating Learning Analytics for Learning Design. In Proceedings of the Tenth International Conference on Learning Analytics & Knowledge, LAK ’20, 569–578. New York, NY, USA: Association for Computing Machinery. ISBN 9781450377126.
  40. 40.Wilson, J.; Ahrendt, C.; Fudge, E. A.; Raiche, A.; Beard, G.; and MacArthur, C. 2021. Elementary teachers’ perceptions of automated feedback and automated scoring: Transforming the teaching and learning of writing using automated writing evaluation. Computers & Education, 168: 104208.
  41. 41.Wu, X.; He, X.; Liu, T.; Liu, N.; and Zhai, X. 2023. Matching exemplar as next sentence prediction (mensp): Zero-shot prompt learning for automatic scoring in science education. In International Conference on Artificial Intelligence in Education, 401–413. Springer.
  42. 42.Zeng, Z.; Li, L.; Guan, Q.; Gasevič, D.; and Chen, G. 2023. Generalizable Automatic Short Answer Scoring via Prototypical Neural Network. In International Conference on Artificial Intelligence in Education, 438–449. Springer.
  43. 43.Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223.
  44. 44.Zhuo, T. Y.; Huang, Y.; Chen, C.; and Xing, Z. 2023. Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity. arXiv preprint arXiv:2301.12867, 12–2.

Citation

MLA
Cohn, C., et al. “A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in Science”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 21, 2024, pp. 23182–90, https://doi.org/10.1609/aaai.v38i21.30364.
APA
Cohn, C., Hutchins, N., Le, T., & Biswas, G. (2024). A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in Science. Proceedings of the AAAI Conference on Artificial Intelligence, 38(21), 23182–23190. https://doi.org/10.1609/aaai.v38i21.30364
Chicago
Cohn, C., N. Hutchins, T. Le, and G. Biswas. 2024. “A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in Science”. Proceedings of the AAAI Conference on Artificial Intelligence 38 (21): 23182–90. https://doi.org/10.1609/aaai.v38i21.30364.
Harvard
Cohn, C. et al. (2024) “A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in Science”, Proceedings of the AAAI Conference on Artificial Intelligence, 38(21), pp. 23182–23190. Available at: https://doi.org/10.1609/aaai.v38i21.30364.
Vancouver
1. Cohn C, Hutchins N, Le T, Biswas G (2024) A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in Science. Proceedings of the AAAI Conference on Artificial Intelligence 38:23182–23190

BibTeX

@article{Cohn_2024, title={A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in Science}, volume={38}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/aaai.v38i21.30364}, DOI={10.1609/aaai.v38i21.30364}, number={21}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Cohn, Clayton and Hutchins, Nicole and Le, Tuan and Biswas, Gautam}, year={2024}, month=Mar, pages={23182–23190} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF