An AI-Powered Sentence-BERT Framework for Automated Assessment of Open-Ended Responses: A Multi-Model Evaluation and Response-Length Bias Analysis

Authors

DOI:

https://doi.org/10.62411/faith.3048-3719-427

Keywords:

AI-Based Assessment, all-MiniLM-L12-v2, Educational Technology, Semantic Text Similarity, Sentence-BERT, Short Answer Grading, Transformers

Abstract

The automated assessment of open-ended responses remains challenging due to the complexity and subjective nature of human grading. A further methodological challenge is determining which pretrained Sentence-BERT model and scoring configuration can provide reliable agreement with human assessment while maintaining computational efficiency. This study presents an AI-powered framework based on Sentence-BERT models for automated grading of short-answer responses, focusing on model performance, reliability, response-length sensitivity, and hybrid scoring. A custom dataset of student responses to open-ended questions, annotated by three independent human raters, was developed and pre-processed. Five Sentence-BERT models, all-MiniLM-L12-v2, all-mpnet-base-v2, stsb-roberta-large, paraphrase-MiniLM-L12-v2, and distilbert-base-nli-stsb-mean-tokens, were evaluated using semantic similarity scoring and multiple metrics, including Pearson and Spearman correlations, Cohen’s Kappa, MAE, RMSE, and ICC. Among the evaluated models, all-MiniLM-L12-v2 provided the most favorable balance of grading performance, reliability (ICC = 0.527), and computational efficiency. Human ratings showed stronger dependence on response length (Pearson r = 0.663) than the evaluated AI models (r = 0.338–0.474), indicating reduced sensitivity to verbosity. Weight-sensitivity analysis showed that the 25% semantic–75% rubric configuration had the lowest continuous-score MAE (1.8499) and RMSE (2.3980) among the tested configurations, supporting the complementary contribution of holistic semantic similarity and rubric-based concept coverage. The selected model and hybrid scoring mechanism were implemented in a web-based prototype providing automated scores and interpretable feedback while retaining instructor review.

Downloads

Download data is not yet available.

Author Biographies

Temidayo Oluwatosin Omotehinwa, Federal University of Health Sciences

Department of Mathematics and Computer Science, Federal University of Health Sciences, Otukpo 972261, Nigeria

Folashade Habibat Omotehinwa, Federal University of Health Sciences

Department of Chemistry, Federal University of Health Sciences, Otukpo 972261, Nigeria

Adeyi Thomas Edeh, Federal University of Health Sciences

Department of Mathematics and Computer Science, Federal University of Health Sciences, Otukpo 972261, Nigeria

Yakubu Aliyu Ibrahim, Federal University of Health Sciences

Department of Mathematics and Computer Science, Federal University of Health Sciences, Otukpo 972261, Nigeria

References

L. Rodrigues, C. Xavier, N. Costa, D. Gasevic, and R. F. Mello, “Is GPT-4 fair? An empirical analysis in automatic short answer grading,” Comput. Educ. Artif. Intell., vol. 8, no. June, p. 100428, Jun. 2025, doi: 10.1016/j.caeai.2025.100428.

L. Zhang, Y. Huang, X. Yang, S. Yu, and F. Zhuang, “An automatic short-answer grading model for semi-open-ended questions,” Interact. Learn. Environ., vol. 30, no. 1, pp. 177–190, Jan. 2022, doi: 10.1080/10494820.2019.1648300.

F. Morley, E. Walland, and C. V. Rodeiro, “Auto-marking short answer questions in science: The foundational years of transformer-based models from BERT to GPT-4,” Int. J. Artif. Intell. Educ., vol. 36, no. 1–2, p. 100005, Mar. 2026, doi: 10.1016/j.ijaied.2026.100005.

S. A. Mahmood and M. A. Abdulsamad, “Automatic assessment of short answer questions: Review,” Edelweiss Appl. Sci. Technol., vol. 8, no. 6, pp. 9158–9176, Dec. 2024, doi: 10.55214/25768484.v8i6.3956.

F. Rodrigues and P. Oliveira, “A system for formative assessment and monitoring of students’ progress,” Comput. Educ., vol. 76, pp. 30–41, Jul. 2014, doi: 10.1016/j.compedu.2014.03.001.

S. Tobler, “Smart grading: A generative AI-based tool for knowledge-grounded answer evaluation in educational assessments,” MethodsX, vol. 12, p. 102531, Jun. 2024, doi: 10.1016/j.mex.2023.102531.

A. Kashi, S. Shastri, A. R. Deshpande, J. Doreswamy, and G. Srinivasa, “A Score Recommendation System Towards Automating Assessment In Professional Courses,” in 2016 IEEE Eighth International Conference on Technology for Education (T4E), Dec. 2016, pp. 140–143. doi: 10.1109/T4E.2016.036.

R. Weegar and P. Idestam-Almquist, “Reducing Workload in Short Answer Grading Using Machine Learning,” Int. J. Artif. Intell. Educ., vol. 34, no. 2, pp. 247–273, Jun. 2024, doi: 10.1007/s40593-022-00322-1.

J. Pecuchova, Ľ. Benko, and M. Drlik, “Automated Grading of Open-Ended Questions in Higher Education Using GenAI Models,” Int. J. Artif. Intell. Educ., vol. 35, no. 6, pp. 3813–3846, Dec. 2025, doi: 10.1007/s40593-025-00517-2.

E. del Gobbo, A. Guarino, B. Cafarelli, and L. Grilli, “GradeAid: a framework for automatic short answers grading in educational contexts—design, implementation and evaluation,” Knowl. Inf. Syst., vol. 65, no. 10, pp. 4295–4334, Oct. 2023, doi: 10.1007/s10115-023-01892-9.

T. O. Omotehinwa, R. A. Ezekiel, M. Onoja, E. C. Omeye, and Y. A. Ibrahim, “Comparative Evaluation of Transformer-Based Models for Automated Short-Answer Grading,” NIPES J. Sci. Technol. Res., vol. 7, no. 1, pp. 843–849, Oct. 2025, doi: 10.37933/nipes/7.4.2025.SI97.

S. Bonthu, S. R. Sree, and M. H. M. Krishna Prasad, “Framework for automation of short answer grading based on domain-specific pre-training,” Eng. Appl. Artif. Intell., vol. 137, p. 109163, Nov. 2024, doi: 10.1016/j.engappai.2024.109163.

J. Garg, J. Papreja, K. Apurva, and G. Jain, “Domain-Specific Hybrid BERT based System for Automatic Short Answer Grading,” in 2022 2nd International Conference on Intelligent Technologies (CONIT), Jun. 2022, pp. 1–6. doi: 10.1109/CONIT55038.2022.9847754.

A. Amalia, M. S. Lydia, M. A. Muchtar, F. Y. Manik, Sinu, and D. Gunawan, “Mitigating Bias and Assessment Inconsistencies with BERT-Based Automated Short Answer Grading for the Indonesian Language,” IAENG Int. J. Comput. Sci., vol. 52, no. 3, pp. 533–545, 2025.

S. Konade, Y. Hirgude, A. Kulkarni, A. Jadhav, and L. Sonar, “Implementation of an Automated Answer Evaluation System,” in 2024 IEEE International Conference on Computing, Power and Communication Technologies (IC2PCT), Feb. 2024, pp. 457–462. doi: 10.1109/IC2PCT60090.2024.10486693.

O. Bolgova, P. Ganguly, M. F. Ikram, and V. Mavrych, “Evaluating large language models as graders of medical short answer questions: a comparative analysis with expert human graders,” Med. Educ. Online, vol. 30, no. 1, Dec. 2025, doi: 10.1080/10872981.2025.2550751.

G. Kuling et al., “Assessment of Short-Answer Questions by ChatGPT in a Medical School Course,” NEJM AI, vol. 3, no. 2, Jan. 2026, doi: 10.1056/AIcs2500239.

M. Kaya and I. Cicekli, “A Hybrid Approach for Automated Short Answer Grading,” IEEE Access, vol. 12, no. June, pp. 96332–96341, 2024, doi: 10.1109/ACCESS.2024.3420890.

J. Osaka, A. Maeda, H. Oka, Y. Mori, T. Ishioka, and H. Suyari, “Reliable and efficient automated short-answer scoring for a large dataset using active learning and deep learning,” Interact. Learn. Environ., vol. 33, no. 6, pp. 3776–3787, Jul. 2025, doi: 10.1080/10494820.2025.2452005.

F. E. A. Hassanein, R. R. Hussein, Y. Ahmed, J. El-Guindy, D. E. Ahmed, and A. Abou-Bakr, “Calibration of AI large language models with human subject matter experts for grading of clinical short-answer responses in dental education,” BMC Oral Health, vol. 26, no. 1, p. 286, Feb. 2026, doi: 10.1186/s12903-026-07665-4.

C. Anghel, E. Pecheanu, A. A. Anghel, M. V. Craciun, and A. Cocu, “ExamQ-Gen: Instructor-in-the-Loop Generation of Self-Contained Exam Questions from Course Materials and Decision-Support Grading,” Computers, vol. 15, no. 3, p. 177, Mar. 2026, doi: 10.3390/computers15030177.

H. M. T. W. Seneviratne and S. S. Manathunga, “Artificial intelligence assisted automated short answer question scoring tool shows high correlation with human examiner markings,” BMC Med. Educ., vol. 25, no. 1, p. 1146, Aug. 2025, doi: 10.1186/s12909-025-07718-2.

Y. Oh Lee, B. Bang, J. Lee, and S. Oh, “Personalized Auto-Grading and Feedback System for Constructive Geometry Tasks Using Large Language Models on an Online Math Platform,” IEEE Access, vol. 14, no. December 2025, pp. 17788–17802, 2026, doi: 10.1109/ACCESS.2026.3657726.

N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 3980–3990. doi: 10.18653/v1/D19-1410.

G. Salton and C. Buckley, “Term-weighting approaches in automatic text retrieval,” Inf. Process. Manag., vol. 24, no. 5, pp. 513–523, Jan. 1988, doi: 10.1016/0306-4573(88)90021-0.

J. Pennington, R. Socher, and C. Manning, “Glove: Global Vectors for Word Representation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2014, pp. 1532–1543. doi: 10.3115/v1/D14-1162.

C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge University Press, 2008. doi: 10.1017/CBO9780511809071.

D. Ortiz Martes, E. Gunderson, C. Neuman, and N. N. Kachouie, “Transformer Models for Paraphrase Detection: A Comprehensive Semantic Similarity Study,” Computers, vol. 14, no. 9, p. 385, Sep. 2025, doi: 10.3390/computers14090385.

H. Gusdevi, A. Setyanto, K. Kusrini, and E. Utami, “Cosine Similarity-Based Evidences Selection for Fact Verification Using SBERT on the FEVER Dataset,” CogITo Smart J., vol. 11, no. 1, pp. 52–66, Jun. 2025, doi: 10.31154/cogito.v11i1.917.52-66.

A. Kumar and H. Kumar, “scaLAR SemEval-2024 Task 1: Semantic Textual Relatednes for English,” in Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), 2024, pp. 902–906. doi: 10.18653/v1/2024.semeval-1.129.

M. Acar Güvendir, A. F. Kılıç, E. Güvendir, and T. Kaçak, “Bridging minds and machines: a comparative study of AI and human rater agreement and reliability in educational assessment,” Educ. Inf. Technol., vol. 31, no. 11, pp. 3781–3804, Jul. 2026, doi: 10.1007/s10639-026-13949-7.

Downloads

Published

2026-09-20

How to Cite

[1]
T. O. Omotehinwa, F. H. Omotehinwa, A. T. Edeh, and Y. A. Ibrahim, “An AI-Powered Sentence-BERT Framework for Automated Assessment of Open-Ended Responses: A Multi-Model Evaluation and Response-Length Bias Analysis”, J. Fut. Artif. Intell. Tech., vol. 3, no. 2, pp. 419–441, Sep. 2026.