Hybrid AI-Based Automated Assessment Model For Long Answer Grading
Keywords:
Automated Assessment, Ensemble Learning, Natural Language Processing, Educational Technology, Semantic Similarity, BERT.Abstract
There are still many difficulties that need to be addressed in automatically grading the long answer responses. Even though various solutions have been offered, most of them fail in balancing between accuracy and the interpretation of the intended meaning by the students as well as the applicability in practice. An ensemble method consisting of five different models, namely TF-IDF, SBERT, BERTScore, Keyword Coverage, and Entity Matching, is suggested in this project. One of the main components of this method is the adaptively weighted models since instead of using a common strategy regardless of question type, a new one is employed depending on each category of questions. Therefore, the adaptively weighted models eliminate the necessity of pseudo labeling in other models, which can lead to erroneous grades. Testing with 1,000 responses on 20 questions provided accuracy of 89% and RMSE of 0.420. Accuracy per question varied between 86% and 92%. Moreover, the method used here is practical for implementation in classrooms since no training data is needed.





