KHUNG ĐÁNH GIÁ CHẤT LƯỢNG CÂU HỎI DỰA TRÊN TRÍ TUỆ NHÂN TẠO

KHUNG ĐÁNH GIÁ CHẤT LƯỢNG CÂU HỎI DỰA TRÊN TRÍ TUỆ NHÂN TẠO

Vũ Xuân Hạnh* hanhvx@hou.edu.vn Trường Đại học Mở Hà Nội 101 Nguyễn Hiền, phường Bạch Mai, Hà Nội, Việt Nam
Đỗ Duy Trinh trinhdd@hou.edu.vn Trường Đại học Mở Hà Nội 101 Nguyễn Hiền, phường Bạch Mai, Hà Nội, Việt Nam
Nguyễn Quang Ánh anhnq@hou.edu.vn Trường Đại học Mở Hà Nội 101 Nguyễn Hiền, phường Bạch Mai, Hà Nội, Việt Nam
Tóm tắt: 
Đánh giá chất lượng câu hỏi là khâu thiết yếu trong đảm bảo giá trị và độ tin cậy của khảo thí, song các phương pháp hiện hành vẫn phụ thuộc nhiều vào chuyên gia, mang tính định tính và khó mở rộng quy mô. Bài viết này đề xuất AF2QE (AI-based Framework for Question Quality Evaluation), một khung tiền kiểm định tích hợp ba tầng nhằm cung cấp các chỉ số chất lượng định lượng, minh bạch và có thể lặp lại. Tầng phân tích ngữ nghĩa sử dụng Sentence-BERT để đo mức độ liên kết của câu hỏi với nội dung giáo trình và chuẩn đầu ra học phần (CLO). Tầng xác thực sư phạm phân loại câu hỏi theo thang Bloom bằng mô hình kết hợp TF-IDF, LinearSVC, tính toán chỉ số bao phủ nội dung và CLO. Tầng kiểm định thống kê ước lượng độ khó và độ phân hóa dựa trên trọng số Bloom và khoảng cách ngữ nghĩa. Thực nghiệm trên 650 câu hỏi thuộc học phần Lập trình Hướng đối tượng cho thấy mô hình phân loại Bloom đạt độ chính xác 92,50% và 84% câu hỏi đồng thời đáp ứng ba tiêu chí về liên kết nội dung, mức nhận thức và vùng độ khó tối ưu. Kết quả bước đầu cho thấy, AF2QE có tiềm năng sàng lọc tiền kiểm định hữu ích, hỗ trợ chuyên gia trong kiểm định chất lượng ngân hàng câu hỏi.
Từ khóa: 
đánh giá chất lượng câu hỏi
tạo sinh câu hỏi tự động
thang phân loại Bloom
tiền kiểm định
chuẩn đầu ra học phần
lí thuyết kiểm thử cổ điển.
Tham khảo: 

[1] Ahmed, A., Kerr, E. & O’Malley, A. (2025). Quality assurance and validity of AI-generated single best answer questions. BMC Medical Education, 25.

[2] Circi, R., Hicks, J. & Sikali, E. (2023). Automatic item generation: foundations and machine learning based approaches for assessments. Frontiers in Education, 8. https://doi.org/10.3389/ feduc.2023.858273.

[3] Gunawan, D., Sembiring, C. A. & Budiman, M. A. (2018). The Implementation of Cosine Similarity to Calculate Text Relevance between Two Documents. 2nd International Conference on Computing and Applied Informatics (ICCAI 2017). DOI: 10.1088/1742-6596/978/1/012120.

[4] He, X., Lin, Z., Gong, Y., Jin, A.-L., Zhang, H., Lin, C., Chen, W. (2024). AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Industry Track, pp.165 190. Mexico City, Mexico.

[5] Hillsley, K. (2025). Question format is the best predictor of item discrimination: a multivariable analysis. Journal of Microbiology & Biology Education, 26(3).

[6] Khojasteh, L., Kafipour, R., Pakdel, F. & Mukundan, J. (2025). Empowering medical students with AI writing co-pilots: design and validation of AI self assessment toolkit. BMC Medical Education, 25. DOI :https://doi.org/10.1186/s12909-025-06753-3.

[7] Kim, Y. H., Kim, B. H., Kim, J., Jung, B. & Bae, S. (2023). Item difficulty index, discrimination index, and reliability of the 26 health professions licensing examinations in 2022, Korea: a psychometric study. Journal of Educational Evaluation for Health Professions, 20. https://doi.org/10.3352/jeehp.2023.20.31.

[8] Makhlouf, K., Amouri, L., Chaabane, N. & EL-Haggar, N. (2020). Exam Questions Classification Based on Bloom’s Taxonomy: Approaches and Techniques. 2020 2nd International Conference on Computer and Information Sciences (ICCIS). Sakaka, Saudi Arabia. DOI: 10.1109/ICCIS49240.2020.9257698.

[9] Moghadamzadeh, A., Salehi, K. & Khodaie, E. (2011). A Comparison the Information Functions of the Item and Test in One, Two and Three Parametric Model of the Item Response Theory (IRT). International Conference on Education and Educational Psychology (ICEEPSY 2011). Istanbul, Turkey.

[10] Mohammed, M. & Omar, N. (2018). Question Classification Based on Bloom’s Taxonomy Using Enhanced TF-IDF. International Journal on Advanced Science, Engineering and Information Technology, 8(4-2), pp.1679–1685. DOI: https://doi. org/10.18517/ijaseit.8.4-2.6835.

[11] Opait, E.-E., Duca, A. & Olariu, M.-E. (2025). Evaluating Generative AI and Human Performance in Question-Answer Validation Tasks. 29th International Conference on Knowledge-Based and Intelligent Information & Engineering Systems (KES 2025), 270, pp.3973-3982. DOI: https://doi. org/10.1016/j.procs.2025.09.522.

[12] Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Hong Kong, China.

[13] Singhania, S., Razniewski, S. & Weikum, G. (2022). Predicting Document Coverage for Relation Extraction. Transactions of the Association for Computational Linguistics (TACL), 10, pp.207-223.

[14] Taheri, E. & Li, N. (2023). L2 Norm-Based Control Regularization for Solving Optimal Control Problems. IEEE Access, 11. DOI: 10.1109/ ACCESS.2023.3331382.

[15] Tam, D. T. & Quynh, P. T. (2024). Alignment between Course Learning Outcomes and Assessments: An Analysis. International Journal of TESOL & Education, 4(2). DOI: https://doi.org/10.54855/ ijte.24422.

[16] Tan, B., Armoush, N., Mazzullo, E., Bulut, O. & Gierl, M. J. (2025). A review of automatic item generation techniques leveraging large language. International Journal of Assessment Tools in Education, 12(2), pp.317-340. https://doi.org/10.21449/ ijate.1602294.

[17] Toraman, Ç., Karadağ, E. & Polat, M. (2022). Validity and reliability evidence for the scale of distance education satisfaction of medical students based on item response theory (IRT). BMC Medical Education, 22. https://doi.org/10.1186/s12909 025-06753-3.

[18] Zhai, X., Haudek, K. C. & Ma, W. (2022). Assessing Argumentation Using Machine Learning and Cognitive Diagnostic Modeling. Research in Science Education, 53, pp.405-424.

Bài viết cùng số