AI-BASED FRAMEWORK FOR QUESTION QUALITY EVALUATION

AI-BASED FRAMEWORK FOR QUESTION QUALITY EVALUATION

Vu Xuan Hanh* hanhvx@hou.edu.vn Hanoi Open University 101 Nguyen Hien street, Bach Mai ward, Hanoi, Vietnam
Do Duy Trinh trinhdd@hou.edu.vn Hanoi Open University 101 Nguyen Hien street, Bach Mai ward, Hanoi, Vietnam
Nguyen Quang Anh anhnq@hou.edu.vn Hanoi Open University 101 Nguyen Hien street, Bach Mai ward, Hanoi, Vietnam
Summary: 
Assessing question quality is an essential step in ensuring the validity and reliability of assessment, yet existing methods remain heavily reliant on experts, are qualitative in nature, and difficult to scale. This paper proposes AF2QE (AI-based Framework for Question Quality Evaluation), a pre-validation framework integrating three tiers to provide quantitative, transparent, and repeatable quality indicators. The semantic analysis tier uses Sentence-BERT to measure the alignment of questions with course content and course learning outcomes (CLO). The pedagogical validation tier classifies questions according to Bloom’s taxonomy using a model that combines TF-IDF and LinearSVC, and calculates content and CLO coverage indices. The statistical validation tier estimates the difficulty and discrimination indices using Bloom weights and semantic distance. Experiments on 650 questions from the Object-Oriented Programming course show that the Bloom classification model achieved an accuracy of 92.50%, and 84% of the questions simultaneously met all three criteria regarding content alignment, cognitive level, and optimal difficulty range. Initial results indicate that AF2QE has potential as a useful pre-screening tool to support experts in validating the quality of question banks.
Keywords: 
Item quality evaluation
automated item generation
Bloom’s taxonomy
pre-deployment validation
course learning outcomes
classical test theory.
Refers: 

[1] Ahmed, A., Kerr, E. & O’Malley, A. (2025). Quality assurance and validity of AI-generated single best answer questions. BMC Medical Education, 25.

[2] Circi, R., Hicks, J. & Sikali, E. (2023). Automatic item generation: foundations and machine learning based approaches for assessments. Frontiers in Education, 8. https://doi.org/10.3389/ feduc.2023.858273.

[3] Gunawan, D., Sembiring, C. A. & Budiman, M. A. (2018). The Implementation of Cosine Similarity to Calculate Text Relevance between Two Documents. 2nd International Conference on Computing and Applied Informatics (ICCAI 2017). DOI: 10.1088/1742-6596/978/1/012120.

[4] He, X., Lin, Z., Gong, Y., Jin, A.-L., Zhang, H., Lin, C., Chen, W. (2024). AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Industry Track, pp.165 190. Mexico City, Mexico.

[5] Hillsley, K. (2025). Question format is the best predictor of item discrimination: a multivariable analysis. Journal of Microbiology & Biology Education, 26(3).

[6] Khojasteh, L., Kafipour, R., Pakdel, F. & Mukundan, J. (2025). Empowering medical students with AI writing co-pilots: design and validation of AI self assessment toolkit. BMC Medical Education, 25. DOI :https://doi.org/10.1186/s12909-025-06753-3.

[7] Kim, Y. H., Kim, B. H., Kim, J., Jung, B. & Bae, S. (2023). Item difficulty index, discrimination index, and reliability of the 26 health professions licensing examinations in 2022, Korea: a psychometric study. Journal of Educational Evaluation for Health Professions, 20. https://doi.org/10.3352/jeehp.2023.20.31.

[8] Makhlouf, K., Amouri, L., Chaabane, N. & EL-Haggar, N. (2020). Exam Questions Classification Based on Bloom’s Taxonomy: Approaches and Techniques. 2020 2nd International Conference on Computer and Information Sciences (ICCIS). Sakaka, Saudi Arabia. DOI: 10.1109/ICCIS49240.2020.9257698.

[9] Moghadamzadeh, A., Salehi, K. & Khodaie, E. (2011). A Comparison the Information Functions of the Item and Test in One, Two and Three Parametric Model of the Item Response Theory (IRT). International Conference on Education and Educational Psychology (ICEEPSY 2011). Istanbul, Turkey.

[10] Mohammed, M. & Omar, N. (2018). Question Classification Based on Bloom’s Taxonomy Using Enhanced TF-IDF. International Journal on Advanced Science, Engineering and Information Technology, 8(4-2), pp.1679–1685. DOI: https://doi. org/10.18517/ijaseit.8.4-2.6835.

[11] Opait, E.-E., Duca, A. & Olariu, M.-E. (2025). Evaluating Generative AI and Human Performance in Question-Answer Validation Tasks. 29th International Conference on Knowledge-Based and Intelligent Information & Engineering Systems (KES 2025), 270, pp.3973-3982. DOI: https://doi. org/10.1016/j.procs.2025.09.522.

[12] Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Hong Kong, China.

[13] Singhania, S., Razniewski, S. & Weikum, G. (2022). Predicting Document Coverage for Relation Extraction. Transactions of the Association for Computational Linguistics (TACL), 10, pp.207-223.

[14] Taheri, E. & Li, N. (2023). L2 Norm-Based Control Regularization for Solving Optimal Control Problems. IEEE Access, 11. DOI: 10.1109/ ACCESS.2023.3331382.

[15] Tam, D. T. & Quynh, P. T. (2024). Alignment between Course Learning Outcomes and Assessments: An Analysis. International Journal of TESOL & Education, 4(2). DOI: https://doi.org/10.54855/ ijte.24422.

[16] Tan, B., Armoush, N., Mazzullo, E., Bulut, O. & Gierl, M. J. (2025). A review of automatic item generation techniques leveraging large language. International Journal of Assessment Tools in Education, 12(2), pp.317-340. https://doi.org/10.21449/ ijate.1602294.

[17] Toraman, Ç., Karadağ, E. & Polat, M. (2022). Validity and reliability evidence for the scale of distance education satisfaction of medical students based on item response theory (IRT). BMC Medical Education, 22. https://doi.org/10.1186/s12909 025-06753-3.

[18] Zhai, X., Haudek, K. C. & Ma, W. (2022). Assessing Argumentation Using Machine Learning and Cognitive Diagnostic Modeling. Research in Science Education, 53, pp.405-424.

Articles in Issue