[1] Ahmed, A., Kerr, E. & O’Malley, A. (2025). Quality assurance and validity of AI-generated single best answer questions. BMC Medical Education, 25.
[2] Circi, R., Hicks, J. & Sikali, E. (2023). Automatic item generation: foundations and machine learning based approaches for assessments. Frontiers in Education, 8. https://doi.org/10.3389/ feduc.2023.858273.
[3] Gunawan, D., Sembiring, C. A. & Budiman, M. A. (2018). The Implementation of Cosine Similarity to Calculate Text Relevance between Two Documents. 2nd International Conference on Computing and Applied Informatics (ICCAI 2017). DOI: 10.1088/1742-6596/978/1/012120.
[4] He, X., Lin, Z., Gong, Y., Jin, A.-L., Zhang, H., Lin, C., Chen, W. (2024). AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Industry Track, pp.165 190. Mexico City, Mexico.
[5] Hillsley, K. (2025). Question format is the best predictor of item discrimination: a multivariable analysis. Journal of Microbiology & Biology Education, 26(3).
[6] Khojasteh, L., Kafipour, R., Pakdel, F. & Mukundan, J. (2025). Empowering medical students with AI writing co-pilots: design and validation of AI self assessment toolkit. BMC Medical Education, 25. DOI :https://doi.org/10.1186/s12909-025-06753-3.
[7] Kim, Y. H., Kim, B. H., Kim, J., Jung, B. & Bae, S. (2023). Item difficulty index, discrimination index, and reliability of the 26 health professions licensing examinations in 2022, Korea: a psychometric study. Journal of Educational Evaluation for Health Professions, 20. https://doi.org/10.3352/jeehp.2023.20.31.
[8] Makhlouf, K., Amouri, L., Chaabane, N. & EL-Haggar, N. (2020). Exam Questions Classification Based on Bloom’s Taxonomy: Approaches and Techniques. 2020 2nd International Conference on Computer and Information Sciences (ICCIS). Sakaka, Saudi Arabia. DOI: 10.1109/ICCIS49240.2020.9257698.
[9] Moghadamzadeh, A., Salehi, K. & Khodaie, E. (2011). A Comparison the Information Functions of the Item and Test in One, Two and Three Parametric Model of the Item Response Theory (IRT). International Conference on Education and Educational Psychology (ICEEPSY 2011). Istanbul, Turkey.
[10] Mohammed, M. & Omar, N. (2018). Question Classification Based on Bloom’s Taxonomy Using Enhanced TF-IDF. International Journal on Advanced Science, Engineering and Information Technology, 8(4-2), pp.1679–1685. DOI: https://doi. org/10.18517/ijaseit.8.4-2.6835.
[11] Opait, E.-E., Duca, A. & Olariu, M.-E. (2025). Evaluating Generative AI and Human Performance in Question-Answer Validation Tasks. 29th International Conference on Knowledge-Based and Intelligent Information & Engineering Systems (KES 2025), 270, pp.3973-3982. DOI: https://doi. org/10.1016/j.procs.2025.09.522.
[12] Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Hong Kong, China.
[13] Singhania, S., Razniewski, S. & Weikum, G. (2022). Predicting Document Coverage for Relation Extraction. Transactions of the Association for Computational Linguistics (TACL), 10, pp.207-223.
[14] Taheri, E. & Li, N. (2023). L2 Norm-Based Control Regularization for Solving Optimal Control Problems. IEEE Access, 11. DOI: 10.1109/ ACCESS.2023.3331382.
[15] Tam, D. T. & Quynh, P. T. (2024). Alignment between Course Learning Outcomes and Assessments: An Analysis. International Journal of TESOL & Education, 4(2). DOI: https://doi.org/10.54855/ ijte.24422.
[16] Tan, B., Armoush, N., Mazzullo, E., Bulut, O. & Gierl, M. J. (2025). A review of automatic item generation techniques leveraging large language. International Journal of Assessment Tools in Education, 12(2), pp.317-340. https://doi.org/10.21449/ ijate.1602294.
[17] Toraman, Ç., Karadağ, E. & Polat, M. (2022). Validity and reliability evidence for the scale of distance education satisfaction of medical students based on item response theory (IRT). BMC Medical Education, 22. https://doi.org/10.1186/s12909 025-06753-3.
[18] Zhai, X., Haudek, K. C. & Ma, W. (2022). Assessing Argumentation Using Machine Learning and Cognitive Diagnostic Modeling. Research in Science Education, 53, pp.405-424.

