CHATGPT’S MATHEMATICAL PROBLEM-SOLVING PERFORMANCE IN GRADE 11 GRAPH THEORY

CHATGPT’S MATHEMATICAL PROBLEM-SOLVING PERFORMANCE IN GRADE 11 GRAPH THEORY

Pham Thanh Dat datpt@hcmue.edu.vn Ho Chi Minh City University of Education 280 An Duong Vuong street, Cho Quan ward, Ho Chi Minh City, Vietnam
Goi Ngoc Tu* goingoctu9at4@gmail.com Ho Chi Minh City University of Education 280 An Duong Vuong street, Cho Quan ward, Ho Chi Minh City, Vietnam
Tran Khanh Anh Kaanh3851D@gmail.com Ho Chi Minh City University of Education 280 An Duong Vuong street, Cho Quan ward, Ho Chi Minh City, Vietnam
Ho Ngoc Bao Tram hongocbaotram2020@gmail.com Ho Chi Minh City University of Education 280 An Duong Vuong street, Cho Quan ward, Ho Chi Minh City, Vietnam
Summary: 
Although ChatGPT’s mathematical problem-solving performance has been extensively investigated across various domains, evaluating its specific proficiency in Graph Theory within the 2018 General Education Curriculum remains a notable research gap that warrants clarification. This study analyzes ChatGPT’s problem-solving capabilities through the detailed assessment of 257 designated grade 11 Mathematics tasks derived from three currently used textbook series, while simultaneously identifying the model’s underlying systematic errors. Empirical results clearly demonstrate significant performance variance across these instructional resources, highlighting five typical error categories the model frequently commits when processing graph data. Ultimately, these findings not only provide practical empirical evidence concerning the true capacity of artificial intelligence in modern mathematics education but also critically assist educators in clearly recognizing the tool’s operational limitations. Consequently, this research establishes a solid foundation for proposing effective and appropriate integration strategies within the ongoing context of educational reforms.
Keywords: 
ChatGPT
graph theory Maths
Problem-Solving Performance
error.
Refers: 

[1] Díaz, B. & Nussbaum, M. (2024). Artificial intelligence for teaching and learning in schools: The need for pedagogical intelligence. Computers & Education, 217, 105071. https://doi.org/10.1016/j. compedu.2024.105071.

[2] Fergus, S., Botha, M. & Ostovar, M. (2023). Evaluating academic answers generated using ChatGPT. Journal of Chemical Education, 100(4), pp.1672-1675. https://doi.org/10.1021/acs.jchemed.3c00087.

[3] Gandolfi, A. (2025). GPT-4 in education: Evaluating aptness, reliability, and loss of coherence in solving calculus problems and grading submissions. International Journal of Artificial Intelligence in Education, 35(1), pp.367-397. https:// doi.org/10.1007/s40593-024-00403-3.

[4] Giannos, P. & Delardas, O. (2023). Performance of ChatGPT on UK standardized admission tests: insights from the BMAT, TMUA, LNAT, and TSA examinations. JMIR Medical Education, 9(1), Article e47737. https://doi.org/10.2196/47737.

[5] Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., ... & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), pp.1-55. https://doi. org/10.1145/3703155.

[6] Kaya, D. & Yavuz, S. (2025). Can generative aI and ChatGPT break human supremacy in mathematics and reshape competence in cognitive-demanding problem-solving tasks?. Journal of Intelligence, 13(4), p.43.

[7] Kortemeyer, G. (2023). Could an artificial-intelligence agent pass an introductory physics course?. Physical Review Physics Education Research, 19(1), p.010132. https://doi.org/10.1103/PhysRevPhysE ducRes.19.010132.

[8] Lê Anh Vinh, Bùi Thị Diển, Lê Quang Quân & Vũ Văn Luân. (2023). Khả năng thực hiện bài kiểm tra định kì môn Toán và môn Ngữ văn cấp Trung học của công cụ ChatGPT: Kết quả nghiên cứu và một số khuyến nghị ban đầu. Tạp chí Khoa học Giáo dục Việt Nam, 19(2), tr.1-10.

[9] Le Thai Bao Thien Trung, Nguyen Minh Dat, Tang Minh Dung, Trinh Van Thanh & Nguyen Minh Nhut. (2025). Promoting teachers’ use of ChatGPT: A case study on generating real-world problems in 10th grade algebra instruction. HCMUE Journal of Science, 22(3), pp.500-511.

[10] Oh, S. (2024). Evaluating Mathematical Problem Solving Abilities of Generative AI Models: Performance Analysis of o1-preview and gpt-4o Using the Korean College Scholastic Ability Test. IEEE Access, 13, pp.1227-1235.

[11] Pardos, Z. A. & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills. PLOS ONE, 19(5), e0304013. https://doi.org/10.1371/ journal.pone.0304013.

[12] Phạm Tất Thành. (2024). Đánh giá kết quả phản hồi của Chatbot Bard trong thực hiện bài thi môn Địa lí kì thi tốt nghiệp Trung học Phổ thông Quốc gia Việt Nam từ năm 2019 đến năm 2023. Tạp chí Giáo dục, 24(4), tr.35-40.

[13] Polverini, G. & Gregorcic, B. (2024). Performance of ChatGPT on the test of understanding graphs in kinematics. Physical Review Physics Education Research, 20(1), 010109. https://doi.org/10.1103/ PhysRevPhysEducRes.20.010109

Articles in Issue