This study examines the development of the ChatGPT-4o model developed by OpenAI to solve university-level probability problems over 8 months. The study is a descriptive study based on the general screening model. Within the scope of the study, 38 probability problems with four different grades and different difficulty levels easy, medium, and hard were modeled, and the answers of ChatGPT-4o were evaluated in terms of accuracy and consistency of the process steps at 8-month intervals. The difference in success between the problem-solving measurements of ChatGPT-4o was not only statistically significant but also practically remarkable. The problem-solving skill of ChatGPT-4o exhibited a statistically significant performance increase at different difficulty levels over time. Furthermore, when all levels were considered collectively, the general success rate improved substantially. The findings reveal that the model significantly increased performance, especially in problems with conditional probability and union/intersection events. At the same time, improvements were observed in the model’s ability to present the solution process in a more explanatory and systematic way. The results reveal ChatGPT-4o’s potential for reasoning during problem-solving processes and demonstrate how this capability has evolved over time. Overall, the integration of ChatGPT-4o models into mathematics education may support students’ fundamental problem-solving skills and enable teachers to allocate classroom time to more complex tasks. However, the model’s limitations should not be overlooked. In this context, further comprehensive research is needed to examine the effects of artificial intelligence–based systems on instructional processes, ensuring their effective, reliable, and pedagogically appropriate integration into mathematics teaching.
Sarıkaya et al. (2026) studied this question.