Journal of Technology and Information Education 2026, 18(1):138-147 | DOI: 10.5507/jtie.2026.009

BARRIERS OF ARTIFICIAL INTELLIGENCE IN SOLVING LOGIC PROBLEMS FOR YOUNGER SCHOOL-AGE STUDENTS

Martina UHLÍŘOVÁ, Jitka LAITOCHOVÁ, Lucie ŠTAFOVÁ, Veronika VESELÁ
Univerzita Palackého v Olomouci, Česká republika

This paper focuses on the analysis of errors in mathematical problem-solving generated by selected artificial intelligence systems, with particular attention to tasks designed for younger primary school pupils. The study is based on a set of 144 problems from the Mathematical Kangaroo competition, with a detailed analysis of the Ecolier category. The tasks were submitted to three generative AI systems (Gemini 2.0 Flash, ChatGPT, and Microsoft 365 Copilot). The study aimed to compare the performance of individual AI systems and to identify the structure and causes of errors in their solutions. The results reveal statistically significant differences between the systems and highlight increased error rates in geometric and context-based logical tasks. These findings are interpreted through the lens of embodied cognition, which helps to explain the discrepancy between formally correct but contextually inadequate AI solutions and the intuitive reasoning strategies employed by children.

Keywords: artificial intelligence, Mathematical Kangaroo, Ecolier category, word problems, primary education.

Received: February 9, 2026; Revised: August 9, 2026; Accepted: February 9, 2026; Published: September 2, 2026  Show citation

ACS AIP APA ASA Harvard Chicago Chicago Notes IEEE ISO690 MLA NLM Turabian Vancouver
UHLÍŘOVÁ, M., LAITOCHOVÁ, J., ŠTAFOVÁ, L., & VESELÁ, V. (2026). BARRIERS OF ARTIFICIAL INTELLIGENCE IN SOLVING LOGIC PROBLEMS FOR YOUNGER SCHOOL-AGE STUDENTS. Journal of Technology and Information Education18(1), 138-147. doi: 10.5507/jtie.2026.009
Download citation

References

  1. Bártek, K., Bártková, E., Mrkvan, K., & Nocar, D. (2025). Možnosti a limity LLM v geometrické přípravě budoucích učitelů 1. stupně základních škol. Elementary Mathematics Education Journal, 7(2). 89-103. https://emejournal.upol.cz/Issues/Vol7No2/Vol7No2_Bartek-et-al.pdf Go to original source...
  2. Boye, J., & Moell, B. (2025). Large language models and mathematical reasoning failures. arXiv. https://arxiv.org/abs/2502.11574
  3. Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., & Steinhardt, J. (2021). Measuring Mathematical Problem Solving With the MATH Dataset. arXiv. https://arxiv.org/abs/2103.03874
  4. Kambhampati, S. (2024). Can large language models reason and plan? Annals of the New York Academy of Sciences, 1534(1), 15-18. https://doi.org/10.1111/nyas.15125 Go to original source...
  5. Lakoff, G., & Núñez, R. E. (2000). Where mathematics comes from: How the embodied mind brings mathematics into being. New York: Basic Books. https://pages.ucsd.edu/~rnunez/ COGS252_Readings/Preface_Intro.PDF
  6. Shapiro, L., & Spaulding, S. (2025) Embodied Cognition. The Stanford Encyclopedia of Philosophy (Summer 2025 Edition). https://plato.stanford.edu/archives/sum2025/entries/ embodied-cognition/
  7. Schoenfeld, A. H. (2016). Learning to Think Mathematically: Problem Solving, Metacognition, and Sense Making in Mathematics (Reprint). Journal of Education, 196(2), 1-38. Go to original source...
  8. Strohmaier, A. R., Van Dooren, W., Seßler, K., Greer, B., & Verschaffel, L. (2025). Large language models don't make sense of word problems: A scoping review from a mathematics education perspective. arXiv. https://arxiv.org/abs/2506.24006
  9. UNESCO. (2023). Guidance for generative AI in education and research. United Nations Educational, Scientific and Cultural Organization. https://unesdoc.unesco.org/ark:/48223/pf0000386693
  10. Xu, W., Wang, J., Wang, W., Chen, Z., Zhou, W., Yang, A., Lu, L., Li, H., Wang, X., Zhu, X., Wang, W., Dai, J., & Zhu, J. (2025). VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models. arXiv. https://arxiv.org/abs/2504.15279
  11. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large models. Advances in Neural Information Processing Systems, 35, 24824-24837. Go to original source...
  12. Wilson, M. (2002). Six views of embodied cognition. Psychonomic Bulletin & Review, 9, 625-636. https://doi.org/10.3758/BF03196322 Go to original source...