CogMath: Assessing LLMs’ Authentic Mathematical Ability from a Human Cognitive Perspective

Published in Proceedings of the 42nd International Conference on Machine Learning (ICML 2025), 2025

CogMath evaluates the mathematical abilities of large language models through three stages inspired by human cognition: problem comprehension, problem solving, and solution summarization. It uses nine fine-grained dimensions to distinguish apparent answer accuracy from genuine mastery.

Paper · arXiv

Recommended citation: Liu, J., Huang, Z., Dai, W., Cheng, C., Wu, J., Sha, J., Li, S., Liu, Q., Wang, S., & Chen, E. (2025). "CogMath: Assessing LLMs’ Authentic Mathematical Ability from a Human Cognitive Perspective." Proceedings of the 42nd International Conference on Machine Learning, 38692–38707.
Download Paper