📄 Research Article
Representation Robustness under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
This study probes how sensitive LLM mathematical problem solving is to the surface representation of an item — a question with direct bearing on Assessment Validity when LLMs are used for scoring or tutoring in STEM Education. Systematically varying representationally equivalent formulations (story problems, word-equations, symbolic equations, and isomorphic paraphrases) across 5 contemporary LLMs, the authors find substantial representational sensitivity: models frequently flip correctness across equivalent formulations, and even subtle paraphrase-level changes degrade performance despite preserved mathematical structure. A second, code-augmented condition constraining models to externalize reasoning as executable Python reveals strong latent capability in weak models but does not uniformly improve robustness — instead failures shift from opaque reasoning errors to protocol and execution violations. The work cautions that treating formulations as interchangeable conflates reasoning errors with interface failures, complicating AI Tutoring and diagnostic uses like LLM Cognitive Diagnosis Handwritten Math. It connects to measurement concerns in Reinforcement Learning Measurement Model Assessment and to reasoning scaffolds in Epistemic Proactivity Math.
Connected Concepts
Connected Articles
Citation
Nath, Graf, Zhang & Zapata-Rivera (2026). Representation Robustness under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving. arXiv:2607.20520. HCI International 2026.