Research Article
Math Education Digital Shadows for facilitating learning with LLMs: Math performance, anxiety and confidence in simulated students and AIs
Synthesis: MEDS (Math Education Digital Shadows) is a 28,000-record dataset from the CogNosco Lab at the University of Trento that maps how 14 large language models reason about and report mathematics under two prompting conditions: as a baseline AI assistant, and as a synthetic human persona — a digital shadow — assembled from weighted socio-demographic attributes and Big Five trait scores. Each record runs four tasks: an open interview about mathematics, three psychometric scales (MSES, AMAS and MSEAQ) with written justifications, behavioral forma mentis networks over 50 cue words, and 18 high-school math problems answered with reasoning and a confidence rating. Validation shows the models are internally consistent and psychologically differentiated, but also family-specific: human-simulated personas carry human-like negative math attitudes and more tangential reasoning, baseline assistants default to inflated Self-Efficacy and suppressed anxiety, and several models — Qwen3 4B Uncensored and Qwen3.5 9B most sharply — report confidence above 0.90 while answering close to chance. MEDS is an observational resource for learning analytics, cognitive scientists and developers of safer AI tutoring, not a proxy for real students.
Key Findings
- A 28,000-record dataset across 14 models and two modes. MEDS contains 28,000 shadow records — 2,000 personas per model, five JSON files per persona across four tasks — generated from 14 open- and closed-weight models spanning the Mistral, Qwen, DeepSeek, Granite, Phi and Grok families. Each persona ran either in human mode (a synthetic persona with demographics and OCEAN traits) or LLM mode (persona set to null, the model answering as itself). Inference took roughly 422 hours of continuous computation on four GPU machines, beyond most research groups' reach without dedicated infrastructure.
- Accuracy and confidence diverge, and the gap runs in both directions. Task 4 contrasted 18 MSES-R multiple-choice problems with self-reported confidence on a 1–5 scale. Grok 4.1 Fast, DeepSeek Chat and several Mistral Small variants were underconfident (accuracy above confidence), Ministral 14B and Anita 24B were reasonably calibrated, while the Qwen family and Ministral 3B were sharply overconfident: Qwen3 4B (Uncensored) and Qwen3.5 9B reported confidence above 0.90 while their accuracy stagnated near 0.55 — the "myopic overconfidence" pattern where incorrect or hallucinated answers are asserted with high certainty.
- Baseline assistants default to a confident, low-anxiety self-portrait. On the psychometric scales, human-mode runs produced wide, plausible variance, whereas baseline LLM runs had negligible variance — Anita 24B (Uncensored), Ministral 14B, Ministral 3B, Phi-4 (Reasoning+) and Qwen3 4B (Thinking) returned identical answers, so no distribution could even be plotted. Where LLM-mode distributions existed they skewed right on MSES self-efficacy, clustered at the low-anxiety end of AMAS, skewed right on the MSEAQ self-efficacy subscale and left on its anxiety subscale.
- Emotional profiles of AI assistants are flatter and less joyful than their simulated humans. Using EmoAtlas on the Task 1 interview, AI assistant profiles showed the highest trust scores in most families (z = 2.64 for Mistral Small 4, against 1.75 for generated "Math Lovers" and 1.27 for "Math Haters") but conspicuously low joy (z = 0.40, below even the 0.43 of the math haters). All Qwen models showed lower positive emotion for AI profiles except trust, and systematically minimized fear; in Ministral 3B the AI profile recorded the lowest positive and highest negative emotion scores of the three groups.
- The same five logical fallacies appear in both prompting modes. Across all 14 models the fallacy taxonomy (DistilBERT fallacy classifier, 85th-percentile confidence cut-off) collapsed onto five categories: fallacy of logic (0.56 synthetic human / 0.58 LLM), fallacy of relevance (0.41 / 0.25), faulty generalization (0.21 / 0.30), appeal to emotion (0.14 / 0.11) and circular reasoning (0.10 / 0.13). The authors read the convergence as a structural property of LLM-generated language rather than a role-specific behavior.
- Reasoning variants are the least fallacious, but not uniformly. Phi-4 (Reasoning+), Qwen3 4B (Thinking) and Ministral 3B exceeded a 60% no-fallacy rate, and the smallest Ministral was the cleanest of its family. Elsewhere the biases were family-specific: Grok 4.1 Fast (Reasoning) recorded the highest fallacy-of-relevance score in baseline mode (0.517), Granite 4 Tiny concentrated almost entirely on circular reasoning in baseline data (0.476), and Qwen3.5 9B inverted between conditions, with low fallacy rates in human mode and high relevance (0.453) plus appeal to emotion (0.305) in LLM mode.
- Cognitive networks reproduce the human science-positive / math-negative split. Behavioral forma mentis networks centered on "mathematics" for Anita 24B connect it negatively to equation, algorithm, computation, proof and difficult — a methodological aversion matching the human learners in earlier forma mentis work — while "science" is surrounded by positive concepts across all 14 models. Censorship mattered: the uncensored Anita 24B assigned about 52% negative valence to mathematics against roughly 35% for its censored counterpart Mistral Small 3.2, and the Qwen family mostly made neutral judgments, with Qwen3.5 9B the outlier at about 20% negative for mathematics.
- Personas are plausible by construction, not sampled from a real population. Attributes were drawn with pseudo-realistic weights (heterosexual labels more likely than LGBTQI+ ones) and status-consistent constraints — education level conditioned on age, marital status on age and orientation, number of children on age, marital status and orientation — with ages from 18–30 plus 35, 40, 50, 60 and 70, and occupations, parental education, migration status, religion, hobbies and favourite or disliked subjects all recorded. Cleaning kept roughly 99% of DeepSeek Chat records but discarded about 55% of Granite 4 Tiny's, using BERT-based semantic alignment at a 0.85 cosine threshold.
- The motivating gap is adoption without evidence. In the 2025 EU-wide Eurostat module, 9.4% of people aged 16–74 and 39.3% of those aged 16–24 reported using GenAI tools for formal education, yet most benchmarks score answers without capturing the socio-cognitive conditions around them. MEDS is offered as an extension of existing math benchmarks rather than a replacement, aimed at learning analytics experts, cognitive scientists, teachers and AI safety researchers.
How MEDS was built
Every run began with a system-level role instruction that fixed its personification mode. Human-mode runs were told to role-play a single human respondent and stay consistent with the persona's demographics and psychological descriptors; LLM-mode runs were told to speak as the language model itself, not to pretend to be human, and to interpret psychometric items in terms of AI functioning rather than human life events. A full JSON schema enforced by response_json_schema() prescribed fixed keys and value constraints (integer ratings 1–5, single-word associations, A–E options with confidence 1–5), which kept records prompt-traceable and machine-parsable.
Four sequential calls followed: Task 1 put seven interview questions about mathematics, math anxiety, prior AI use and specific procedures (second-order equations, stationary points, PCA); Task 2 administered the Mathematics Self-Efficacy Scale (9 items), the Abbreviated Math Anxiety Scale (9 items) and the 28-item Mathematics Self-Efficacy and Anxiety Questionnaire with reverse-valence items remapped; Task 3 collected free associations and 1–5 valence ratings for 50 cue words in two batches of 25; Task 4 asked for solutions to the 18-item MSES-R problems subscale with step-by-step reasoning and a confidence score.
Selection of the 14 models was a compromise between open-weight availability, API availability, representativeness across LM Studio, ollama and vLLM, and computational resources. Cleaning was the most fragile stage: because models rephrase JSON keys, keys were matched to the original questions by BERT Base Uncased cosine similarity, with 0.85 selected after testing 0.85, 0.90 and 0.95, and unrecoverable files discarded and regenerated. Network records required associations for at least 75% of the 50 cue words. Validation then checked schema integrity, persona consistency and cross-task completeness before any analysis, and the whole JSON dataset plus the persona-generation weights are published openly for reuse and extension.
Math performance, confidence and prompt-conditioned bias
Task 4 is the part that reads like a conventional math benchmark: per-question accuracy, averaged per model, set against min-max rescaled confidence. Its value is in the calibration curve rather than the leaderboard, because accuracy and certainty come apart in ways that matter when a model is tutoring: underconfident families undersell correct answers, calibrated families (Ministral 14B, Anita 24B) stay aligned, and overconfident families convert errors into confident teaching. For learning analytics the consequences are concrete — a dashboard comparing performance with confidence can flag misplaced certainty on specific problem types, and fallacy rates by topic can identify where human oversight remains most necessary.
The hallucination framing matters most for mathematics education, where a wrong step is not a stylistic slip. Comparing both prompting modes isolates the effect of the persona constraint from the model itself: human-simulated runs introduced more tangential reasoning (relevance 0.41 vs 0.25), while baseline runs overgeneralised more (faulty generalization 0.30 vs 0.21), and the dominant fallacy of logic was essentially invariant across conditions. Across the whole dataset, the authors argue, this convergence points to argumentative homogeneity in generated text rather than a distinct "human-like" reasoning style, which is a caution for anyone treating persona prompting as a route to authentic student reasoning.
Simulated student shadows and their grounding in real student data
A digital shadow, as the paper defines it, is deliberately weaker than a digital twin: it records what a system does across controlled conditions through unidirectional data acquisition, without updating or influencing an external counterpart. A Math Education Digital Shadow is therefore the structured math psychological and problem-solving record of one model under a specified persona, prompting mode and four-task battery — a trace to be analyzed, not an agent to be run. This is what makes latent human variables such as Self-Efficacy, anxiety and semantic association explicit and experimentally manipulable inside a Simulation rather than merely observed.
Grounding comes from comparison rather than from real participants. Task 3 networks are read against the human forma mentis results of Stella and colleagues (2019), which found students associating science positively and its mathematical building blocks negatively; MEDS reproduces that bias, extends it to 14 models, and adds the uncensored-versus-censored contrast that past GPT-based studies could not make. Task 2 distributions are compared against human psychometric profiles from the source instruments, and Task 1 emotional profiles against published cognitive network findings on math anxiety. No real students were recruited or simulated from individual records, so these are distributional analogies, not replications — a limit the authors state plainly when they call the dataset "not a proxy for human opinion".
Anxiety and confidence modeling and what it is intended to support
The three psychometric instruments cover complementary ground: MSES and the MSEAQ self-efficacy subscale measure belief in one's ability to solve mathematical problems, AMAS and the MSEAQ anxiety subscale measure the emotional cost of engaging with them. Running all of them in both modes turns the usual single score into a profile, and profile variance becomes the signal: wide in human mode, near-zero in baseline mode, with the baseline models projecting systematically higher Self-Efficacy and lower anxiety than their simulated peers. Since self-reports from a model are behavioral traces rather than testimony, the authors use the term calibration and reserve judgment about whether these numbers describe anything internal.
That framing is what the authors mean by psychometrically aware evaluation. Intended uses include teachers and professors gauging whether a model is psychologically appropriate before students use it; developers of tutoring systems building adaptive support that responds to a learner's anxiety level; safety researchers auditing which families are suitable for anxious or struggling students; and measurement researchers using the score distributions and confidence-accuracy gaps to baseline other synthetic-participant datasets. Persona metadata (gender, age, migration status) also enables stereotype testing, checking whether demographic attributes shift confidence or anxiety scores in ways that would disadvantage particular Learners. Scores can be filtered by city and country to compare models against locally relevant profiles, and the release includes datasets for text classifiers over Tasks 1 and 4 — a route towards moderation of anxiety-signaling content and preference data for better-calibrated explanations.
What this means for practice
- Researchers. Run baseline mode as your control before reading any persona-conditioned score: baseline runs frequently returned identical scores across personas and scales, which is evidence of mode sensitivity rather than a psychological measurement.
- Researchers. Treat a model's stated confidence as a score with no calibration warranty, since MEDS records systematic overconfidence: a high self-efficacy or low anxiety output says nothing about how a real student would answer the same items.
- Educators. Use the anxiety and self-efficacy profiles to judge whether a model is psychologically appropriate for your students before they work with it, not to describe your own class — the authors call MEDS an observational resource and assertively not a stand-in for student data.
- Learning analytics designers. Filter scores by city and country before comparing a model with the students you serve, because persona attributes are weighted, pseudo-realistic distributions rather than a sampling frame.
- Edtech designers. Check the family, mode and version of the system you deploy rather than the one you benchmarked: uncensored and censored variants diverge, discard rates ranged from roughly 1% to about 55%, and refusal behavior and reasoning styles shift between versions.
Limitations
The overriding caveat is simulation validity. MEDS was generated entirely by LLMs and has no relation to any actual individual; the personas are synthetic constructions with weighted, pseudo-realistic attribute distributions rather than a sampling frame, so no finding generalizes to real students, classrooms or schools. Persona-based prompting produces behaviorally differentiated outputs, but differentiation is not fidelity: item 4 above shows AI profiles that are emotionally flatter and less joyful than their simulated humans, and item 5 shows fallacy categories that do not change when the model is told to be human, which means human-like responses in this dataset cannot be assumed to carry the reasoning style of real learners.
Second, the psychometric results depend on models answering about themselves. Baseline runs frequently returned identical scores across personas and scales, which is evidence of mode sensitivity rather than a psychological measurement; where distributions do exist, interpreting "math anxiety" in an AI assistant rests on analogy with the human instrument. Third, data quality is uneven across the model set: the discard rate ranged from roughly 1% to about 55%, so cross-model comparisons carry unequal effective sample sizes, and the BERT key-alignment threshold of 0.85 was chosen empirically for maximum recoverable structure rather than from a validation study of its own error rate.
Fourth, the psychological battery is narrow. Three instruments, 50 cue words and 18 high-school problems cannot cover the range of mathematical content, task formats or classroom contexts that matter for K-12 and Higher Education practice, and the reasoning-verbosity-to-fallacy-rate relationship the authors flag as important is left for future work. Fifth, the ecosystem moves faster than the dataset: model versions, refusal behavior and reasoning styles all shift, so the snapshot describes 14 specific systems rather than the families they stand for. The authors position MEDS accordingly — an observational resource for auditing prompt-conditioned GenAI behavior, assertively not a stand-in for student data — and note that because it is synthetic, using it carries no ethics-review or Privacy constraint, which is precisely why it cannot answer the questions about real students it is often recruited for.
Connected Concepts
- Simulating Students — the core method: LLM personas standing in for learners and analyzed as shadows
- Anxiety and Stress — math anxiety as the affective dimension MEDS tries to make measurable in models
- Self-Efficacy — the belief dimension carried by MSES and the MSEAQ self-efficacy subscale
- Large Language Models (LLMs) — the 14 systems across six families whose behavior the dataset records
- Psychometrically Aware AI — psychometric scales adapted to AI respondents, with mixed distributions as the signal
- Trust Calibration — the accuracy-versus-confidence gap and myopic overconfidence
- Benchmark — MEDS as a multidimensional extension of score-only math benchmarks
- Network Analysis — forma mentis networks linking math concepts with emotional valence
- Learning Analytics — performance, confidence and fallacy data as a source of discrepancy indicators
- Bias Mitigation — family-specific biases, uncensored-versus-censored differences and stereotype testing
- Hallucination Risk — confident mathematical errors in a tutoring context
- Limitations in AIEd Research — simulation validity and the inference limits the authors state explicitly
Connected Articles
- Towards Valid Student Simulation with Large Language Models — What makes student simulation with LLMs valid, and where it fails
- Simulating Students with Large Language Models: A Review of Architecture, Mechanisms, and Role Modelling in Education with Generative AI — Review of architectures and mechanisms for simulating students
- Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators — Whether simulator personas reproduce real misconceptions or flatter the prompt
- The Absent Cognitive Baseline: Theorizing a Structural Gap in AI-Native College Students' Academic Self-Assessment — The missing baseline problem in AI-native students' self-assessment
- The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows — Argumentative failure and misattributed competence in LLM reasoning
- AI-Powered Math Tutoring: Platform for Personalized and Adaptive Education — A personalized adaptive math tutoring platform built on LLMs
- Taming the Black Box: Design Principles for Rule-Integrated LLM Tutoring Systems in Primary School Mathematical Problem Solving — LLM maths tutoring for younger learners with rule-based constraints
- Beyond checking: verification quality, reliance calibration, and learning in generative AI-assisted higher education — Verification, reliance calibration and learning in GenAI-assisted study
- MathBuddy: Affective Math Tutoring — Affective math tutoring that responds to learner emotion
Citation
Esposito, N., Tricarico, A., Porzio, L., Aghazadeh Ardebili, A., & Stella, M. (2026). Math Education Digital Shadows for facilitating learning with LLMs: Math performance, anxiety and confidence in simulated students and AIs. arXiv:2604.27618.