On this page

Synthesis: AMPT is Carnegie Learning's chat-based tool for personalising math story problems: a student tells an LLM about their interests in a short conversation, and the tool then writes a problem wrapped around those interests. Norberg, Murphy and Ritter built it over 15 months through six 30-minute co-design sessions with 93 middle-school students, most from economically disadvantaged, majority-Black urban districts. Students chose between Wacky, Directed and Expansive interest-gathering modes, named the characters themselves, then accepted or requested revisions to a generated problem and rated it. Across the work students created 331 problems over 7 math domains, rating their own problems 4.17 of 5 stars and agreeing strongly (M = 4.73) that naming the characters mattered to them. Their reported pre-to-post math attitudes moved unevenly: perceived ability and interest and value did not change significantly, but sense of belonging in mathematics rose 4.79% (Cohen's d = 0.48, p = .01), driven by the Membership and Acceptance subscales. In a MATHia pilot, 15,495 students rated deployed problems, 61% favourably, with student-authored problems among both the most and least liked.

Key Findings

  1. The tool was developed iteratively with 93 middle-school students. AMPT was designed and tested over 15 months through six 30-minute co-design sessions with 7th and 8th graders from 4 metropolitan school districts and 1 rural district. The metropolitan districts reported 96–100% of students economically disadvantaged, 86–96% identifying as Black and 4–5% as Hispanic; the rural district reported 90% white and 43% economically disadvantaged. Students participated outside regular school hours.
  2. Students produced a large corpus of personally contextualised problems. Co-design students created 331 problems, M = 3.64 per student (SD = 1.68), spanning 7 distinct math domains including probability, ratios, linear equations and basic algebra. Students rated their own problems M = 4.17 stars out of 5 (SD = 0.93).
  3. Students reported a representation gap they cared about closing. On a 5-point scale, students said their current math class's story problems represented their culture at only M = 2.36 (SD = 1.14) and their interests at M = 2.77 (SD = 1.39), while reporting it was more important to see their interests reflected (M = 4.08, SD = 1.16) than their culture (M = 3.01, SD = 1.43); interests were judged more important than culture, t(73) = 7.26, p < .001.
  4. Student-authored problems were preferred to traditional ones. Students preferred the problems they and their peers authored over traditional problems, M = 4.40 (SD = 0.97, n = 92); wanted their problems used in educational math software, M = 4.58 (SD = 0.89, n = 83); and judged peer problems in Arena mode more interesting than the traditional problems they see in class, M = 4.39 (SD = 0.92, n = 41).
  5. Naming characters and choosing topics were the features students valued most. Students agreed broadly that naming the characters was important to their experience, M = 4.73 (SD = 0.85, n = 73). Across two cohorts (n = 44) the most common topic choice was "None of these" at 21%, then Fantasy Worlds 14%, Music 13%, Video Games 11%, Food 10%, History 8% and Sports 7%; Careers was chosen for only 4% of problems.
  6. Belonging rose, but interest and perceived ability did not. For three cohorts surveyed before and after AMPT (n = 43), perceived ability did not change significantly and interest and value combined rose only marginally, 3.56%, p = .08. Sense of belonging in mathematics increased by 4.79% (Cohen's d = 0.48, t(41) = 2.76, p = .01), an effect driven by the Membership and Acceptance subscales rather than Affect, Desire-to-Fade or Trust.
  7. At scale, reactions were favourable but mixed. In a MATHia pilot, 71 approved student co-authored problems were embedded in the software and 15,495 students rated new and old problems; rating was optional and 61% of the time students declined. Among problems that were rated, 61% received a thumbs up. Student-authored problems accounted for 8 of 10 of the most well-liked problems in the two domains tested, yet the least-liked probability problem was student-authored and received just 51% thumbs up despite its author's 5-star rating.
  8. A small pilot linked peer authorship to lower error odds. In a preliminary pilot study, the odds of error were lower for peer-authored math story problems than for the comparison problems, z = 5.23, p < .001. The authors report the current field study is ongoing with a similar early trend, and plan to test whether expressed interest in a problem context moderates those error rates.

How AMPT works, end to end

AMPT is a problem-posing tool, not a problem bank: the student authors the context and an LLM authors the mathematics. The pipeline has three stages. In interest gathering, a chat window fills the screen and the student converses with an LLM in one of three modes. In Wacky and Directed modes students use a point-and-click interface to choose a topic, two characters, a place and an activity; in Directed mode the options are generated by the LLM to be consistent with the student's earlier choices and produce cohesive scenarios, while Wacky mode bypasses the LLM during interest gathering and draws on a human-curated, deliberately random pool, allowing unconventional combinations and avoiding direct student–LLM conversation for families or schools that want that. In Expansive mode the student selects only the topic and then types free responses to increasingly detailed AI questions, with the AI prompted first to show comprehension and then to probe the aspects of the topic that interest the student. Conversation length is configurable, and in practice four exchanges proved enough to gather the needed detail without losing the student's attention. All modes end with the student suggesting names for the characters and confirming or refining the AI's understanding of their interests. Character naming was ceded to students precisely because early cohorts found the chat repetitive when AMPT reused similar topics and names.

In problem creation, an LLM first summarises the student's interests to focus the writing agent, which receives that summary, the full conversation transcript and the target math learning domain. Administrators select the learning domain in advance, and new domains can be added through an admin interface; mathematical standards and requirements live in the prompt, with fine-tuning on the full set of MATHia problems deliberately avoided so that existing topics would not bias the model against the student's vision. AMPT generates five candidate problems, evaluates each separately for interest overlap, adherence to the domain instructions, audience appropriateness and coherence, using reasoning LLMs plus machine learning components such as toxic-bert for toxic content and type-token ratio for semantic overlap, then a reasoning LLM (o3-mini) selects the best candidate in light of the whole history. In student evaluation, the student sees the problem with buttons to accept it or request a revision; freeform revision requests pass first to an intermediary LLM that turns vague input into usable instructions, for example reading "what about Amber?" as a request to develop what a named character is doing. Accepting prompts a 5-point star rating, and 4- or 5-star problems become candidates for MATHia, Carnegie Learning's intelligent tutoring system, after human review.

The two deployment loops close the design: a student-facing Arena mode shows peer-authored problems for thumbs up or thumbs down with no maths attached, gathering broad-grain interest signals, and problems embedded in MATHia carry a thumbs up/down control that lets tens of thousands of students rate a story while their error and success data are logged. Interestingly, half of the students seeing peer-authored problems were told a peer wrote them and half were not, making peer authorship itself an experimental manipulation rather than only a design feature.

Study design and measures

The work is a mixed design reported across several partially overlapping samples rather than one trial. Development was iterative and qualitative in spirit: six structured co-design sessions over 15 months in which students interacted with AMPT and then answered questions about their experience and feature preferences, with prompts and features revised between rounds. Four of those sessions (n = 43 students from the four metropolitan districts) added survey batteries measuring perceived ability, interest and value, and sense of belonging before and after interacting with AMPT, with the belonging findings reported in detail by Norberg and colleagues in 2025. A wider post-session feedback survey captured the culture and interest representation items (n = 73), the preference and Arena items (n = 92, 83 and 41), and the naming item (n = 73), all on 5-point Likert scales with 5 as strong agreement. The scaled component is observational and correlational rather than experimental: 71 approved co-authored problems were embedded in MATHia, ratings were voluntary, and the authors report ratings given before a student began solving a problem specifically so that difficulty could not colour the perception. ethics approval came from an external committee, parents consented and minors assented, and the study was funded by the Gates Foundation.

Results on learning, interest and belonging

The headline attitudinal result is directional rather than global. Compared with traditional problems, students preferred their own and their peers' problems, wanted them deployed in their software, and rated the characters' names as central to the experience, all with means above 4.3 on a 5-point scale. What did not move was more telling: interest and value in mathematics rose only marginally, 3.56%, p = .08, and perceived ability was flat. The one significant pre-to-post change was a 4.79% rise in sense of belonging, Cohen's d = 0.48, p = .01, concentrated in the Membership and Acceptance subscales, the parts of the construct closest to belonging and being valued. The authors note that belonging may precede interest and that stable perceptions of mathematics plausibly need sustained rather than single-session effort to shift.

The scaled data are a more honest picture of fine-grained personalisation. Overall 61% of rated problems earned a thumbs up, students declined to rate 61% of the time, and peer-authored problems were simultaneously among the best and the worst received: 8 of 10 of the most-liked problems in algebra and probability, but also the least-liked probability problem, which used a student's own name and their teacher's name and earned 51% approval from 5 stars of self-rating. The problems students loved most (68%, 69% thumbs up in probability; 62%, 63% in algebra) share what the authors call a coherent narrative where the student's background motivates the mathematics rather than decorating it, whereas the weakest cases either rested on a specific name the author cared about or introduced a topic that never entered the math.

Personalisation, grain size and student agency

The paper frames its contribution through three moderators it borrows from Walkington and colleagues: depth, the extent to which the story integrates with the student's interests; grain size, whether a problem is pitched at a broad category like sports or at one student; and ownership, how much control the student feels. AMPT is an attempt to reach the far end of all three at once, most clearly visible in the problems students wrote about eating Ro-Tel at a named local mall, a principal and a math teacher, the Kanto region of the Pokémon universe, and regional dishes like fufu and poi. The topic data complicate the assumption that connecting maths to careers and lived experiences is the route to utility value: students chose Careers for only 4% of problems and gravitated instead to Fantasy Worlds, Music, Video Games and Food, and even when they picked no topic at all their scenarios often went fantastical (flying talking pickles, a knight rescuing a princess from a dragon, martians).

The student agency story has two layers. One is the ownership students exercised over the context, including character names, topics and freeform revisions, and the authors read the belonging gain as most plausibly rooted in representation, membership and acceptance rather than in generic enjoyment, while conceding that without a control group other factors cannot be excluded. The other is an honest weakness the authors name themselves: problem posing in AMPT is "currently more of a creative writing exercise than a mathematics learning opportunity", because students never engage with the formula or numbers. That protects them from the frustration problem posing can produce, but it also lets errors through, illustrated by a llama problem where the animals roar. Because the authors believe depth requires the context to be integrated with the mathematical construct, they propose future Scaffolding that exposes students to a light version of the constraints, for example that both actors must share an activity or that the problem concerns change over time, so that students select interests that connect to the maths without needing deep prior knowledge of the domain.

Limitations and open questions

Several limitations qualify every claim here. The attitudinal study has no control group, so the belonging shift is causally ambiguous even though it appeared across multiple groups and at different points in the semester. The dose is small: 30-minute sessions, short conversations (typically four exchanges) and, in the survey sample, a single encounter, so the authors themselves note that personalisation interventions are usually too brief to reliably build durable interest, and that stable perceptions may need sustained engagement. Novelty is a live alternative explanation for both the positive ratings and the belonging rise, particularly given that students often reacted to their own names appearing in a maths problem. The deployment results come from a pilot of only 71 problems across two domains, with voluntary ratings that 61% of students skipped, a response profile that could bias the thumbs-up proportion in either direction; the earlier error-rate result rests on a small preliminary pilot reported elsewhere. Finally, the tool is model-dependent. Reported sessions used OpenAI's GPT-4o or GPT-4.1, with o3-mini selecting among candidate problems, and AMPT's panel supports most OpenAI and Anthropic models, so behaviour, verbosity, bias and problem quality will vary with the model in use and with the prompt revisions that are still ongoing. Added to these is a scaling tension the authors do not resolve: the very features that make a problem meaningful to its author, an unfamiliar local dish, a low-frequency word or a teacher's name, can make it useless or harder to read for the students who later meet it.

Connected Concepts

  • Personalized Learning — the design goal, pursued here by individualising the story context rather than the task sequence
  • Learner Agency — student control over topic, characters, names and revisions, the paper's central explanatory construct
  • Generative AI — the technology that makes interest-matched problem generation feasible at scale
  • Large Language Models (LLMs) — the writing, clarifying and selecting agents behind candidate problems
  • Conversational AI — the chat interface through which AMPT elicits interests
  • Motivation — situational interest theory and utility value as the motivational frame for context personalisation
  • Student Engagement — the outcome personalisation is meant to move, measured here through liking and belonging
  • Learner Identity — representation, membership and acceptance as the belonging components that shifted
  • Math Education — story problems, reading demands and the discipline context of the study
  • Intelligent Tutoring — MATHia as the deployment vehicle and the source of large-scale rating and error data
  • Culturally Relevant Pedagogy — the representation gap students reported and the equity rationale for AMPT
  • Problem Solving — the dual linguistic and numerical processing demands of math story problems

Connected Articles

Citation

Norberg, K., Murphy, A., & Ritter, S. (2026). AMPT: A tool for personalizing math learning with generative AI. In GenAI in novel educational applications. Preprint; the peer-reviewed version is forthcoming.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.