Research Article
Simulating Learners' Task-Selection Strategies and System Constraints in Mastery Learning
Synthesis: Noh, Chowdhary, Ooge, Aleven and Borchers (2026) ask what happens when the shared control that intelligent tutoring systems increasingly grant learners collides with the efficiency that mastery-based progression is supposed to deliver. Rather than run the experiment on students, they simulate it: simulated learners driven by models fitted to two real tutoring datasets choose their own skills a thousand times over, under eight task-selection strategies and four system constraints. One strategy — Minimize Worst Case Loss, which picks whichever skill risks the least — produced roughly thirty times the overpractice of the others. A constraint that narrows the choice to skills just below mastery cut that by an order of magnitude, and left efficient learners untouched. The paper's case is methodological as much as empirical: simulation over logged interaction data is a cheap way to find out when and for whom guardrails on learner choice are worth the autonomy they cost.
Overview
Mastery learning is a promise about competence: a learner advances only after demonstrating they know a prerequisite. Intelligent tutoring systems operationalize it with performance-based criteria, and the cost is cumulative and easy to miss. Because a multi-step problem exercises several skills at once — the equation problems in this study average 5.67 skills each, the graph problems 8.57 — a problem chosen to practice one already-mastered skill drags redundant practice of others along with it. The paper calls the resulting waste overpractice: time on task that produces no proportional learning.
Set against that, ITS increasingly hand learners shared control over what to practice next. The motivation is well established, drawing on self-determination theory and supported by documented gains in engagement, perceived autonomy, metacognitive reflection and persistence. But prior research also shows that learners adopt a diversity of strategies — targeting their strengths, targeting their weaknesses, interleaving across skills, or blocking on one skill until it is done — and that some of these choices delay progression systematically. A choice architecture that is good for autonomy in general therefore need not be good for efficiency in particular.
The awkward question is what to do about that, and the authors note it is expensive to answer the obvious way. Testing corrective constraints in real classrooms is slow, risks exposing students to worse conditions, and requires re-running an intervention per constraint. So the paper substitutes simulation instead of live deployment: fit learner models to real interaction logs, let simulated learners make the same choices real ones make, and watch which strategies and constraints produce waste. The framing question is deliberately two-sided — how do different task-selection strategies affect mastery-based efficiency, and can system constraints repair the damage without harming the learners who were already fine?
Study Design & Method
Two datasets, two domains. Both come from PSLC DataShop, and both are real classroom data rather than synthetic corpora. The equation-solving corpus (sets #5549 and #5604) is IRB-approved middle-school data from the APTA tutoring system for linear equations: grades six to eight across two public schools in the eastern United States, 261 students, over 10,000 step-level interactions, averaging 5.67 ± 2.47 skills per problem. The graph-interpretation corpus (#5360) is classroom transaction data from the Mathtutor system, 97 ninth-grade students in a U.S. high school who alternated paper practice with Mathtutor across three units on linear graphs: 31 problems, 13 skills, and a denser 8.57 ± 8.41 skills per problem. The two differ deliberately in problem length and skill composition, which gives the framework a cross-domain test of whether a result is a property of the model or of the domain.
Two learner models. Step-level performance is generated by the Additive Factors Model, which predicts correctness as a logistic function of student ability plus additive skill difficulty and practice effects — the standard educational data mining view of learning in tutoring logs. Knowledge state is tracked with Bayesian knowledge tracing, a hidden Markov model over a binary learned/unlearned state updated by Bayes' rule after each response from prior knowledge, learning rate, guess rate and slip rate. Both are fitted to the real datasets, and the fitted parameters drive everything the simulated learners do; a skill counts as mastered once the posterior passes a 95 percent threshold. BKT defaults follow TutorShop: pinit = 0.25, plearn = 0.22, pguess = 0.2, pslip = 0.1. This is student modeling used in reverse — not to predict what students will do, but to generate plausible students who then make choices the researchers want to study.
The simulation loop. Each run lets 1,000 simulated learners progress through the skill set until mastery, in a cycle that mirrors a shared-control tutor: between problems the learner selects a skill according to a strategy (optionally narrowed by a constraint), the system selects a problem exercising that skill, the learner attempts the multi-step problem with per-step performance sampled from the model, and knowledge tracing updates after every step. Learners choose between problems but never within them — a boundary that matters for interpreting the whole study.
Eight strategies and four constraints. Four strategies come from the prior literature (Strength Targeting, Weakness Targeting, Interleaving, Blocking) plus Random as a control. Three outcome-informed rules were added: Maximize Usual Case Improvement and Maximize Usual Case Outcome, which both pick the highest projected post-practice mastery, and Minimize Worst Case Loss, a risk-averse rule selecting the skill with the smallest possible mastery loss. Constraints come in two families. Task-selection constraints narrow the selectable pool without prescribing a choice — closer-to-mastery restricts learners to skills just under the threshold, further-from-mastery to lower-proficiency skills. Problem-selection constraints bias which problem is delivered once a skill is chosen: prefer-easier weights lower-difficulty problems, prefer-harder weights higher-difficulty ones, sampled from a normalized multinomial. Everything is compared against an unconstrained baseline, and the simulation and analysis code is open source.
Outcome. Overpractice — continued practice of already-mastered skills, aggregated across all simulated learners and skills — is the primary metric, compared across conditions with means, standard deviations and Cohen's d.
Key Findings
One risk-averse strategy is the outlier, and it is not close. Unconstrained, Minimize Worst Case Loss produced overpractice of 30.28× in equation-solving and 28.93× in graph-interpretation, against means of 1.733 ± 0.383 and 3.223 ± 0.050 across the other seven strategies. The authors' explanation is structural rather than motivational: skills a learner already knows best sit closer to the mastery threshold, so choosing the safest-looking option repeatedly steers practice into territory where redundancy is most likely.
Weakness targeting is the efficient default. In the equation-solving domain, Weakness Targeting and Maximize Usual Case Improvement produced the lowest overpractice, 1.26× and 1.29×. Strength Targeting was the second-worst strategy there — a useful reminder that "practice what you are good at" and "practice what you are bad at" are not symmetric choices in a system that sequences by mastery estimates.
Constraints repair the worst case. Guiding loss-averse learners toward skills closer to mastery reduced overpractice from 30.28 to 1.83 in equation-solving (Cohen's d = 1.87) and from 28.93 to 3.22 in graph-interpretation (d = 1.86) — bringing the worst strategy down to the level of the good ones. The further-from-mastery variant still helped substantially, to 12.88 and 14.46, but plainly less. Problem-selection constraints worked as well: under Prefer Harder, Minimize Worst Case Loss fell from 30.28 to 1.98 (d = 1.86), again levelling it with strategies that needed no constraint at all, by raising exposure to more challenging problems.
The repair is selective, which is the design point. Blocking, Interleaving, Strength Targeting and Weakness Targeting held their overpractice levels steady across every constraint condition. The corrective effect lands on maladaptive behavior and nowhere else — so a system could apply the constraint without degrading the learners who were already choosing well, which is exactly the property a real adaptive guardrail would need.
Domain structure shapes the result. The equation-solving domain — multi-step, highly recursive, with problems admitting multiple solution paths and heavy skill repetition — showed greater variability in overpractice and a stronger response to Prefer Harder than graph interpretation did. Whatever the framework measures, part of it is a property of how the domain's problems decompose into skills.
What this means for practice
- Instructors. Screen for the risk-averse pattern in a shared-control tutor: a learner who repeatedly picks whichever skill risks the least generated 30.28 times the overpractice of the other seven strategies in equation-solving and 28.93 times in graph interpretation, so treat that choice signature as a signal to intervene rather than as harmless caution.
- Designers. Bound learner choice conditionally instead of uniformly: restricting the selectable skills to those just below the mastery threshold cut the worst strategy's overpractice from 30.28 to 1.83 times in equation-solving (Cohen's d = 1.87) and from 28.93 to 3.22 in graph interpretation (d = 1.86), while Blocking, Interleaving, Strength Targeting and Weakness Targeting held steady across every constraint condition.
- Designers. If narrowing the skill pool is impractical, bias problem delivery toward harder problems instead — Prefer Harder lowered Minimize Worst Case Loss from 30.28 to 1.98 times in equation-solving (d = 1.86) by raising exposure to more challenging problems.
- Researchers. Use the simulation-before-deployment template: fit an Additive Factors Model and Bayesian knowledge tracing to logged interaction data, then pre-test sequencing and constraint policies across more than one domain before any student meets them, keeping the simulated learners parameter-grounded rather than generative.
- Instructors. Treat the constraint as a human-in-the-loop design decision, not a hard system rule: because strategy diversity is real and consequential, constrain behavior that is maladaptive rather than the average learner, and stay alert to whether visible constraints erode the autonomy that motivated shared control.
Limitations
- No student took part in the experiment: every result comes from simulation runs of 1,000 learners whose behavior is generated by models fitted to two PSLC DataShop datasets — 261 middle-school students in equation-solving and 97 ninth-grade students in graph interpretation — so no observed learner behavior is tested.
- Task selection always comes first and never changes: learners re-select a skill between problems according to a static strategy, and the authors leave dynamic combinations of strategies to future work.
- Motivational, metacognitive and affective factors are deliberately excluded, so the model has no way to represent a learner who gives up, gets bored, or becomes more confident.
- Alternative paradigms are not modeled — direct problem selection, multi-task selection, or fully system-driven sequencing — and the two datasets cover two domains only, so the framework's behavior under other skill structures is untested, and by the authors' own account it remains open whether learners perceive differing constraint levels or whether visible constraints erode the autonomy benefits that motivated shared control.
Citation
Noh, H., Chowdhary, A., Ooge, J., Aleven, V., & Borchers, C. (2026). Simulating Learners' Task-Selection Strategies and System Constraints in Mastery Learning. arXiv preprint (v2, 25 May 2026).