Research Article
Mechanics Cognitive Diagnostic: Testing Fine-Grained Learning Objectives in Introductory Physics
Synthesis: Le and colleagues build the Mechanics Cognitive Diagnostic (MCD-v2), a cognitive-diagnostic computerized adaptive test that reports student mastery of fine-grained learning objectives during instruction, and test whether the physics assessments courses already administer can supply its item bank. Working from research-based assessments (RBAs) — the 30-item Force Concept Inventory, the 47-item Force and Motion Conceptual Evaluation and the 25-item Energy and Momentum Conceptual Survey — they defined 14 learning objectives from OpenStax textbooks and AP Physics standards under evidence-centered design, mapped items to objectives with a Q-matrix, and refined that mapping with the DINA cognitive-diagnostic model on 24,394 posttest responses drawn from 807 courses at 79 institutions. The FCI and EMCS fit the model well, the FMCE only marginally, and classification accuracy met or exceeded low-stakes formative benchmarks for 19 of the 22 objective–assessment combinations that had data. The result is a working 14-objective item bank assembled from instruments courses already use, plus a precise account of where the evidence is weakest — three nested energy objectives that share most of their items and violate the DINA model's independence assumption — and a roadmap to a 35-objective diagnostic covering a full semester at roughly two objectives per week.
From Retrospective Inventories to In-Instruction Diagnosis
The RBAs at the center of physics education research were built to measure learning after instruction, and that design has become their limitation. The FCI, FMCE and EMCS are fixed-length pretest–posttest instruments: a pretest says how prepared students were at the start, a posttest says how the semester went, and neither tells an instructor which specific ideas need attention next week. Black and Wiliam's formative-assessment tradition — effect sizes of 0.4 to 0.7 across more than 250 studies, among the largest in educational research — requires exactly what this format withholds: timely, actionable evidence about student thinking that instructors can act on and students can use, rather than Feedback arriving after the students who generated it have finished the course.
Two further obstacles stand between an RBA score and usable diagnostics. First, administering a fixed-length instrument more often is impractical without a large item bank, since repeated exposure to the same items degrades both validity and test security — the problem Yasuda et al.'s shorter, chained adaptive tests address. Second, and more fundamental, no agreed mapping exists from individual RBA items to specific constructs. Exploratory factor analysis failed to recover the FCI authors' own proposed sub-categories, and subsequent work has proposed structures ranging from two to nine dimensions without consensus, so a score cannot reliably indicate which topic a student has missed.
The authors' answer is to sidestep the factor-analytic question and adopt an instructional structure instead: learning objectives (LOs), defined to align with the goals and pacing that organize a physics course. Standards-based and competency-based grading already organize introductory physics this way, and one implementation raised course grades while lowering D and F grades and withdrawals, with the largest gains for women and first-generation students. When diagnostic feedback is expressed in the same objectives that guide instruction, instructors can act on it directly. The MCD is framed as the first cognitive diagnostic computerized adaptive test (CD-CAT) in physics; its predecessor, MCD-v1, mapped the same three assessments onto four broad skills (vectors, conceptual relationships, algebra, visualizations) and delivered mastery feedback during the semester, but those four attributes lumped conceptually distinct reasoning processes together, so a low "visualisation" score could not say whether graph reading or force-diagram interpretation was the problem.
Defining 14 Objectives and Mapping the Item Bank
The 14 objectives were built from OpenStax university and AP physics textbooks and from the AP Physics 1 and AP Physics C – Mechanics standards, sized at approximately two per week of instruction so that a quiz covering one week's objectives needs about 12 items and 12 minutes — brief enough to fit inside regular class time. They span vectors; one- and two-dimensional kinematics; free fall; free-body diagrams; Newton's second and third laws; kinetic and potential energy; work and conservation of energy; and momentum, impulse and momentum conservation. An expansion to 35 objectives is planned.
Coding was done by three researchers with physics and physics-education backgrounds, with at least two coding each item independently and all three resolving discrepancies by discussion. The result is a Q-matrix — a binary item-by-objective table specifying which objectives an item requires — which serves as the study's evidence model. The DINA model then assumes a conjunctive relationship: a student must have mastered every objective an item requires to answer it correctly, with slipping and guessing parameters absorbing real-world inconsistency. Because DINA makes no allowance for one skill compensating for another, and because CD models in general lose parameter recovery and classification accuracy as attributes multiply (most applications stop at three to eight), the authors' decision to work with 14 objectives is itself the study's test: can diagnostic specificity be bought without giving up accuracy?
Data came from the LASSO platform, which administers RBAs online, scores them and returns results to instructors. From 69,449 responses the authors removed completions under five minutes and retained each student's first attempt on the most recent posttest, leaving 24,394 posttest responses — 15,371 FCI, 7,033 FMCE and 1,990 EMCS — from 315 algebra-based, 460 calculus-based and 32 other courses across 79 institutions. FCI item 29 was excluded for poor psychometric performance, consistent with prior factor analyses. Three-parameter logistic IRT models were also fitted with the mirt package, but only to document item statistics for replication; DINA remains the measurement model. Model fit was assessed with RMSEA2 (≤ 0.05 indicating good fit) and SRMSR (≤ 0.07), and per-objective classification accuracy was judged against cognitive-diagnostic benchmarks of ≥ 0.9 good and ≥ 0.8 acceptable for low-stakes formative use.
Key Findings
- Content-expert coding was largely confirmed by the data: across the three assessments only 104 of 754 possible item–objective codings (14 percent) were even suggested for revision by DINA, and the coders adopted 20 of them, changing just 2.7 percent of all codings. Adoption was uneven — 13 percent of suggestions for the FCI, 14 percent for the FMCE and 33 percent for the EMCS.
- The one documented example of a data-driven correction is instructive: FCI item 1 (comparing fall times for balls of different mass) was initially mapped to Free Fall alone, and DINA flagged Newton's Second Law as additionally required — a change the coders adopted once they agreed that rejecting the "heavier falls faster" expectation requires recognizing that gravitational force scales with mass while acceleration does not.
- Coverage was broad but uneven: FCI items mapped to 7 of the 14 objectives, FMCE items to 8, and EMCS items to 7. Across all three, Newton's Second Law (36 items) and 1D Kinematics (33) were best covered, followed by Free-Body Diagrams (18), Vectors (16) and Newton's Third Law (15), while energy and momentum objectives drew 6–14 items each, largely from the EMCS.
- Most items are diagnostically simple: 38 items (38 percent) assessed a single objective, 39 (39 percent) assessed two, 20 (20 percent) assessed three, and only four items required four or more. The EMCS supplied all of the most complex items — two requiring four objectives, one requiring five and one requiring six — complexity absent from the FCI and FMCE.
- Model fit was good for the FCI (RMSEA2 = 0.033, SRMSR = 0.055) and the EMCS (RMSEA2 = 0.022, SRMSR = 0.037), but marginal for the FMCE (RMSEA2 = 0.065, SRMSR = 0.087), which fell outside both criteria.
- Classification accuracy was adequate or better for 19 of the 22 objective–assessment combinations with data (42 combinations are possible but each assessment covers only a subset): 13 reached the good threshold (≥ 0.90) and six the acceptable band (0.80–0.90).
- The three exceptions were all EMCS energy objectives — Potential Energy (0.675), Conserve Energy (0.705) and Kinetic Energy (0.745) — while the EMCS momentum objectives, drawing on the same instrument and a similar number of items, reached acceptable-to-good accuracy (Linear Momentum 0.860, Impulse 0.820, Conserve Momentum 0.917).
- The finer objective structure improved fit over MCD-v1's four broad skills on every assessment: FCI RMSEA2 fell from 0.048 to 0.033, EMCS from 0.028 to 0.022, and FMCE from 0.090 to 0.065 — with the FMCE remaining the sole marginal-fit instrument under both independently coded attribute structures.
- The combined bank covers all 14 objectives with at least six items each, meeting the four-to-six-items-per-attribute minimum that Simulation work associates with stable DINA estimation, and it mixes single-objective items (precise targeting of isolated mastery) with multi-objective items (evidence about knowledge integration) more richly than any single RBA.
Why the FMCE and the Energy Objectives Underperform
Both weaknesses are traced to instrument design rather than modeling artifact, and the diagnoses are unusually specific.
The FMCE's marginal fit has a structural cause: 42 of its 43 scored items are grouped into chained or blocked sets that share a scenario stem, versus 13 of 30 for the FCI and none at all for the EMCS. Item chaining creates local item dependence within a block, which violates the DINA model's assumption of conditional independence given a student's mastery profile. Prior work also finds FMCE items engaging more overlapping dimensions than FCI items, so the misfit reflects the instrument rather than the coding.
The energy objectives' low accuracy has two separable causes. The first is item noise: the mean DINA slipping parameter is 0.349 on the EMCS against 0.138 on the FCI and 0.136 on the FMCE, meaning students the model classifies as having mastered an item's objectives still miss it often, which weakens every response's evidential value. The second is that item noise alone cannot explain why energy objectives fail while momentum objectives succeed on the same instrument. The distinguishing factor is item overlap. Eight EMCS items require all three energy objectives, and any two of them share roughly 70 percent of their items (Jaccard overlap 0.67–0.73), whereas any two momentum objectives share only 18–33 percent and just one item requires all three. When objectives share most of their items, few responses depend on one objective without the others, so the model cannot separate mastery of kinetic energy from mastery of potential energy or conservation. The authors argue this reflects the conceptual nesting of the objectives — conservation of mechanical energy presupposes both kinetic and potential energy — meaning the items violate DINA's conjunctive, independent-attribute assumption rather than the coding being wrong.
Sample size compounds the structural problem. DINA slipping and guessing estimates stabilize only in large samples, typically well over 2,000 responses, and the EMCS contributed the smallest dataset at 1,990, so its estimates are the least stable of the three.
The paper considers a range of responses rather than claiming the problem solved. Larger EMCS samples, which will accumulate through LASSO, should stabilize the noise component. The structural component needs a more flexible model: the generalized DINA model relaxes the all-or-nothing requirement so partial mastery of required objectives raises the probability of a correct response, and hierarchical diagnostic classification models go further by encoding prerequisite orderings — kinetic and potential energy as prerequisites for conservation of energy. Both are more demanding to estimate (G-DINA fits up to 2^Kj parameters per item requiring Kj objectives) and will become feasible as energy-objective data accumulate. Merging the three energy objectives into one was considered and rejected on three grounds: the three match the source texts' chapter-pacing (the energy unit spans about two weeks and four objectives), instructors still need a mastery decision for each concept even when teaching them together, and the overlap is a property of the current item bank rather than of the objectives themselves — the planned expansion to 35 objectives will add energy-specific items and should reduce it.
What this means for practice
The headline contribution is a reframing of what existing instruments are for. RBA items were not written for objective-level diagnosis, yet a Cognitive Diagnosis model applied to them yields diagnostic inferences that no single instrument can produce, which means courses can gain fine-grained feedback without adopting an unfamiliar assessment or authoring new items. For instructors, the practical implication is that a cognitive diagnostic — not a factor-analytic score — is what converts an existing concept inventory into something usable mid-semester.
For the measurement community, the study is a clean demonstration that psychometric validity for diagnostic claims is local, not global: an instrument can be a good measure of overall mechanics understanding (the EMCS fits DINA well) and simultaneously a poor vehicle for separating three specific objectives (accuracy 0.675–0.745), and the diagnosis of which objectives and why is the deliverable. The finding that fit improved for every assessment when the same items were recoded against a finer objective structure also bears on a long-running dispute: to the extent that factor-analytic disagreement about RBA dimensionality reflects under-specified constructs, an instructional structure chosen a priori can outperform explored structures for adaptive delivery. Two cautions travel with it, both stated by the authors: the assessment is currently framed for low-stakes formative use rather than high-stakes decisions, and the diagnostic feedback is only as fair as a DIF analysis that has not yet been run.
Limitations
The limitations are framed as gaps in the evidence for the student models. DINA cannot represent relationships between objectives, which is precisely what the nested energy objectives require. The analysis uses posttest data only, so it cannot evaluate mastery development during instruction — and the authors stress that extending to multiple time points requires establishing measurement invariance first, because item slipping, guessing or objective requirements can shift between administrations, and apparent gains could otherwise reflect changed item behavior rather than learning. No differential item functioning analysis was run across demographic groups, even though prior work documents DIF on several FCI items and a five-factor analysis attributing much of it to real knowledge differences rather than bias; fairness of objective-level diagnosis across populations is an open question.
Three practical questions are left untested altogether: whether instructors can interpret objective-level feedback, whether it changes their teaching decisions, and whether students who receive it learn more. The authors make a design case for each — objectives aligned with weekly pacing map onto decisions instructors already make, the bank draws on assessments courses already administer, and the LASSO platform already reaches hundreds of courses — but treat the claims as hypotheses. Future work also needs to close coverage gaps for the remaining 21 planned objectives, for which the FCI, FMCE and EMCS hold no items at all (rotational mechanics, mathematical reasoning), with candidate instruments already identified, and to add new items through online calibration that embeds unscored items alongside the operational bank so the item bank can grow without pausing testing.
Connected Concepts
- Cognitive Diagnosis — DINA-based mastery classification on 14 learning objectives
- Physics Education — first CD-CAT in physics, built from the FCI, FMCE and EMCS
- Educational Measurement — evidence-centered design, Q-matrix validation and model-fit evidence
- Item Response Theory — 3PL calibration used alongside the cognitive-diagnostic measurement model
- Formative Assessment — the design goal is actionable feedback during instruction, not after it
- Automated Assessment — algorithmic scoring that returns an objective-level mastery profile
- Adaptive Learning — adaptive item selection matched to the student's estimated state
- STEM Education — mechanics as the testbed for competency-based diagnostic assessment
- Assessment Validity — student models and evidence models tested for separable, recoverable objectives
Connected Articles
- Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach — Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
- Bayesian cognitive diagnosis optimizes personalized learning paths via mediation of cognitive load and Hidden Markov Model state transitions — Bayesian cognitive diagnosis optimizes personalized learning paths via mediation of cognitive load and Hidden Markov Model state transitions
- Beyond ID Embeddings: Process-Grounded Language Modeling for Cognitive Diagnosis — Beyond ID Embeddings: Process-Grounded Language Modeling for Cognitive Diagnosis
- The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory — The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
- Applying IRT to distinguish between human and generative AI responses to multiple-choice assessments — Applying IRT to distinguish between human and generative AI responses to multiple-choice assessments
- Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study — Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study
- Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results — Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results
- Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention Study — Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention Study
- Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes — Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes
- A Decade of Reflection and Thematic Review on Artificial Intelligence's Impact on Educational Measurement — A Decade of Reflection and Thematic Review on Artificial Intelligence's Impact on Educational Measurement
Citation
Le, V., Nissen, J. M., Morphew, J. W., Chang, H. H., & Van Dusen, B. (2026). Mechanics Cognitive Diagnostic: Testing Fine-Grained Learning Objectives in Introductory Physics. arXiv preprint arXiv:2609.09584.