On this page

Synthesis: Across the centuries-long migration of mechanical work from human to tool — fingers → pencil → calculator → Python → agent — Engelhardt argues that what a learner must know is tool-invariant. Five pillars (inputs/outputs, method concept, terminology, sensemaking, operating the tool) stay stable while only their content and weight shift. Because AI-generated simulations are opaque and bespoke (validated by no community), verification becomes the load-bearing skill and the artifact no longer certifies the student. The constructive response pairs AI-free in-class coding quizzes with oral defenses of comment-stripped, AI-assisted work, gated on a verification dimension — design prescriptions that largely await Fall 2026 cohort validation.

Note on type: This is a framework / position essay (first-person, draws on the author's teaching practice), not a controlled empirical study. Its claims are argued and illustrated, with open questions explicitly flagged as untested. It is tagged confidence: high for fidelity to the source and internal rigor, but the design prescriptions are the author's and largely await validation (Fall 2026 cohort).

Summary

Agentic AI — systems that write, run, and revise simulation code from natural-language specs — is the latest step in a centuries-long migration of mechanical work from human to tool. Engelhardt argues that what a learner must know is tool-invariant: across tools (fingers → pencil → calculator → Python → agent), the requirements are stable; only their content and weight shift. The paper organizes these into five pillars, argues that sensemaking/verification is now the load-bearing skill (because AI-generated artifacts are bespoke and unvalidated, unlike socially-validated libraries), and draws the Assessment consequence: when artifacts can be generated on demand, the artifact no longer certifies the student. The proposed response: AI-free in-class coding quizzes (measure white-box residue) + oral defenses of comment-stripped, AI-assisted work (measure orchestration), with a verification gate that must pass regardless of total score.

Key Findings

  1. Tool-invariance of the five pillars. Competent use of any computational method at any scale requires inputs/outputs, a method concept, terminology, sensemaking, and tool operation. Content changes most for tool operation (penmanship → prompting) and least for sensemaking, but weight moves the other way: sensemaking becomes load-bearing under agentic AI.
  2. Terminology is now an input channel. When a tool accepts natural language, every choice a vague specification leaves open is silently ceded to the tool — "integrate this ODE with adaptive Runge–Kutta and verify energy conservation" and "make the ball bounce right" yield categorically different artifacts.
  3. Tool operation has inverted shape. Operating NumPy required narrow, deep skill; directing an agent requires broad, shallow "interactional expertise" (conversing, judging plausibility, directing work) across many ecosystems.
  4. Verification is newly load-bearing. A library FFT is opaque but socially validated; an AI simulation is opaque and bespoke — a population-of-one artifact whose failures arrive disguised as successes. The verification burden ecosystems amortized now lands on each student, for each artifact, every time.
  5. Principle of validation authority. Delegation is safe exactly where the delegator retains the ability to validate the outputs. "Anyone can build software with AI" holds only where correctness is observable in use (dashboards, apps); a physics simulation's correctness must be established by disciplinary checks.
  6. Two-instrument assessment. AI-free in-class coding quizzes (white-box residue) plus ten-minute oral defenses of comment-stripped AI-assisted work, scored on a five-dimension rubric with a verification gate that must reach "functional" regardless of total.

The five pillars (tool-invariant)

  1. Inputs and outputs — what goes in, what comes out, conditions of validity; problem posing.
  2. Method concept — a working model of what the method does, including its knobs (step size, tolerance) and characteristic failure modes (instability, divergence, overfitting, aliasing).
  3. Terminology — precise disciplinary vocabulary; now an input channel, since vague natural-language specs silently cede choices to the tool.
  4. Sensemaking — judging whether outputs make sense and establishing correctness (verification/validation). Always present; becomes load-bearing under agentic AI.
  5. Operating the tool — the actuation skill: once penmanship, now directing an agent (broad, shallow "interactional expertise" rather than narrow deep syntax).

The pillars organize into a workflow: Specify → Predict → Delegate → Verify → Interpret; Iterate, threaded by calibrated reliance (how much verification is owed, given tool + task) — a competency that threads through trust in automation.

Why verification is newly load-bearing (the core argument)

Opacity was never the problem — validation is. A library routine (FFT, linear algebra) is opaque but socially validated (published algorithms, decades of testing, millions of users). An AI-generated simulation is opaque and bespoke: a population-of-one artifact from a stochastic process whose competence frontier is "jagged and invisible," whose failures arrive disguised as successes (running code, smooth plots, confident prose). The verification burden that ecosystems amortized across a community now lands on each student, for each artifact, every time.

Principle of validation authority: delegation is safe exactly where the delegator retains the ability to validate the outputs. A computational-physics course now exists to train validation authority. Corollary: "anyone can build software with AI" holds only where correctness is observable in use (dashboards, apps); a physics simulation's correctness must be established by disciplinary checks. This reframes AI Literacy for science as the trained capacity to specify checks and judge evidence — a form of Metacognition made concrete.

What must remain human (non-delegables)

Posing the problem · choosing & owning the physical model/assumptions · the pre-execution prediction · specifying the checks · final epistemic responsibility ("the AI said so" is never a justification). Items 1, 2, 5 are constitutive of doing science, not claims about current AI capability, and do not weaken as models improve. These map directly onto Assessment Validity concerns: the inference from artifact to student collapses precisely where these human elements are delegated.

Assessment design (the constructive response)

  • Proxy collapse: traditional "write code → submit report" grading died because the artifact no longer certifies the student (Goodhart's law; Kortemeyer's assessment alarm). Supervised formats survive.
  • The product is the student's ability to explain and defend artifacts in the discipline's language — certified, as at the Ph.D. level, by oral defense.
  • Two instruments: (1) AI-free in-class coding quizzes in a lockdown browser (assess the white-box phase / coding residue); (2) ten-minute oral defenses of AI-assisted work, with code comments stripped beforehand so understanding can't be performed by reading borrowed narration. The defense probes, live and adaptively: walkthrough of uncommented code, plot interpretation, and verification probes ("why should I believe this?", "what did the AI decide that you didn't?").
  • Verification gate: the rubric scores five dimensions (code comprehension, method understanding, physics model/terminology, interpretation, verification); the verification dimension must reach "functional" for the defense to pass, regardless of total.
  • Scalability: honest arithmetic — ~28 contact-hours of defenses/semester for ~15 students (less than grading 15 reports, more informative); degraded modes (spot-defenses, TA-led, paired) named with costs. "It doesn't scale" is "partly the point" — an equity concern for under-resourced institutions.

Teaching practices

  • White-box, then black-box (Buchberger): study a method transparently (hand-code the 15-line Euler integrator, watch it fail at large dt) before delegating. Dissolves the "must code vs. need not code" debate — both true at different rungs, tied to Scaffolding theory.
  • Error injection: give students subtly-wrong agent output (sign error, dt too large, wrong potential); grade the diagnosis. Trains reading code one didn't write — a core response to over-reliance.
  • Motivation over prohibition: a guardrailed tutor helped engaged students but was useless to answer-seekers ("an unguardrailed model is two browser tabs away"). Framing: homework is the gym, not the job; AI is a forklift at the gym. Design for Motivation, not bans — echoing Self-Regulated Learning and Help-Seeking theory.
  • Term "comprehension debt" (gap between code a system contains and code its maintainers understand) imported from software engineering as a risk of AI-assisted production, connecting to Cognitive Offloading.

What this means for practice

  • Instructors. Measure coding fluency only under AI-free conditions: weekly in-class coding quizzes in a lockdown browser capture the white-box residue once artifacts can be generated on demand.
  • Instructors. Measure orchestration separately with ten-minute oral defenses of comment-stripped AI-assisted work, probing uncommented code, plot interpretation and verification questions such as what the AI decided that the student did not.
  • Instructors. Gate the defense on verification: score the five rubric dimensions (code comprehension, method understanding, physics model and terminology, interpretation, verification), but require verification to reach the functional level for a pass regardless of total - a transferable model for keeping human accountability central when only one of the two instruments is used.
  • Instructors. Teach white-box before black-box by having students hand-code the 15-line Euler integrator and watch it fail at large step size, then grade the diagnosis of injected errors (a sign error, too-large dt, the wrong potential) instead of banning AI.
  • Instructors. Budget and defend the format with the author's arithmetic - ten minutes times fifteen students comes to two and a half contact-hours per defended assignment, and at roughly eleven defended assignments per term the course projects to some 28 contact-hours a semester, with spot-defenses, TA-led and paired modes as named degraded options - and treat the equity gap for institutions that cannot staff human-scale assessment as an argument to take to administrators.

Limitations

  • This is a framework/position essay, not an empirical study: the five pillars, the load-bearing claim about verification and the two-instrument assessment design are argued and illustrated, and the author's own prescriptions were still awaiting validation with students in Fall 2026.
  • The first-person evidence comes from the easy case - a week-long June 2026 PICUP workshop after which one professor built an LMS dashboard and a comment-stripping tool without writing the code - and the author notes that a dashboard's correctness is observable in use, unlike a physics simulation's.
  • The single empirical probe of the guardrailed tutor was an end-of-semester survey in a sophomore computational course of about fifteen students that drew exactly one response, which the author describes as one instructor's reading of one semester.
  • The scalability arithmetic is partly estimated: the ten-minute defense length is measured from a semester of per-assignment defenses in two courses, but the fifteen-student cohort and the roughly 28 contact-hours per semester are projections, and no plan is offered for a 300-seat course.

Citation

Engelhardt, L. (2026). A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.