On this page

Synthesis: LaTA is a privacy-preserving autograder that grades LaTeX homework with a locally hosted open-weight reasoning model (gpt-oss:120b on a single Mac Studio), so no student work leaves the instructor's machine — the FERPA problem that blocks most automated grading deployments disappears rather than being managed. Deployed across a full 200-student mechanical-engineering course, its instructor-confirmed error rate held at roughly 0.02-0.04% per rubric line item, and the author reports better exam performance and large self-assessed confidence gains against his previous traditionally graded cohort. The paper is notable for how carefully it refuses to over-claim: the exam gain bundles three changes at once, and the author says so explicitly rather than attributing it to the autograder.

Overview

Upper-division STEM coursework with derivations is expensive to grade: a term of assignments and exams routinely consumes hundreds of TA hours of first-pass grading. The obvious fix — send student work to a commercial LLM API — runs into FERPA and institutional data-governance rules, because student educational records end up at a third-party provider. LaTA takes the other route: run an open-weight reasoning model on hardware the instructor owns, and use instructor-authored rubrics with reference solutions to keep the grading decision interpretable and auditable.

The paper is a deployment and program-evaluation study rather than a benchmark. Its contribution is showing that the local route can carry a full-replacement deployment at real course scale — every submission of every homework graded by the pipeline, not a subset — while keeping the cost, privacy, and reliability story honest.

Study Design & Method

  • System architecture. A four-stage pipeline (ingest → segment → grade → report) grading LaTeX-native submissions against YAML rubrics with binary per-item scoring and instructor-authored reference solutions. A regex segmenter splits top-level chunks and falls back to gpt-oss:20b when it cannot; the grader is gpt-oss:120b, with responses validated against a strict Pydantic schema and prompts that wrap student text in untrusted-input delimiters.
  • Hardware and cost. All inference ran on one Apple Mac Studio (M3 Ultra, 256 GB unified memory) in the instructor's lab, with gpt-oss:120b and gpt-oss:20b the only active models. Cost is $0 marginal per assignment; per-submission grading took 1-3 minutes, aggregating to 4-8 hours of wall-clock time per homework set for the whole cohort.
  • Course and enrollment. ME 373 at Oregon State University, Winter 2026: eight homework sets across weeks 1-9 with an enrollment of about 200 students. Homework submitted in LaTeX and graded end-to-end by LaTA; instructor rubric authoring took 30-60 minutes per homework once the binary decomposition was internalized.
  • Corrections and disputes (two tiers). Tier 1 was a per-assignment corrections pass: 90% of students submitted corrections, and because corrections mode regrades the entire resubmission rather than a diff, a student who fixed one problem also had the rest regraded. Tier 2 was a Gradescope regrade request handled by the instructor — roughly 5-10 requests per assignment across the quarter.
  • Evidence streams (three, deliberately triangulated). Operational data plus a regrade audit; an anonymous post-term student survey (N = 159) with Likert items and free responses; and a quasi-experimental between-cohort exam comparison against the same instructor's Winter 2025 cohort.
  • Comparison cohorts. Winter 2026 enrollment 200 with 182 sitting the final exam, against Winter 2025 enrollment 181 with 157 sitting the final. Same instructor, textbook, weekly schedule, and exam structure; about two-thirds of exam problems were held identical and the replacement third was judged slightly harder, which the author treats as biasing the comparison against the new cohort. No inferential statistics are reported for this comparison, by design.
  • Confidence instrument. Block 1 measured pre/post confidence on four course-level learning objectives with 5-point Likert items, collected as a single-administration retrospective pre-test. Because the survey was anonymous, pre and post distributions were treated as independent and tested with Mann-Whitney U rather than a paired test.

Key Findings

  • Reliability was high in operational terms. Across the quarter, instructor-confirmed per-rubric-item error rates held at approximately 0.02-0.04%, derived from the regrade audit: every submission passes through the grader three to six times, and only about 5-10 regrade requests per assignment reached the instructor.
  • The privacy claim is structural, not procedural. No component of the grading pipeline sends student work off the machine, so the FERPA problem is removed by architecture; the paper releases the code under AGPLv3.
  • Feedback volume funded other teaching. The TA hours released by autograding were redirected into office hours, with Winter 2026 TA office-hour coverage roughly 3× that of Winter 2025 — an operational consequence that matters independently of grading accuracy.
  • Exam performance moved in the intended direction. The LaTA-graded cohort outperformed the previous one by approximately 11% on the midterm and 8% on the final exam.
  • The exam delta is a composite effect, and the author says so. Three changes were bundled between cohorts — LaTA autograding with LaTeX-native homework, the corrections workflow, and tripled TA office hours — and a single year of post-hoc data cannot separate them. The paper makes the composite attribution explicit and states that disentangling the three requires a multi-section or multi-year replication.
  • Confidence gains were large on every objective. Survey responses (N = 159) showed differences of at least +1.49 Likert points on every stated learning objective, with p < 10⁻²⁷ on every comparison.
  • The confidence instrument has a known bias. A retrospective pre-test avoids response-shift bias but tends to inflate apparent gains, and the unpaired Mann-Whitney analysis is conservative only in the p-value sense. The author argues the magnitudes should not be read as clean pre/post differences, while noting the between-cohort exam delta moves the same direction.

What this means for practice

  • Instructors. Author the rubric yourself in binary per-item form with a reference solution for each item before grading: the paper attributes the 0.02-0.04% instructor-confirmed error rate per rubric line item to rubric design, not to model quality.
  • Instructors. Budget 30-60 minutes of rubric authoring per homework once the binary decomposition is internalized, and plan for a corrections pass — 90% of students submitted corrections, and because corrections mode regrades the entire resubmission rather than a diff, fixing one problem re-exposes the rest of the work to new grading errors.
  • Software developers. When FERPA or data-residency rules bar a cloud grader, run the open-weight model on hardware the instructor owns: gpt-oss:120b on a single Mac Studio (M3 Ultra, 256 GB unified memory) carried a 200-student course at $0 marginal cost, with 1-3 minutes of wall-clock time per submission.
  • Instructors. Redirect the TA hours that autograding releases rather than treating them as savings — with the freed time the course ran roughly 3× the TA office-hour coverage of the previous year, which is an instructional gain independent of grading accuracy.

Limitations

  • The deployment is a single instructor, single course, single year: Winter 2026 ME 373 (about 200 students enrolled, 182 sitting the final) is compared with the same instructor's Winter 2025 cohort (181 enrolled, 157 sitting the final), so nothing here separates the tool from the course or the year.
  • The exam improvement — approximately 11% on the midterm and 8% on the final — bundles three changes (LaTA autograding, the corrections workflow, and roughly tripled TA office hours), and the author states that one year of post-hoc data cannot disentangle them.
  • Confidence results come from a retrospective pre-test on an anonymous survey (N = 159) analyzed with unpaired Mann-Whitney U tests; the design avoids response-shift bias but tends to inflate apparent gains, so the author says the +1.49-point magnitudes should not be read as clean pre/post differences.
  • No inferential statistics are reported for the between-cohort exam comparison, by design, and the author names the single-coder thematic analysis of free responses as a limitation of the qualitative evidence.

Citation

Rodríguez, J. A. (2026). LaTA: A drop-in, FERPA-compliant local-LLM autograder for upper-division STEM coursework. Submitted to Computers & Education.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.