Research Article
REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
Synthesis: REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models advances the Automated Grading frontier by solving a fundamental trust problem: even accurate AI graders are unusable if educators cannot verify their reasoning. Standard Large Language Models (LLMs)-based graders operate as black boxes, while earlier Concept Bottleneck Models (CBMs) offer interpretability but fail at modeling rubric dimensions, ordinal score semantics, and noisy human annotations. REC-CBM introduces three innovations: (1) a rubric-aware concept encoder that learns concept-specific representations aligned with actual grading rubrics, (2) an ordinal pairwise calibration objective that preserves score ordering (e.g., 'poor' < 'fair' < 'good'), and (3) a latent error-correction module that denoises concept predictions while maintaining full interpretability. Experiments demonstrate consistent improvements in both grading accuracy and concept-level reasoning faithfulness over baselines. This work directly addresses Assessment Validity concerns raised in Generative AI as a Design Variable: An Evidence-Centered Framework for Principled Governance in STEM Assessment and complements Confidence Estimation in Automatic Short Answer Grading with LLMs by adding the interpretability dimension. The rubric-aware design aligns with Formative Assessment needs and Scaffolding principles, and the error-correction approach resonates with work on Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education.
What this means for practice
- Instructors. Adopt a rubric-aligned grading model where grades must be defensible: REC-CBM reached the best task accuracy and macro-F1 within every backbone block across three benchmarks and the best concept F1 in 14 of 15 model–dataset cells, while exposing every rubric dimension it used.
- Instructors. Audit at the dimension level, not the whole grade: substituting lower-confidence concept labels moved the predicted grade downward (wrong and random labels cost 30–45 accuracy points at the largest k), while oracle labels preserved or improved it, so a single rubric judgment can be inspected and overridden.
- Instructors. State your rubric as explicit numbered dimensions with ordinal levels before deployment — the concept encoder learns one head per dimension, and the paper's rubrics use 7–8 concepts on three- or five-level ordinal scales.
- Researchers. Expect the concept-supervision signal to be load-bearing: removing the rubric-aware concept encoder was the most damaging single ablation, and grading quality saturates by roughly four concepts and encoder width of 384 or more.
- Researchers. Treat auto-generated concept labels as measurements with error, not as ground truth — the error-correction module exists precisely because annotator disagreement, overlapping dimensions, and open-ended ambiguity make rubric labels noisy.
Limitations
- The three grading benchmarks (Mohler, 2,273 samples; ASAP 2.0, 17,292; MOCHA, 31,069) ship no concept-level annotations, so all concept labels were generated by GPT-4o and Gemini-2.5-pro and then reviewed by only three domain experts — the bottleneck's supervision inherits that annotation error.
- The educator intervention study is a simulation: oracle, wrong, and random concept labels were substituted programmatically on a frozen second-stage head with no real instructors grading in a loop.
- Evaluation covers three English benchmarks with seven to eight rubric concepts each, and the authors leave multilingual assessment, domain-specific rubrics, and educator-in-the-loop rubric refinement to future work.
- The framework is sensitive to tuning: the learning-rate search had to be restricted to a narrow 1e-5 to 1e-4 band because larger rates destabilize rubric-aware token attention before calibration can take effect.
Citation
Chengshuai Zhao, Fan Zhang, Kumar Satvik Chaudhary, Yiwen Li, Lo Pang-Yun Ting, Ying-Chih Chen, Huan Liu (2026). REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading. arXiv preprint.