📄 Research Article
A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol
Extends a prior validated protocol for classifying open-ended teaching-evaluation feedback by thematic category and sentiment, introducing a durability and cross-language transfer benchmark. Institutions collect far more teaching feedback than they read; automated classification makes it actionable.
Sits at the intersection of Feedback Loop, Automated Grading, and Teacher Role support in Higher Ed. It relates to Formative Assessment and the AI Feedback Quality stub, providing a reusable benchmark for scaling educator feedback analysis and connecting to Faculty Development.
Key Findings
Study Design & Method
The benchmark re-runs a previously validated protocol — built from a documented annotation guide, intra-annotator reliability measurement, stratified cross-validation, and held-out evaluation on a Spanish institutional corpus with a frozen-encoder design — on the original Spanish data across three representation generations: sparse lexical features, frozen transformer embeddings, and prompted large language models. The sentiment task is then transferred to English with a balanced 45,000-comment corpus contrasted with an aspect-labeled educational dataset. Paired comparisons are treated as descriptive rather than inferential: headline weighted-F1 differences are reported with descriptive bootstrap confidence intervals (95% percentile, 2,000 resamples), where an interval covering zero indicates no descriptive separation. Spanish arms use a 233-comment split and English LLM arms use a 1,500-comment subset.
Implications for AI in Education
Institutions collect far more open-ended teaching-evaluation feedback than they read, and automated classification makes that corpus actionable for improving teaching. The benchmark's central message is that a validated protocol can remain useful as representation methods advance: a frontier model wins only the hardest thematic task in Spanish, and on sentiment — in both languages — it shows no descriptive separation from economical alternatives. For Educational NLP and Benchmark practice, this argues for reporting paired comparisons descriptively and for treating model selection as a cost-performance deployment question rather than chasing frontier models by default. The LIME-based auditability check also demonstrates a lightweight way to keep classifications inspectable and aligned with the annotation guide, supporting responsible use of automated feedback analysis in faculty-facing systems.
Connected Concepts
Connected Articles
Citation
Esteban U. Vega Barajas (2026). A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol. arXiv:2607.11873. arXiv preprint.