Research Article
A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol
Synthesis: Re-running a previously validated Spanish teaching-feedback classification protocol across three representation generations — sparse lexical features, frozen transformer embeddings, and prompted LLMs — shows that the procedure survives as models change, while the choice of model is a cost-and-auditability deployment decision. The 2026 frontier model wins the hardest Spanish thematic task but buys no sentiment advantage over a cheaper model in either Spanish or English.
Institutions collect far more open-ended teaching-evaluation feedback than they read, and automated classification is what makes that corpus actionable for improving teaching. This work sits at the intersection of Feedback Loop, Formative Assessment, Automated Grading, and Teaching support in Higher Education, and it doubles as a reusable Benchmark for educational NLP — connecting to Educational Development, Educational Measurement, and the AI Feedback Quality stub.
Overview
The starting point is a 2026 thesis by the same author (Vega Barajas, 2026), which built a procedure an institution could rely on for a Spanish-language institutional corpus: a written annotation guide over four thematic categories and three sentiment classes, an explicit reliability measurement, stratified five-fold cross-validation for model selection, and a held-out test set left untouched to the end. With a Spanish BERT model (BETO) as a frozen feature extractor feeding classical classifiers, the prior study reached weighted F1 of 0.80 for thematic classification (macro 0.75) and 0.85 for sentiment — under three self-imposed constraints: a small labeled set, limited compute, and the need for an auditable pipeline. Peak accuracy was never the goal.
That result invites two objections, and the paper is organized around answering them. First, the method looks dated: representation methods have moved from sparse lexical features through contextual transformer embeddings to instruction-tuned LLMs, so is a frozen-2019 design still competitive or has it been superseded? Second, the protocol was validated on Spanish only — does it transfer to English educational text, and what changes when the language changes?
Study Design & Method
The carried-over protocol. Four parts, reused without modification. A written annotation guide defines four thematic categories — methodology (teaching methods and course delivery), evaluation (assessment practices), interaction (student–educator treatment), and attendance/engagement — plus three sentiment classes (positive, negative, neutral). Reliability is measured as test-retest intra-annotator agreement: a single annotator re-labeling a 100-comment sample one week apart, yielding Cohen's κ = 0.82 for thematic and 0.88 for sentiment. Crucially, the paper stresses these values estimate one annotator's stability over time and are not inter-annotator agreement. Model selection uses stratified five-fold cross-validation, and a 20% held-out test set is reserved for final reporting only. The same guide and procedure apply to every arm and both languages, so the comparison tests the protocol, not a particular model.
Spanish data. 44,707 valid open-ended comments from the SIIAU teaching-evaluation system at Universidad de Guadalajara, collected across three cycles (2022B, 2023A, 2023B) and anonymized. A single annotator labeled a subset of 1,165 comments; a 20% held-out set of 233 comments is reserved for final reporting. In that subset the sentiment classes are imbalanced — negative largest at 46.5%, positive 39.7%, neutral 13.7%.
Model arms: same protocol, three representation generations. (1) A sparse-lexical anchor: TF-IDF features with a linear support vector classifier (C = 0.1), the continuity baseline carried from the prior study. (2) A frozen-embedding 2019-generation arm using BETO (dccuchile/bert-base-spanish-wwm-uncased) mean-pooled frozen embeddings. (3) A modern multilingual-encoder arm using mean-pooled frozen embeddings from intfloat/multilingual-e5-base, feeding the same classifiers. (4) LLM arms: instruction-tuned models prompted zero- and few-shot with the annotation guide's label definitions, evaluated on the identical held-out items at three points — a cheap hosted model (Claude Haiku 4.5), the current frontier (Claude Opus 4.8), and a small open-weights model run locally (Qwen2.5-1.5B-Instruct, evaluated on Spanish on CPU and English on a consumer GPU). A parameter-efficient fine-tuning arm is a staged extension and is out of scope. For each arm the paper also reports inference cost, latency, and auditability.
English transfer data. Three openly-licensed resources. The primary corpus is a balanced 45,000-comment sentiment sample drawn from a large public RateMyProfessor teaching-evaluation release (9.54M reviews, CC BY-NC-SA per its Baidu AI Studio page). From the raw file the authors decode escaped byte sequences, drop comments under 15 characters and exact duplicates, remove reviews without a star rating, and take a class-balanced sample by uniform per-class reservoir sampling at a fixed seed — 15,000 comments per class, split 31,500 train / 6,750 development / 6,750 test, mean comment length 231 characters. Sentiment labels are star-derived under the same three-class scheme (positive ≥ 4 stars, negative ≤ 2, neutral otherwise). Because a star-derived label is a convenience signal rather than ground truth, the rule was validated against EduRABSA's review-level human sentiment labels: on the 3,000 RateMyProfessor teacher reviews sharing the star scale, the rule agrees with human labels at Cohen's κ = 0.66 with 82% accuracy, with the neutral class weakest because human-neutral reviews scatter across star levels. SIGHT (public YouTube lecture comments) is used only as a register-contrast resource, and synthetic LLM-generated "student feedback" datasets are excluded by design as primary corpora.
Evaluation discipline. Weighted F1 is the primary metric, with macro F1 to expose class imbalance and per-class precision, recall, and F1; because the classes are imbalanced, accuracy is not the headline metric. Paired comparisons are reported as descriptive evidence of consistency, never as significance tests: McNemar's test on held-out predictions plus 95% percentile bootstrap confidence intervals over 2,000 resamples, where an interval covering zero indicates no descriptive separation. The Spanish arms use the 233-comment split; the English LLM arms use a 1,500-comment subset, so those values are not directly comparable to the full 6,750-comment classical and encoder arms.
Key Findings
- The protocol is durable across representation generations. On the Spanish 233-comment held-out set (single stratified split, seed 20260616): TF-IDF + LinearSVC reached thematic wF1 0.602 / macro 0.459 and sentiment wF1 0.814; frozen BETO-2019 reached 0.744 / 0.681 / 0.838; the modern e5 encoder reached 0.740 / 0.678 / 0.859; Haiku 4.5 zero-shot 0.854 / 0.832 / 0.923 and few-shot 0.841 / 0.817 / 0.922; Opus 4.8 zero-shot 0.874 / 0.854 / 0.884 and few-shot 0.886 / 0.864 / 0.910; Qwen2.5-1.5B zero-shot 0.307 / 0.354 / 0.761 and few-shot 0.626 / 0.567 / 0.819.
- The biggest gain is in thematic macro F1, which rises from 0.46 for TF-IDF to about 0.68 for the BETO encoder and then to roughly 0.82–0.86 for the LLM arms — newer representations mostly rescue the rare thematic classes.
- The frontier model shows no sentiment advantage in Spanish. On the hardest thematic task Opus 4.8 posts the best figure, but on sentiment the cheap Haiku 4.5 zero-shot reaches 0.923 against Opus 4.8 zero-shot at 0.884 (head-to-head McNemar b = 13, c = 4 in Haiku's favor), and across the benchmark Opus cost about 7.7× as much as Haiku.
- A task/label ceiling may explain the missing frontier gain. Weighted F1 already sits near 0.92 on this 233-item Spanish set, so a ceiling — more than capability — could account for the frontier model's failure to pull ahead.
- Per-class weakness is concentrated in the residual and rare classes. For the frozen BETO-2019 arm, the neutral sentiment class is the weakest (F1 0.39; precision 0.46, recall 0.34, n = 32) while positive reaches 0.95 and negative 0.87; thematically, attendance/engagement is lowest (F1 0.61, n = 27) against methodology 0.82, interaction 0.68, and evaluation 0.62.
- Sentiment transfers to English, but the ranking changes. On the full 6,750-comment test split the modern e5 encoder is the strongest arm at wF1 0.729 (macro 0.729); TF-IDF trails at 0.675. With LLM arms on the 1,500-comment subset: Haiku 4.5 zero-shot 0.671 / few-shot 0.733; Opus 4.8 zero-shot 0.672 / few-shot 0.723; Qwen2.5-1.5B zero-shot 0.558 / few-shot 0.620. Few-shot examples are what lift the LLMs to encoder parity.
- The frontier model brings no English advantage either. Opus and Haiku sit within a point of each other at every setting, and neither closed LLM leads English by the descriptive margin both show on Spanish — no significance is claimed.
- The encoder lead over the LLMs is indicative, not separable. The e5 encoder's 0.729 carries a bootstrap interval of [0.718, 0.740] on the full split, while Haiku few-shot's 0.733 carries [0.710, 0.756] on the smaller subset; the intervals overlap. Other headline comparisons: Spanish sentiment Haiku-0 − Opus-0 +0.040 [0.007, 0.075]; Spanish thematic BETO − Opus-few −0.142 [−0.218, −0.079]; English sentiment e5 − TF-IDF +0.054 [0.043, 0.066]; English sentiment Haiku-f − Opus-f +0.010 [−0.006, 0.026].
- English is ceiling-bound by label noise. Because a classifier cannot exceed its labels' agreement, the star-versus-human κ of 0.66 caps the English sentiment task, unlike the gold-labeled Spanish setting.
- A LIME auditability check aligns local explanations with the annotation guide. On held-out Spanish predictions from the classical TF-IDF and e5 arms, grosero and amable drive the interaction class, evaluación and calificación drive evaluation, and faltaba and tarde drive attendance/engagement; on sentiment, excelente and muy drive positive while no, falta, and grosero drive negative, and the neutral class leans on hedging cues such as pero and sin embargo — consistent with its lower per-class F1. This is the inspectable auditability the classical arms provide and the prompt-bound LLM arms do not.
- Stated limitations. All Spanish gold labels come from one annotator, so the κ values are test-retest intra-annotator stability, not inter-annotator agreement; a second annotator is the single most important next step. The Spanish durability ordering is read from one stratified split (seed 20260616, 233 held-out items) with descriptive bootstrap intervals but no split variation, so robustness across resampled or repeated splits is future work. Spanish and English results are not directly comparable — they differ in label source (gold versus star-derived), domain, and training-set size, with label quality the dominant confound; a down-sampled English training set matched to the Spanish size is left to future work. English sentiment labels are heuristic and their noise is quantified but not eliminated. School, department, and state fields in the public review collection are unreliable parsing artifacts, so no per-institution or per-region English claims are made. Register and level gaps apply: the English data are higher-education reviews and SIGHT is public lecture-video comments, and neither is K-12 family feedback. Cross-language thematic transfer is a documented category alignment, not a validated classifier, and its mapping is imperfect. The English corpus is scraped, self-selected, commercial-platform data — appropriate for a research benchmark with provenance stated rather than for a public-institution corpus — and the Spanish institutional corpus is not redistributed. No demographic or subgroup claim is made and no equity audit is run: equity is out of scope by design.
What this means for practice
- Faculty developers. Adopt the protocol before the model: fix the annotation guide, the stratified splits, and the held-out set first, since the same procedure carried unchanged from TF-IDF (Spanish thematic wF1 0.602) through frozen BETO (0.744) to Opus 4.8 (0.886) across every representation arm tested.
- Faculty developers. Start with a near-free frozen encoder before paying for a frontier model: BETO reached 0.744 thematic and 0.838 sentiment weighted F1 on the Spanish held-out set, and on sentiment the cheaper Haiku 4.5 zero-shot (0.923) beat Opus 4.8 zero-shot (0.884) while costing about 7.7× less.
- Faculty developers. Measure annotator reliability before trusting any classifier and state what the number means: the reported κ = 0.82 (thematic) and 0.88 (sentiment) come from one annotator re-labeling 100 comments a week apart and are test-retest stability, not inter-annotator agreement.
- Instructors. Route classifier output to aggregate review and triage only, never to personnel decisions: the models produce a noisy signal and even the validation labels are imperfect proxies, with star-derived English sentiment labels agreeing with human labels at only κ = 0.66 and 82% accuracy.
- Software developers. Keep the classifier inspectable where deployment is institutional, because the LIME-based audit tied predictions to the annotation guide's own terms (evaluación, faltaba, grosero, pero, sin embargo) on the classical arms while the prompt-bound LLM arms offered no comparable check.
Limitations
- All Spanish gold labels come from a single annotator, so the κ values estimate one annotator's stability over time rather than inter-annotator agreement; the paper names a second annotator as the single most important next step.
- The Spanish durability ordering is read from one stratified split (seed 20260616, 233 held-out comments) with descriptive bootstrap intervals only and no split variation, and the English LLM arms use a 1,500-comment subset against 6,750 for the classical and encoder arms, so those values are not directly comparable.
- Spanish and English results differ in label source, domain, and training-set size, with label quality the dominant confound; English sentiment labels are star-derived and their noise is quantified (κ = 0.66) but not eliminated, which caps the English task.
- The English corpus is scraped, self-selected RateMyProfessor data from a commercial platform whose school, department, and state fields are unreliable parsing artifacts, and it covers higher-education reviews only — no K-12 family feedback; no equity audit was run and no subgroup claim is made.
Citation
Esteban U. Vega Barajas (2026). A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol. arXiv preprint.