Research Article
MuTSE: A Human-in-the-Loop Multi-use Text Simplification Evaluator
Synthesis: MuTSE (Roscan et al., 2026) tackles a critical methodological gap in text simplification for Intelligent Tutoring Systems (ITS) and language-learning applications: how to systematically evaluate Large Language Models (LLMs)-generated simplifications across many prompting strategies and model architectures without drowning researchers in high-dimensional comparisons. It pairs an asynchronous P × M generation pipeline with a novel tiered semantic alignment engine — biased by a real-time linearity heuristic (λ) — so evaluators can visually trace how each source sentence is transformed across every prompt–model permutation, then score the results on custom, pedagogically meaningful criteria.
Key Findings
- Parallel comparison workflow: MuTSE executes
Pprompts ×Mmodels concurrently (bounded by the slowest model's latency, ~O(max(tᵢ))) and presents all outputs in a unified, side-by-side, toggleable matrix — removing the need to run or orchestrate inference manually. - Tiered semantic alignment engine: A 3-tier cascade (multilingual SBERT embeddings → TF-IDF with character n-grams → normalized positional fallback) plus a user-adjustable linearity bias (λ, 0–2, default 0.5) maps source sentences to simplified counterparts, suppressing false-positive alignments and enabling CPU-only deployment.
- Evaluation-agnostic human annotation: No fixed rubric — evaluators build custom rating scales (binary through continuous 100-point) with relative weights, normalized into a weighted performance percentage, sidestepping the field's contested Likert-scale conventions.
- NLP-ready data export: Every session persists locally and exports as standardized JSON/CSV, directly supporting fine-tuning, reward-model training, and exploratory data analysis.
The Challenge of Text Simplification in Education
Adapting complex texts into accessible reading material is a core component of Intelligent Tutoring and Language Learning applications. Large Language Models now give educators robust generative frameworks to tailor texts to specific proficiency targets such as CEFR A2 or B1. Yet evaluating those outputs remains fragmented and labor-intensive: researchers fall back on static scripts or notebooks, while educators are confined to standard conversational chat interfaces — neither supports systematic, multi-dimensional evaluation of prompt–model permutations.
Standard automated metrics such as BLEU and SARI (via developer-centric toolkits like EASSE) drive large-scale benchmarking, but they aggregate performance into single opaque numbers and lack the granular, visual interpretability educators need to validate multi-reference transformations. Related human-in-the-loop annotation tools (e.g. TS-ANNO) are built for post-hoc corpus creation rather than live, concurrent multi-model generation. MuTSE deliberately bridges both worlds — combining the concurrent generation of modern LLMs, the metric extraction of programmatic toolkits, and the visual mapping of annotation platforms into one accessible environment.
MuTSE: Architecture and Design
MuTSE rests on a decoupled, asynchronous client-server architecture: a Python/FastAPI backend orchestrates parallel LLM generation via Together AI's serverless endpoints, while a Vue.js 3 frontend handles multi-dimensional comparison and client-side alignment. A lightweight, local JSON persistence layer keeps deployment portable for individual educators.
Semantic Alignment Engine
The core innovation is a hierarchical fallback strategy that pairs semantically congruent sentences without the false positives that plague plain cosine similarity:
- Primary semantic level: a condensed 384-dimensional multilingual transformer (paraphrase-multilingual-MiniLM-L12-v2) computes the base cosine-similarity matrix.
- Secondary lexical level: if embeddings fail or return null vectors, a hybrid TF-IDF representation using word- and character-level n-grams captures morphological similarity.
- Tertiary positional level: a structural fallback aligning sentences purely by normalized sequence position.
The linearity bias (λ) acts as a structural regularizer, offsetting the reduced discriminative capacity of compressed embedding models. Because the final alignment graph is recomputed on the client, adjusting λ yields instantaneous visual feedback without redundant server calls.
Customizable Annotation and Metrics
MuTSE is deliberately evaluation-agnostic — it ships with zero predefined metrics. Through the Settings module, evaluators define their own dimensions, arbitrary rating scales, and relative impact weights. Alongside manual scoring, the interface badges real-time textual diagnostics: word frequency, sentence count, average length, compression ratio, Flesch-Kincaid grade level, and Flesch Reading Ease — all computed dynamically with language-specific hyphenation. This unification of automated readability metrics, visual semantic mapping, and inline scoring accelerates consistent multi-dimensional comparison.
Connection to LLMs in Education
As Generative AI becomes prevalent in ITS, text simplification confronts:
- Prompting strategy variability: the same model, prompted differently, produces materially different simplifications.
- Architecture differences: GPT, Claude, and specialized open-weight models (Llama 3, DeepSeek V3, Qwen 2.5) diverge in output quality and style.
- Evaluation challenge: linguistic metrics fail to capture pedagogical quality, and the field lacks standardized evaluation terminology.
MuTSE addresses these by letting evaluators toggle prompts and models on the fly, visually trace alignments, and detect conversational artifacts or semantic hallucinations via inline cosine-similarity scores — moving beyond aggregate benchmarks toward reproducible, human-in-the-loop qualitative assessment.
What this means for practice
-
Researchers. Compare prompt–model permutations instead of single averages. MuTSE runs P prompts × M models concurrently and presents every output in a side-by-side, toggleable matrix, so differences between prompting strategies and model architectures stay visible rather than collapsing into one score.
-
Researchers. Define your rating dimensions and weights before you start, then cross-check the scores against the system's real-time diagnostics — compression ratio, Flesch-Kincaid grade level, and Flesch Reading Ease. MuTSE ships with zero predefined metrics and normalizes custom scales (binary through continuous 100-point) into a weighted performance percentage, which sidesteps contested Likert conventions and lets the readability metrics flag conversational artifacts and semantic hallucinations.
-
Software developers. Adopt the three-tier alignment cascade (multilingual SBERT embeddings, TF-IDF with word- and character-level n-grams, then normalized positional fallback) to suppress false-positive sentence alignments, and keep the deployment CPU-only by using the 384-dimensional MiniLM embedding tier.
-
Software developers. Expose the linearity bias λ (range 0-2, default 0.5) as a user-adjustable control recomputed on the client, so evaluators get instant visual feedback without redundant server calls, and export each session as structured JSON and CSV — the JSON preserving sentence-to-annotation mappings for fine-tuning or reward-model training, the CSV for exploratory analysis.
-
Instructors. Drive simplification from real-time proficiency estimates rather than a fixed corpus level: dynamic simplification belongs inside an Adaptive Learning loop, with Learner Modeling and Adaptive Instruction supplying the reading-level target, an educator verifying the output before learners see it, and quality weights (meaning preservation, fluency, concept coverage) set locally rather than accepted by default. Route those texts through human review (Human-in-the-Loop) so the cascade supports Scaffolding and Sociocultural Learning instead of stripping key concepts.
Limitations
- MuTSE's local JSON persistence does not scale to concurrent multi-user deployments; a relational database would be needed for large collaborative annotation campaigns.
- Cloud-based model access removes local GPU requirements but still imposes initial environment-configuration friction.
- The alignment cascade is optimized for monolingual simplification; cross-lingual syntactic restructuring may not respect monotonic sentence order, so extending it to machine translation requires recalibrating λ.
Citation
Roscan, R.-A., Petre, G., Dumitran, A.-M., & Dumitran, A.-L. (2026). MuTSE: A Human-in-the-Loop Multi-use Text Simplification Evaluator.