On this page

Synthesis: PromptDecipher addresses a critical gap in AI tutor deployment: teacher quality assurance. A formative study of 121 chatbots created by instructors in an "AI for Educators" MOOC revealed that educators authoring AI tutoring chatbots virtually never systematically test them before student deployment — a finding with serious implications for SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems and educational quality, particularly for K–12 learners. The system shifts the authoring paradigm from abstract system-prompt writing to direct correction-based interaction: teachers edit undesirable bot responses in a live simulated chat, and an automated pipeline analyzes the correction, infers the pedagogical intent, proposes a targeted system prompt rewrite, and validates it across previously passed test scenarios. This bridges the Teaching gap between classroom practitioner and AI system designer — a tension also explored in Modeling AI-TPACK in Practice: Insights from Teachers'' Multi-Agent Workflow Design, which found that effective AI integration requires systems thinking beyond simple tool use. By embedding testing directly into the authoring workflow, PromptDecipher scaffolds teachers in roles they would otherwise skip, resonating with the Agentic AI paradigm of AI-scaffolded work. The design also mitigates the kind of diagnostic failures identified in Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most, where LLMs struggle precisely where feedback matters most.

Key Findings

  1. In a formative study of 121 teacher-authored chatbots, nearly all instructors successfully specified learning content but virtually none systematically tested their bots before publishing them to students.
  2. PromptDecipher repositions the core authoring activity: instead of writing an abstract system prompt, teachers directly edit undesirable bot responses in a live simulated student chat.
  3. The Reverse Prompting Pipeline infers pedagogical intent from each correction, proposes a minimal system prompt rewrite, and runs regression verification across all previously passed test scenarios.
  4. Publication is gated behind at least one completed test-correct-verify cycle, embedding quality assurance structurally into the workflow rather than relying on voluntary compliance.

The Formative Study: Unsystematic Teacher QA

The paper's motivating problem comes from a formative study of 121 chatbots created by instructors in an "AI for Educators" MOOC. While teachers successfully specified learning content, nearly all failed to engage in any systematic testing before publication — a critical concern when the resulting bots are deployed to real learners, including K–12 students. The authors trace this to a mismatch between the interface these chatbot platforms provide — a raw text editor for a system prompt — and teachers' existing mental models. Teachers are deeply experienced in giving corrective Feedback on student work, but have no prior frame for specifying AI behavior in natural language. Prior work on end-user prompt engineering documents that non-experts approach prompting opportunistically rather than systematically and struggle to translate observations about undesired outputs into concrete prompt requirements. The core problem is that effective Intelligent Tutoring authorship demands teachers simultaneously act as learning designers, AI interaction designers, and QA engineers — roles far beyond their typical experience.

PromptDecipher System Design

PromptDecipher is a web-based system that restructures the authoring workflow around a direct correction-based interaction rather than abstract prompt writing. A teacher creates a new bot, selects a foundation model (OpenAI, Anthropic, or Google), and may optionally upload course materials. Rather than drafting a system prompt from scratch, the teacher is directed to a test environment: a simulated student chat. They select a student profile (e.g., "expected path," "struggling learner," "off-topic input"), read the bot's response, and either mark it as passing or edit it to reflect the desired behavior. Submitting an edit triggers the Reverse Prompting Pipeline. The system additionally supports direct prompt editing via templates and an AI-assisted discussion panel, though the correction-based pipeline is the primary contribution.

The Reverse Prompting Pipeline

Each teacher correction triggers a three-stage automated pipeline:

  1. Diff analysis. An Large Language Models (LLMs) compares the original and corrected responses to infer the teacher's pedagogical intent (e.g., "the bot should ask a follow-up question rather than give the answer directly").
  2. Prompt rewrite. A targeted addition or modification to the system prompt is proposed and shown to the teacher for review as a tracked diff.
  3. Regression verification. The revised prompt is automatically evaluated across all previously passed test cases; any regression is flagged for the teacher's attention before they can proceed.

Because teachers must complete at least one such cycle before publishing, QA is structurally embedded in the workflow. This test-correct-verify cycle enforces Pedagogical Safety as a first-class activity and scaffolds teachers in roles they would otherwise skip, exemplifying Scaffolding through tool design rather than instruction.

Demonstration

The paper presents an interactive demonstration plan in which attendees author an AI tutoring bot end-to-end using provided laptops. In a short setup, the attendee creates a bot, enters a brief description of the learning context (e.g., "a Socratic tutor for introductory statistics"), and selects a foundation model. They then select a simulated student profile, read the bot's initial response, and edit it to reflect what they wish the bot had said, watching the Reverse Prompting Pipeline run live — showing the inferred intent, the proposed prompt update, and regression check results. After iterating through one additional scenario, they publish the bot and receive a shareable link to interact with it as a student.

What this means for practice

  • Instructors. Use authoring tools that let you correct a simulated bot response rather than hand-write a system prompt: making the response the unit of work converts prompt engineering into the corrective Feedback teachers already give student work.
  • Faculty developers. Require a completed test-correct-verify cycle before any tutoring bot goes live, because the formative study found that nearly all of 121 teacher-authored chatbots were published with no systematic testing.
  • Faculty developers. Spend professional development on learning-design judgment and quality assurance rather than prompt-writing technique, since effective Generative AI adoption in education may depend more on authoring environments aligned with teachers' existing skills than on prompt training.
  • Instructors. Test a bot against explicit student profiles — expected path, struggling learner, off-topic input — and correct the responses you dislike before students ever meet the bot, which is the QA role the raw-prompt interface leaves out.
  • Faculty developers. Plan to evaluate the pipeline rather than assume it works: PromptDecipher's deployment in an "AI for Educators" MOOC with hundreds of higher-education instructors is scheduled for fall 2026 and is what will supply the usage data on testing rates and prompt quality.

Limitations

  • This is a three-page demonstration paper: PromptDecipher has not been evaluated with users, and the authors' planned MOOC deployment in fall 2026 is the study that will test whether correction-based authoring increases testing rates.
  • The motivating evidence is a formative study of 121 chatbots from a single "AI for Educators" MOOC that documents the absence of testing but compares no alternative authoring interface and runs no comparison condition.
  • The pipeline delegates intent inference and prompt rewriting to an Large Language Models (LLMs) and validates changes against previously passed test scenarios, with no reported benchmark of how accurately the inferred intent matches what the teacher intended.
  • The conference demonstration collects only informal feedback from attendees, so no measured usability or quality-assurance outcome is reported.

Citation

Koyama, M., Xiao, R., & Stamper, J. (2026). PromptDecipher: Supporting AI Tutor Authoring Through Editable Simulated Interactions. In Proceedings of the 13th ACM Conference on Learning @ Scale (L@S '26).

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.