On this page

Synthesis: CODE-GEN is a dual-agent RAG (Retrieval-Augmented Generation)-based Agentic AI system for generating and validating coding-comprehension multiple-choice questions, evaluated by 6 SMEs across 7 pedagogical dimensions (N=288 questions, 2,016 rating pairs). AI excels at criteria-matching and computational verification (concept alignment 98.6%, code validity 95.5%), but human expertise remains essential for distractor quality (79.9%) and pedagogically rich feedback — providing an evidence-based division of labor for Human-in-the-Loop educational content generation. Its distinctive move is treating the automated Validator's judgment as an empirical object of study rather than an assumed capability.

Venue: AIED 2026 (short paper) ArXiv: 2604.03926

Overview

CODE-GEN (Context-aligned, Output-validated, Dual-agent, Expert-guided GENeration) is a Human-in-the-Loop Agentic AI system for generating contextually grounded multiple-choice coding comprehension questions. It integrates RAG (Retrieval-Augmented Generation) with a dual-agent architecture separating question generation from quality validation.

Key Findings

  1. AI reliably automates technically grounded dimensions. Human-validated success rates are high where correctness maps to computational verification and explicit criteria: concept alignment 98.6%, question stem clarity 97.9%, code validity 95.5%, correct answer validity 92.0%, and correct-answer feedback 92.4%.
  2. Human expertise remains essential for pedagogical judgment. Distractor quality (79.9% success, 15.6% failure) and distractor feedback quality (86.1%) are the weakest dimensions, because assessing whether distractors target common Misconceptions about AI and whether feedback elaborates underlying concepts requires instructional judgment, not just computational correctness.
  3. Tool augmentation works. An Arithmetic Expression Evaluator and a Sandboxed Python Runner substantially enhance both generation quality and validation reliability for computationally grounded dimensions.
  4. Automated evaluators must be validated, not assumed. Large Language Models (LLMs)-based critique agents exhibit hallucination, bias, and inconsistent judgment; CODE-GEN explicitly compares the Validator against human SMEs, showing that without such validation, errors at the evaluation stage risk being amplified rather than corrected.
  5. Systematic failure patterns. False positives (approving syntactically valid but instructionally shallow distractors and surface-mechanics feedback) and false negatives (misinterpreting answer schemas, confusing answer value with option position) reveal where the Validator's judgment breaks down.
  6. SME–Validator agreement is consistent. Agreement rates across all dimensions range 82.5%–98.4%, with five of seven dimensions showing standard deviations ≤3.8%, indicating SMEs applied comparable standards despite reviewing different question sets.

Architecture

  1. RAG Pipeline: Instructional materials (learning objectives, example questions, code) are parsed with a domain-specific chunking strategy that preserves semantic coherence (line-by-line parsing with docstring recognition, explicitly preserving objectives, sample questions with answer options, and Python code), embedded via OpenAI text-embedding-3-small, and indexed in a FAISS vector store with L2 distance. On generation, nearest-neighbor retrieval (IndexFlatL2) injects relevant examples into the Generator's prompt.
  2. Generator Agent (GPT-4.1): Produces MCQs with stem, executable code, four answer options, and explanatory feedback. Augmented with an Arithmetic Expression Evaluator tool for deterministic computation, mitigating known LLM arithmetic weaknesses.
  3. Validator Agent (GPT-5-mini): Independently assesses each question across seven pedagogical dimensions: question stem clarity, code validity, concept alignment, correct answer validity, distractor quality, correct answer feedback quality, and distractor feedback quality. Uses an Arithmetic Expression Evaluator and a Sandboxed Python Runner for code execution verification. Model selection for both agents was driven by comparative experiments (Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5-mini, GPT-4.1) on novelty, correctness, and latency.

Evaluation

  • 6 SMEs (three men, three women) evaluated 288 AI-generated questions
  • 2,016 human-AI rating pairs (SME agreement/disagreement with Validator)
  • 131 qualitative feedback instances

Key Results

Dimension Success Rate Failure Rate
Concept Alignment 98.6% 0.3%
Question Stem Clarity 97.9% 2.1%
Code Validity 95.5% 3.1%
Correct Answer Feedback 92.4% 2.1%
Correct Answer Validity 92.0% 1.4%
Distractor Feedback Quality 86.1% 9.4%
Distractor Quality 79.9% 15.6%

Division of Labor Findings

AI strengths (reliable automation):

  • Computational verification and explicit criteria matching
  • Concept alignment via RAG grounding
  • Code syntax and output verification via tool augmentation

Human-essential dimensions (require oversight):

  • Designing pedagogically meaningful distractors that target common student misconceptions
  • Providing feedback that elaborates on underlying concepts, not just surface mechanics
  • Interpreting structured answer representations (Validator sometimes confused option position with answer value)

Failure Patterns

  • False positives: Validator approved distractors that were syntactically valid but instructionally shallow; approved feedback that described surface mechanics without deeper elaboration
  • False negatives: Validator misinterpreted answer schemas (confusing answer value with option position); internal inconsistency where textual analysis affirmed correctness but binary classification contradicted it

Significance for automated-assessment and automated-question-generation

CODE-GEN demonstrates that agentic AI with RAG grounding and tool augmentation can serve as scalable first-line quality control for Automated Assessment item generation. The explicit evaluation of the Validator against human judgment — rather than assuming automated evaluation is reliable — provides an evidence-based framework for determining where AI can be safely delegated and where Human-in-the-Loop oversight must be maintained.

What this means for practice

  • Designers. Delegate first the dimensions where correctness is computationally verifiable: human-validated success reached 98.6% for concept alignment, 97.9% for stem clarity, and 95.5% for code validity.
  • Designers. Keep expert reviewers on the pedagogical dimensions — distractor quality was the weakest at 79.9% success with 15.6% failure and distractor feedback at 86.1% — because these require anticipating common Misconceptions about AI, not checking correctness.
  • Designers. Validate an automated critique agent against experts before trusting its verdicts: SME agreement ranged 82.5%–98.4%, and the Validator both approved instructionally shallow items and confused answer value with option position.
  • Designers. Augment both agents with deterministic tools (an arithmetic expression evaluator and a sandboxed Python runner) instead of relying on the model for computation and code execution.
  • Researchers. Treat the evaluator's judgment as an empirical object and keep the disagreement cases; the 131 qualitative feedback instances are a usable seed set for pedagogical alignment work.

Limitations

  • Six SMEs (three men, three women) rated 288 generated questions, yielding 2,016 human–AI rating pairs; the rater pool is small and all of them taught introductory programming.
  • The study is confined to introductory Python: the authors note the architecture is not domain-specific, but no other subject area was tested.
  • Items were rated by experts rather than deployed in a course, so effects on student learning or assessment validity in authentic settings are unmeasured.
  • The Generator (GPT-4.1) and Validator (GPT-5-mini) are specific commercial models, so dimension-level success rates may not transfer to other backbones.

Citation

Duan, X., Nwanganga, F., & Wang, C. (2026). CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.