Research Article
CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation
Synthesis: CODE-GEN is a dual-agent RAG (Retrieval-Augmented Generation)-based Agentic AI system for generating and validating coding-comprehension multiple-choice questions, evaluated by 6 SMEs across 7 pedagogical dimensions (N=288 questions, 2,016 rating pairs). AI excels at criteria-matching and computational verification (concept alignment 98.6%, code validity 95.5%), but human expertise remains essential for distractor quality (79.9%) and pedagogically rich feedback — providing an evidence-based division of labor for Human-in-the-Loop educational content generation. Its distinctive move is treating the automated Validator's judgment as an empirical object of study rather than an assumed capability.
Venue: AIED 2026 (short paper) ArXiv: 2604.03926
Overview
CODE-GEN (Context-aligned, Output-validated, Dual-agent, Expert-guided GENeration) is a Human-in-the-Loop Agentic AI system for generating contextually grounded multiple-choice coding comprehension questions. It integrates RAG (Retrieval-Augmented Generation) with a dual-agent architecture separating question generation from quality validation.
Key Findings
- AI reliably automates technically grounded dimensions. Human-validated success rates are high where correctness maps to computational verification and explicit criteria: concept alignment 98.6%, question stem clarity 97.9%, code validity 95.5%, correct answer validity 92.0%, and correct-answer feedback 92.4%.
- Human expertise remains essential for pedagogical judgment. Distractor quality (79.9% success, 15.6% failure) and distractor feedback quality (86.1%) are the weakest dimensions, because assessing whether distractors target common Misconceptions about AI and whether feedback elaborates underlying concepts requires instructional judgment, not just computational correctness.
- Tool augmentation works. An Arithmetic Expression Evaluator and a Sandboxed Python Runner substantially enhance both generation quality and validation reliability for computationally grounded dimensions.
- Automated evaluators must be validated, not assumed. Large Language Models (LLMs)-based critique agents exhibit hallucination, bias, and inconsistent judgment; CODE-GEN explicitly compares the Validator against human SMEs, showing that without such validation, errors at the evaluation stage risk being amplified rather than corrected.
- Systematic failure patterns. False positives (approving syntactically valid but instructionally shallow distractors and surface-mechanics feedback) and false negatives (misinterpreting answer schemas, confusing answer value with option position) reveal where the Validator's judgment breaks down.
- SME–Validator agreement is consistent. Agreement rates across all dimensions range 82.5%–98.4%, with five of seven dimensions showing standard deviations ≤3.8%, indicating SMEs applied comparable standards despite reviewing different question sets.
Architecture
- RAG Pipeline: Instructional materials (learning objectives, example questions, code) are parsed with a domain-specific chunking strategy that preserves semantic coherence (line-by-line parsing with docstring recognition, explicitly preserving objectives, sample questions with answer options, and Python code), embedded via OpenAI text-embedding-3-small, and indexed in a FAISS vector store with L2 distance. On generation, nearest-neighbor retrieval (IndexFlatL2) injects relevant examples into the Generator's prompt.
- Generator Agent (GPT-4.1): Produces MCQs with stem, executable code, four answer options, and explanatory feedback. Augmented with an Arithmetic Expression Evaluator tool for deterministic computation, mitigating known LLM arithmetic weaknesses.
- Validator Agent (GPT-5-mini): Independently assesses each question across seven pedagogical dimensions: question stem clarity, code validity, concept alignment, correct answer validity, distractor quality, correct answer feedback quality, and distractor feedback quality. Uses an Arithmetic Expression Evaluator and a Sandboxed Python Runner for code execution verification. Model selection for both agents was driven by comparative experiments (Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5-mini, GPT-4.1) on novelty, correctness, and latency.
Evaluation
- 6 SMEs (three men, three women) evaluated 288 AI-generated questions
- 2,016 human-AI rating pairs (SME agreement/disagreement with Validator)
- 131 qualitative feedback instances
Key Results
| Dimension | Success Rate | Failure Rate |
|---|---|---|
| Concept Alignment | 98.6% | 0.3% |
| Question Stem Clarity | 97.9% | 2.1% |
| Code Validity | 95.5% | 3.1% |
| Correct Answer Feedback | 92.4% | 2.1% |
| Correct Answer Validity | 92.0% | 1.4% |
| Distractor Feedback Quality | 86.1% | 9.4% |
| Distractor Quality | 79.9% | 15.6% |
Division of Labor Findings
AI strengths (reliable automation):
- Computational verification and explicit criteria matching
- Concept alignment via RAG grounding
- Code syntax and output verification via tool augmentation
Human-essential dimensions (require oversight):
- Designing pedagogically meaningful distractors that target common student misconceptions
- Providing feedback that elaborates on underlying concepts, not just surface mechanics
- Interpreting structured answer representations (Validator sometimes confused option position with answer value)
Failure Patterns
- False positives: Validator approved distractors that were syntactically valid but instructionally shallow; approved feedback that described surface mechanics without deeper elaboration
- False negatives: Validator misinterpreted answer schemas (confusing answer value with option position); internal inconsistency where textual analysis affirmed correctness but binary classification contradicted it
Significance for automated-assessment and automated-question-generation
CODE-GEN demonstrates that agentic AI with RAG grounding and tool augmentation can serve as scalable first-line quality control for Automated Assessment item generation. The explicit evaluation of the Validator against human judgment — rather than assuming automated evaluation is reliable — provides an evidence-based framework for determining where AI can be safely delegated and where Human-in-the-Loop oversight must be maintained.
What this means for practice
- Designers. Delegate first the dimensions where correctness is computationally verifiable: human-validated success reached 98.6% for concept alignment, 97.9% for stem clarity, and 95.5% for code validity.
- Designers. Keep expert reviewers on the pedagogical dimensions — distractor quality was the weakest at 79.9% success with 15.6% failure and distractor feedback at 86.1% — because these require anticipating common Misconceptions about AI, not checking correctness.
- Designers. Validate an automated critique agent against experts before trusting its verdicts: SME agreement ranged 82.5%–98.4%, and the Validator both approved instructionally shallow items and confused answer value with option position.
- Designers. Augment both agents with deterministic tools (an arithmetic expression evaluator and a sandboxed Python runner) instead of relying on the model for computation and code execution.
- Researchers. Treat the evaluator's judgment as an empirical object and keep the disagreement cases; the 131 qualitative feedback instances are a usable seed set for pedagogical alignment work.
Limitations
- Six SMEs (three men, three women) rated 288 generated questions, yielding 2,016 human–AI rating pairs; the rater pool is small and all of them taught introductory programming.
- The study is confined to introductory Python: the authors note the architecture is not domain-specific, but no other subject area was tested.
- Items were rated by experts rather than deployed in a course, so effects on student learning or assessment validity in authentic settings are unmeasured.
- The Generator (GPT-4.1) and Validator (GPT-5-mini) are specific commercial models, so dimension-level success rates may not transfer to other backbones.
Citation
Duan, X., Nwanganga, F., & Wang, C. (2026). CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation.