On this page

Synthesis: This paper introduces TurtleAI, a Benchmark containing 823 tasks curated from real-world visual programming in the Turtle Graphics domain, evaluating how well vision-language models (VLMs) perform on education-oriented visual programming. Most prior work focuses on visual programming for productivity; the authors find that current VLMs struggle significantly on these tasks, and that fine-tuning on synthetic data yields about a 20% improvement — informing programming education and multimodal AI evaluation.

Abstract

Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focuses on visual programming for productivity; it remains unclear how well current VLMs perform on education-oriented visual programming and what factors limit their performance. To bridge this gap, we introduce TurtleAI, a benchmark containing 823 tasks curated based on real-world visual programming tasks in the Turtle Graphics domain. Solving these tasks requires models to perceive geometric patterns, reason about spatial relationships, and synthesize Python code that faithfully reproduces geometric patterns. We evaluate 20+ VLMs, including GPT-5, GPT-4o, and Qwen2-VL-72B, and find that they struggle significantly, with most achieving success rates below 30%. To address these limitations, we propose a data generation technique that requires only a small set of seed samples. Fine-tuning Qwen2-VL-72B on the resulting synthetic data yields an improvement of about 20% on real-world tasks. Failure analysis reveals that GPT-4o struggles with spatial reasoning and precise visual replication, whereas fine-tuning primarily improves the alignment between visual perception and code generation — a contribution to programming education, multimodal models, and benchmarking AI in education.

What this means for practice

  • Designers. Treat VLM-generated Turtle Graphics code as a prototype rather than a classroom feature: on the real-world dataset of tasks designed for students in grades 3-6, even the top-performing model, o3, reaches only a 40.2% symbolic success rate, and the strongest open-source base model, Qwen2.5-VL, reaches only 6.80% on the full benchmark.
  • Instructors. Sequence Turtle Graphics content from Basic Geometry, where models perform best, toward Spiral tasks, the most challenging category for all models because they demand long-horizon sequential control; this tells you where student support should concentrate.
  • Designers. When a tutoring feature needs broader Multimodal AI capability, fine-tune on generated data: fine-tuning Qwen2-VL-72B on TurtleAI-Datagen output improved real-world task performance by over 20%, and the pipeline expands a seed set of only 10 image-code pairs into a training set of 738,126 samples.
  • Instructors. Expect model assistance to be least reliable on hand-drawn student work: fine-tuning improved performance mainly on DSReal and DSSyn while DSCraft performance remained low.
  • Researchers. Use the 823-task benchmark, composed of 102 real-world, 102 hand-drawn, and 619 synthetic tasks, to compare models on education-oriented visual programming rather than productivity-oriented tasks.

Limitations

  • The evaluation framework normalizes drawing comparison to be invariant to size, translation, and line width, which the authors acknowledge may discard geometric variations essential to inverse-graphics tasks; code quality is likewise reduced to compactness, measured only by length ratio.
  • No systematic ablation or comparison against other data synthesis techniques was run, so the contribution of individual TurtleAI-Datagen stages is not isolated.
  • Domain coverage is narrow: the 823-task benchmark is confined to Turtle Graphics, and only 102 tasks come from real-world student drawings while 619 are synthetic.
  • Fine-tuned models still struggle on hand-drawn inputs (DSCraft), which the authors attribute to synthetic training data that focuses on clean renderings and lacks the noise and distortions of human drawings.

Citation

Wen, C., & Staub, J. (2026). TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.