Yu, Lu, Si et al. (77 authors, 2026) โ Shanghai Jiao Tong University, SII, GAIR. Open-source benchmark.
AcademiClaw: Student-Sourced AI Agent Benchmark
Yu, Lu, Si et al. (77 authors, 2026) โ Shanghai Jiao Tong University, SII, GAIR. Open-source benchmark.
Core Contribution
AcademiClaw is a bilingual benchmark of 80 complex, long-horizon tasks sourced from university students' real academic workflows โ homework, research projects, competitions, and personal projects โ that they found current AI agents unable to solve. It extends the OpenClaw ecosystem beyond assistant-level tasks into academic-level agent evaluation.
Benchmark Design
Task Sourcing & Curation
230 student-submitted candidates โ rigorous expert review โ 80 final tasksTasks are authentic: real problems students couldn't solve with current AI agents25+ professional domains: olympiad-level math, linguistics, GPU-intensive RL, full-stack system debugging16 tasks require CUDA GPU execution โ testing hardware-accelerated AI capabilitiesExecution & Scoring
Each task runs in an isolated Docker sandboxScored by multi-dimensional rubrics using 6 complementary techniquesIndependent 5-category safety audit provides behavioral analysis beyond task completionBilingual (Chinese/English)Key Results
6 frontier models testedBest pass rate: 55% โ no model exceeds thisSharp capability boundaries across task domainsDivergent behavioral strategies between modelsDisconnect between token consumption and output quality โ more tokens โ better resultsWhy This Matters for AI in Education
AcademiClaw flips the evaluation paradigm: instead of researchers designing artificial benchmarks, students define what AI agents should be able to do. This aligns evaluation with real educational needs:
1. Authentic task validity: Tasks reflect genuine academic workflows, not synthetic proxies
2. Capability gap diagnosis: The 55% ceiling reveals where AI agents still fail students
3. Token-output disconnect: Challenges the assumption that more compute solves academic problems โ relevant to AI tutoring cost/benefit analysis
4. Safety in academic contexts: The 5-category safety audit surfaces risks specific to educational AI deployment
Limitations
Contributor pool concentrated at a single institution (SJTU) โ limits cultural and disciplinary diversityGPU-intensive tasks (16/80) require specialized hardware, limiting reproducibilityBilingual but primarily Chinese university contextSafety audit releases aggregate statistics only, not full violation tracesOpen Questions
How would the 55% pass rate change with iterative refinement or multi-agent collaboration?Would results differ at non-Chinese universities with different academic workflows?Can the benchmark be adapted for K-12 or professional training contexts?What does the token-output disconnect imply for AI tutoring systems that bill by token usage?Connected Concepts
Automated Question GenerationPedagogical LLM TrainingSocratic MethodMath EducationPrompt EngineeringHuman In The Loop AIAffective TutoringAutomated Essay ScoringConnected Articles
Codify Socratic Tutoring Programming โ Codify: An Intelligent Socratic Tutoring System for Programming EducationAgentic AI Education Scoping Review โ Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent ParadigmAI Tutor Behavioral Evaluation โ The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor EffectivenessLets Chat Chatbot Outreach 2026 โ Let''s Chat: Leveraging Chatbot Outreach for Improved Course PerformanceAI Generated Feedback Higher Ed โ Artificial intelligence and feedback in university education: effectiveness and student perceptionsScheu Mobile Chatbot Journaling Motivation 2026 โ Designing a mobile chatbot-based learning journaling system for intrinsic motivation and engagementCitation
Yu, J., Lu, P., Si, W., Lu, H., Wu, J., Tao, K., et al. (2026). AcademiClaw: When Students Set Challenges for AI Agents. arXiv:2605.02661.