AcademiClaw: When Students Set Challenges for AI Agents

Created: 2026-05-11 | Tags: benchmarkhigher-edllmgenerative-aistudent-experience

Yu, Lu, Si et al. (77 authors, 2026) โ€” Shanghai Jiao Tong University, SII, GAIR. Open-source benchmark.

Core Contribution

AcademiClaw is a bilingual benchmark of 80 complex, long-horizon tasks sourced from university students' real academic workflows โ€” homework, research projects, competitions, and personal projects โ€” that they found current AI agents unable to solve. It extends the OpenClaw ecosystem beyond assistant-level tasks into academic-level agent evaluation.

Benchmark Design

Task Sourcing & Curation

Execution & Scoring

Key Results

Why This Matters for AI in Education

AcademiClaw flips the evaluation paradigm: instead of researchers designing artificial benchmarks, students define what AI agents should be able to do. This aligns evaluation with real educational needs:

1. Authentic task validity: Tasks reflect genuine academic workflows, not synthetic proxies 2. Capability gap diagnosis: The 55% ceiling reveals where AI agents still fail students 3. Token-output disconnect: Challenges the assumption that more compute solves academic problems โ€” relevant to AI tutoring cost/benefit analysis 4. Safety in academic contexts: The 5-category safety audit surfaces risks specific to educational AI deployment

Connection to the Wiki

Limitations

Open Questions

Related Pages

Citation

APA: Yu, J., Lu, P., Si, W., ... & Liu, P. (2026). AcademiClaw: When Students Set Challenges for AI Agents. arXiv:2605.02661.