🧠 AI Ed Wiki

Synthesis: This paper reports a blind Turing Test evaluating leading LLMs on three Italian professional legal examinations: the Bar exam, Judges exam, and Notary exam. LLMs generated full written papers that were anonymised and graded by expert examiners using real examination rubrics. Results show marked variance across models and tasks: some LLMs match or exceed human passing thresholds on certain exam sections, while all models struggle with tasks requiring deep legal reasoning, jurisdiction-specific knowledge, and nuanced argumentation. The study highlights both the promise and the current limits of LLMs in high-stakes professional assessment contexts, raising implications for AI's role in legal education and certification.

The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and anonymously evaluated by expert examiners, using the same criteria applied in real examinations. Results reveal marked differences across both models and tasks. While some LLMs match or exceed top human performance in adversarial legal argumentation and doctrinal analysis, all models fail in the notary exam, which requires goal-directed legal planning under strict formal and substantive constraints. Beyond ranking models, the study identifies task-specific strengths, limitations and

reports a blind Turing Test evaluating leading LLMs on three Italian professional legal examinations: the Bar exam, Judges exam, and Notary exam. LLMs generated full written papers that were anonymised and graded by expert examiners using real examination rubrics. Results show marked variance across models and tasks: some LLMs match or exceed human passing thresholds on certain exam sections, while all models struggle with tasks requiring deep legal reasoning, jurisdiction-specific knowledge, and nuanced argumentation. The study highlights both the promise and the current limits of LLMs in high-stakes professional assessment contexts, raising implications for AI's role in legal education and certification.

Connected Concepts

  • Benchmark
  • Human In The Loop AI
  • Formative Assessment
  • Automated Essay Scoring
  • Automated Question Generation
  • AI Ed Evaluation
  • Open Source
  • CS Education
  • Connected Articles

  • Machines Misread Pedagogical Quality — Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation
  • Cotal Formative Assessment Scoring 2026 — CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
  • Automated Formative Assessments A Level Sciences — The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences
  • LLM Computational Thinking Physics 2026 — Using LLMs to Detect Growth in Computational Thinking in Introductory Physics
  • Ground Truth Reliability AIED — Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
  • Tutoring Effectiveness Index — The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals
  • Citation

    Bertoli, Germana et al. (2026). What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries. arXiv:2608.06166.