On this page

Synthesis: This study investigates how middle-school science teachers evaluate AI-generated assessment questions and how they perceive the role of collaborative evaluation in professional development. Sixty science teachers reviewed open-ended questions on lower- and higher-order thinking skills produced by ChatGPT, rating them through their disciplinary, pedagogical, and curricular judgment. The mixed-methods findings show that collaborative evaluation helps teachers apply conceptual-precision judgment to AI content and surface risks such as reinforcing Misconceptions. The paper underscores the value of professional development that positions teachers as critical evaluators of AI-generated material rather than passive consumers.

Background

As GenAI becomes integrated into education, AI-generated content risks reinforcing misconceptions and propagating imprecise reasoning — particularly in disciplines that demand conceptual precision. Question writing for assessment therefore requires critical evaluation by educators, but little is known about how teachers exercise that judgment collectively.

Key Findings

  1. Sixty middle-school science teachers reviewed ChatGPT-produced questions spanning lower-order and higher-order thinking skills.
  2. Teachers drew on disciplinary, pedagogical, and curricular judgment to rate AI-generated content, applying conceptual-precision criteria.
  3. Collaborative evaluation emerged as valuable professional development, helping teachers surface and address the risk that AI content reinforces misconceptions.
  4. A mixed-methods design combined quantitative ratings with qualitative insights on teachers' perceptions.
  5. The findings position teachers as critical evaluators of AI material, supporting critical engagement with GenAI in the classroom.

Implications

The study supports professional development designs that build teachers' capacity to critically evaluate AI-generated content, strengthening teacher judgment and guarding against the propagation of misconceptions through AI tools. It also connects to Formative Assessment design, where the quality of AI-generated questions depends on informed educator oversight.

Connected Concepts

Connected Articles

Citation

Gat, I., Usher, M., & Barak, M. (2026). Teachers' Collaborative Evaluation of AI-Generated Content: Insights from a Professional Development Workshop. EdArXiv preprint.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.