On this page

Synthesis: Doğru and Faulconer (2026) tested ChatGPT as a virtual teaching assistant (VTA) in an introductory undergraduate biology laboratory course, comparing its responses to student-generated questions about an enzyme-activity lab exercise against those of a cohort of graduate teaching assistants (TAs). Evaluated by subject-matter experts and students, human TAs were more accurate and demonstrated more substantial teaching effectiveness (instructor voice, understanding, helpfulness). Notably, students preferred the computer-generated response 40% of the time and could identify an AI-generated response only 45% of the time. The authors conclude ChatGPT can help lift the burden of basic questions from TAs and instructors, but students must understand its error rate and a safety net is needed to prevent proceeding on incorrect information — a real safety concern in the biology laboratory.

Design

  • Setting: introductory undergraduate biology laboratory course.
  • Method: student-generated questions about a lab exercise on enzyme activity were separately given to ChatGPT and a cohort of graduate TAs.
  • Evaluation: responses were rated by subject-matter experts (SMEs) and students on content accuracy and teaching effectiveness (instructor voice, understanding, helpfulness).

Findings

  • Accuracy: human TAs provided more accurate responses than ChatGPT.
  • Teaching effectiveness: human TAs demonstrated more substantial teaching effectiveness across instructor voice, understanding, and helpfulness — significant on most items for both SME and student ratings.
  • Preference and detection: students preferred the computer-generated response 40% of the time, and could identify an AI-generated response only 45% of the time — a meaningful share of AI output was both preferred and undetected.
  • SMEs: ChatGPT's overall content accuracy was rated acceptable but with a notable error rate on specialized biology terminology and concepts.

The biology-specific challenge

The authors note LLMs may struggle with specialized terminology and concepts in biology, leading to incorrect or misleading information — a concern amplified in a laboratory context where inaccurate information can present safety risks. They draw parallels to prior findings that ChatGPT gives incomplete/misleading answers to physics questions and errors on complex tasks.

What this means for practice

  • Instructors. Limit a ChatGPT VTA to basic, low-stakes questions and keep a human safety net: in a biology laboratory, incorrect information can present safety risks.
  • Instructors. Tell students the error rate plainly — students preferred the AI response 40% of the time and identified it only 45% of the time, so preference and fluency are no signal of accuracy.
  • Instructors. Route specialized terminology and concepts to human TAs; subject-matter experts rated ChatGPT's accuracy acceptable but flagged a notable error rate on specialized biology content.
  • Instructional designers. Design the VTA to offload routine questions while building in fact-checking prompts, so students do not substitute AI answers for laboratory reasoning and critical thinking.
  • Instructional designers. Point students to the human TA when they need help understanding a process, since human TAs showed stronger instructor voice, understanding, and helpfulness.

Limitations

  • Participants came from a single institution — 65 pre-service science teacher candidates plus 8 subject-matter experts — which limits the generalizability of the findings.
  • The design is exploratory and cross-sectional with no control group; the data cover May–June 2023 only and cannot show change over time.
  • The comparison rests on eight student-generated questions, and the authors note that the limited quantity and the subjectivity of question selection may affect generalizability.
  • Students may lack the expertise to judge content accuracy, and authorship misclassification — only about 45% of AI answers were correctly identified — may itself have shaped the ratings.

Citation

Doğru, M. S., & Faulconer, E. K. (2026). ChatGPT as a virtual laboratory teaching assistant in undergraduate biology. Research in Science Education, 56, 379–399.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.