On this page

Overview

Schmucker and Moore examine whether Item-Writing Flaw (IWF) rubrics — a domain-general, textual approach to evaluating test items without student data — have predictive validity for empirical Item Response Theory (IRT) parameters. Traditional validation relies on resource-intensive pilot testing; IWF rubrics offer a scalable pre-deployment alternative. The study analyzes 7,126 multiple-choice questions across STEM subjects (physical science, mathematics, life/earth sciences), using an automated approach (including LLM-based coding) to annotate items.

Key Findings

  • Item-Writing Flaw rubrics show predictive validity for empirical IRT parameters — the presence of flaws relates to item difficulty and discrimination.
  • The IWF approach offers a scalable, pre-deployment evaluation that does not require student data, complementing or partially substituting pilot testing.
  • The method is applied across STEM domains, supporting its domain-general utility.
  • Automated (LLM-assisted) coding enables annotation of large item banks (7,126 questions).

Implications for Practice

  • For assessment developers: IWF rubrics allow early identification of problematic items before costly pilot testing, and predict IRT difficulty/discrimination.
  • For psychometricians: Combining IWF-based screening with empirical IRT analysis can improve item-bank quality efficiently.
  • For researchers: Automated LLM-based coding makes large-scale item-quality evaluation tractable.

Connected Concepts

Connected Articles

Citation

The impact of item-writing flaws on difficulty and discrimination in item response theory — Schmucker, R., & Moore, S. (2026). Computers and Education: Artificial Intelligence, 11, 100632.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.