Responsible assessment in the AI era β assessment grounded in learners' sociocultural contexts and designed to generate valid, trustworthy, context-specific inferences from accumulated evidence, not one-shot outputs. This Stanford Accelerator for Learning + ETS white paper (McGee, Thille, Choi, Ercikan & Hau, 2026, distilled from a January 2026 convening of ~100 education leaders) argues generative AI has broken the assumption that final products measure human capability: learners can produce high-quality artifacts without the underlying learning, AI scoring introduces construct-irrelevant variance, and the gap between what is measured and what matters is widening. The field's response is a shift from testing events to systems of inference β continuous and formative-assessment, portfolio- and conversation-based evidence (socratic-tests-conversational-assessment), authentic-assessment in real tasks, and human-in-the-loop design β paired with validity infrastructure for automated-grading, shared definitions of emerging constructs like ai-literacy, and sustained attention to equity, transparency, and trust.
Nneka J. McGee, Candace Thille, Ikkyu Choi, Kadriye Ercikan & Isabelle C. Hau (2026) β Stanford Accelerator for Learning, Stanford University, with support from ETS.
π Full text (PDF) Β· local
Summary
The report synthesizes a future-focused convening (January 29, 2026) on how assessment should evolve as AI reshapes learning, work, and measurement. It defines responsible assessment as assessment that is grounded in individuals' sociocultural contexts and designed to produce valid, trustworthy, context-specific inferences from accumulated evidence β a shift toward continuous, context-rich, developmentally oriented practices that leverage AI responsibly (building on Johnson, 2025).
Why traditional assessment is evolving. Three interconnected pressures: (1) Misalignment with how learning actually occurs β learning is a process, but assessment remains event-based; conditions that support learning (risk-taking, mistake-making) are absent in assessment mode, and assessment misses the collaborative, AI-supported pathways students actually take. (2) The increasing influence of AI β GenAI means outputs alone can no longer serve as reliable indicators of human capability, so assessment "risks measuring technological proficiency rather than human skill." AI scoring (the most widely used AI application in assessment) can capture construct-irrelevant features (e.g., rewarding language competency instead of scientific reasoning), which becomes bias when it systematically advantages or disadvantages particular learners. Other validity threats: construct underrepresentation, failures of generalization, training data that don't represent authentic child learning (especially K-12), and calibration differences between what humans and AI attend to in the same rubric. Meanwhile AI literacy and "durable skills" (critical thinking, creativity, curiosity, collaboration, agency, adaptability) rise in importance, and workforce shifts (30% of jobs automated within four years per McKinsey; 39% of core skills changing by 2030 per WEF) demand assessments of the skills workers will actually need. (3) A growing gap between what we measure and what matters β measurement signals what systems value; assessment shapes curriculum and public judgments of quality. Advances in data availability (stealth assessment grounded in evidence-centered design, continuous assessment, process/ambient/longitudinal data) expand what can be observed, but more data does not automatically produce better insight, and privacy, learner awareness, and overinterpretation loom. Human-centered skills remain underdefined β AI literacy and adaptability lack shared operational definitions, though the OECD/EC AILit Framework is emerging (feeding the 2029 PISA Media and AI Literacy assessment). Existing/emerging tools: cognitive interviews, live demonstrations, role-playing, scenario-based assessments, portfolios, conversation-based assessments (ECD + AI agents), virtual robotics tasks, and game-based assessment in naturalistic settings (e.g., Guess What?).
What system change requires. (1) Testing events β systems of inference: summative results often arrive too late to help; a development-centered orientation elevates formative assessment and AI agents that generate questions and evidence on demand. (2) Infrastructure and validity for AI-powered assessment: data architectures, interoperable platforms, and sustained human investment are prerequisites β "without infrastructure, even the most advanced AI-supported assessment cannot function effectively"; LLM scoring of constructed responses may be less reliable and more costly once validity-evidence curation is counted, and requires more extensive validity evidence than traditional scoring. (3) Expanding what counts as evidence: relational and social learning (Vygotskian social constructivism, teacherβstudent relationships, collaborative problem-solving) is poorly captured by traditional systems; a broader evidentiary base β multiple sources, inclusive demonstrations (leadership shown through caring for siblings, not just sports captaincies) β is needed. (4) Extending formative assessment: embedded feedback loops, peer formative assessment compared against AI-generated feedback, culturally responsive AI apps (e.g., MOSAIC), and a five-step framework for integrating AI into formative assessment β while acknowledging AI bias, inaccurate information, and the ethics of passive data collection ("the more data we're able to collect passively, the harder those questions are going to become").
Advancing responsible assessment. Reflection: responsibility begins with the people assessment serves β humanizing assessment, keeping humans in the loop, co-designing with diverse stakeholders, recognizing learner uniqueness, and contextualizing design. Trust through transparency: trust depends on what a system can reliably do (Brunskill's microwave analogy β users need faith in safeguards, not full inner mechanics), demonstrated value ("give educators what they don't have"), technical documentation (ETS's automated-scoring docs), adversarial testing and efficacy studies, and limits on transparency (protecting mechanics to preserve trustworthiness; AI extrapolation to small cultural groups produces nonsensical outputs). Action across four stakeholder groups: education systems should pilot continuous embedded assessment, adopt portfolio/competency-based approaches, reduce reliance on high-stakes one-time testing, build educator capacity, and engage learners/families; researchers should clarify constructs, advance ecological validity, develop validity standards for AI-assisted assessment, and leverage new evidence forms responsibly (e.g., the MAGIC Project measuring curiosity and inquiry); developers should design for interpretability, support learning processes not just outputs, enable human-in-the-loop workflows, and avoid fully automated high-stakes decision-making; funders should invest in new constructs and measurement models, shared infrastructure (longitudinal data systems, synthetic data environments), translation, and coordinated mechanisms.
Conclusion. Assessment must evolve without losing fairness and validity. High-stakes assessment retains an essential role, but assessment must also show how learners progress over time. "AI can expand what is feasible, but it cannot be held accountable for the consequences of assessment decisions. People can."
Connections to the Wiki
- Assessment validity & AI scoring β the report's catalog of validity threats (construct-irrelevant variance, underrepresentation, generalization, calibration) is a policy-level complement to the wiki's empirical evidence on automated-grading failures (e.g., llm-handwritten-math-grading, ai-scoring-language-bias-physics, machines-misread-pedagogical-quality).
- Formative assessment β the call to extend formative-assessment with AI (peer feedback vs. AI feedback, culturally responsive tools) connects to ai-generated-feedback-higher-ed and feedback-loop research.
- Conversation-based assessment β ECD-based AI-agent dialogue assessment aligns with socratic-tests-conversational-assessment and intelligent-tutoring design.
- Construct definitions β the underdefined AI literacy / durable-skills problem echoes ai-literacy debates and educational-theory work on what AI-era competencies mean operationally.
- Equity & trust β sociocultural responsiveness, bias in AI scoring, and human accountability map to equity, human-in-the-loop, and ai-ed-evaluation.
Related Pages
- assessment β Assessment is at a moment of change; the report reframes validity and purpose for the AI era
- assessment-validity β Validity threats from AI scoring (construct-irrelevant variance, calibration, generalization)
- formative-assessment β Extending formative assessment with AI feedback loops and embedded evidence
- automated-grading β AI scoring is the most widely used AI application in assessment, with new validity costs
- ai-literacy β AI literacy as an emerging construct in need of operational definition
- authentic-assessment β Live demonstrations, portfolios, and scenario-based assessments as authentic evidence
- equity β Socioculturally responsive assessment and bias in AI-mediated scoring
- human-in-the-loop β Keeping humans in the loop for high-stakes AI assessment decisions