---
source_url: https://acceleratelearning.stanford.edu/conference/responsible-assessment-in-the-ai-era/
ingested_date: 2026-08-03
sha256: adeeb44444037b653e409c26cafecfcf371f31b5e5f1a3e635c16e3e036aefbf
---

# Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference

RESPONSIBLE ASSESSMENT
            IN THE AI ERA
               Key Insights from a
       Future-Focused Conference




                                 Nneka J. McGee
                                  Candace Thille
                                      Ikkyu Choi
                                 Kadriye Ercikan
                                  Isabelle C. Hau
                   Matthew S. Johnson (Reviewer)
Table of Contents

1 .Introduction: Assessment at a Moment of Change    5

1.1 Opening the Lens on Responsible Assessment       5

2. Why Traditional Assessment is Evolving            7

2.1 Misalignment with How Learning Actually Occurs   7

2.2 The Increasing Influence of AI                   8

2.3 A Growing Gap Between What We Measure
     and What Matters                                10

3. What System Change Requires: Rethinking
   Assessment                                        12

3.1 Testing Events to Systems of Inference           12

3.2 Building Infrastructure and Validity
    for AI-Powered Assessment                        13

3.3 Expanding What Counts as Evidence                13

3.4 Extending the Role of Formative Assessment       14

4. Advancing a Focus on Responsible Assessment       15

4.1 Reflecting on Responsible Assessment             15

4.2 Establishing Trust through Transparency          15

4.3 Translating Insight into Action                  16

Conclusion                                           19

Glossary                                             20

Endnotes                                             21
Acknowledgments
We would like to express our appreciation to the following individuals for their support and contributions
to the Responsible Assessment in the AI Era convening:

Daniel L. Schwartz the I. James Quillen Dean, and Nomellini & Olivier Professor of Educational
Technology at Stanford Graduate School of Education, and The Halper Family Faculty Director of the
Stanford Accelerator for Learning.

Amit Sevak Chief Executive Officer, ETS.

Candace Thille Associate Professor, Stanford Graduate School of Education, Faculty Director, Adult and
Workforce Learning Initiative, Stanford Accelerator for Learning.




Collaborating Organization
The convening was supported by Educational Testing Service (ETS), an organization whose mission
focuses on “advancing the science of measurement to power human progress.”




Recognition of Contributions
We would also like to extend our gratitude to panelists, moderators, poster session presenters, and
breakout session contributors. Their discussions informed the themes, tensions, and future-forward
moves reflected throughout this paper.


Panelists
Bryan Brown, Stanford University                             Matthew S. Johnson, ETS*
Emma Brunskill, Stanford University                          Victor Lee, Stanford Graduate School of Education
Sara Caldwell, Open AI                                          and Stanford Accelerator for Learning
Tony Chan, Former President, KAUST*                          Lydia Liu, ETS
Kristen DiCerbo, Khan Academy                                Temple Lovelace, Assessment for Good,AERDF
Ben Domingue, Stanford Graduate School of                    Nancy Otero, Gates Foundation
   Education*                                                Kadriye Ercikan, ETS
Maureen Heymans, Google*                                     Jesse R. Sparks, ETS*


Moderators
Mille Garcia, California State University                    Darienne Hudson, United Way, Southeastern MI
Kevin Gutherie, ITHAKA                                       Beverly Daniel Tatum, Spelman College




*panelist or presenter and breakout session contributor
**graduate student at Stanford GSE and breakout session contributor
              Poster Session Presenters
              ETS                                                          Stanford Graduate School of Education
              Ikkyu Choi*                                                  Yunsung Kim**
              Patrick Kyllonen*                                            Hansol Lee**
              Matthew S. Johnson*                                          Eunjung Myoung
              Sandip Sinharay*                                             Rojas Pino
              Diego Zapata-Rivera*                                         Marcos Santiago

              *panelist or presenter and breakout session contributor
              **graduate student at Stanford GSE and breakout session contributor



              We thank all attendees who joined and contributed to the spirit of inquiry. Their engagement reminded
              us that responsible assessment in the AI era is not the work of one organization or discipline, but a shared
              effort across communities committed to learning.




              Authors
              Nneka J. McGee, J.D., Ed.D., Project lead and principal co-author for the “Responsible Assessment in the
              AI Era: Key Insights from a Future-Focused Conference” white paper.

              Candace Thille, Ed.D., Associate Professor, Stanford Graduate School of Education, Faculty Director,
              Adult and Workforce Learning Initiative, Stanford Accelerator for Learning.

              Ikkyu Choi, Ph.D., Senior Research Scientist, ETS

              Kadriye Erkican, Ph.D., Senior Vice President of Global Research, ETS; President, ETS Canada Inc.

              Isabelle C. Hau, Executive Director, Stanford Accelerator for Learning

              Recommended Citation: Nneka J. McGee, Candace Thille, Ikkyu Choi, Kadriye Ercikan, and Isabelle C.
              Hau. “Responsible Assessment in the AI Era: Key Insights from a Future-Focused Convening,” Stanford
              Accelerator for Learning, Stanford University, 2026.

              Reviewer: Matthew S. Johnson, Ph.D., Principal Research Scientist, ETS Research Institute



              Disclaimer: This paper reflects the perspectives of the authors and includes quotes from the “Responsible
              Assessment in the AI Era” convening. It does not necessarily represent the views of Stanford University, the
              Stanford Accelerator for Learning, ETS, participants or collaborating organizations.

              A note on AI: This paper was developed and written by the principal authors, who reviewed and fact-
              checked all content. ChatGPT (GPT 5.5) was used sparingly for minor editorial refinement. Sembly AI was
              used to transcribe and summarize breakout session discussions.




4 Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
1. Introduction: Assessment at a Moment of Change

“H
        ow do we shape the current moment?” The
        question, from Stanford Professor and Stanford
        Accelerator for Learning Faculty Co-Director
Candace Thille, framed a broader discussion about                   “We want to have conversations
how to approach assessment in the era of artificial                  that are rich and deep about how
intelligence (AI). In search of answers, the Stanford
Accelerator for Learning, in collaboration and with                  research, policy, technology and
support from ETS, hosted the “Responsible Assessment                 education come together.”
in the AI Era” convening on January 29, 2026.
                                                                        — Amit Sevak, CEO of ETS
“We want to have conversations that are rich and deep
about how research, policy, technology and education
come together,” Amit Sevak, CEO of ETS, said in his
opening remarks. The convening drew approximately
100 leaders from research, technology, K-12 schools,             1.1 Opening the Lens on Responsible
and higher education institutions seeking to advance             Assessment
the science and practice of responsible assessment
in the age of AI. There was a consensus among them:              The conference took place at a critical juncture for
Assessment is at a moment of change, and AI is a major           assessment. There are growing calls for assessments
driver of that transformation.                                   to become more socioculturally responsive to learners’
                                                                 cultural, educational, and societal contexts in order to
Panelists shared wide-ranging perspectives on                    optimize engagement, motivation and performance.
evidence, rethinking readiness, socioculturally                  Through personalization, AI may help address some of
responsible assessment, and building trustworthy AI.             these goals.
Through poster sessions, scholars presented findings
from their latest studies. Contributors expressed their          At the same time, the integration of AI into education
viewpoints at breakout sessions on topics ranging from           and assessment raises questions about which
AI policy to extracting construct evidence. While this           competencies learning and assessment should
paper primarily reflects the insights and perspectives           prioritize, and how to leverage the potential of AI while
shared by speakers at the event, it also draws on                minimizing the risks of exacerbating opportunity
relevant research and related work to contextualize the          gaps for individuals from different societal and
contributors’ dialogue. It summarizes key takeaways              cultural backgrounds. The conference brought these
from the convening and includes a glossary of terms              concepts together by focusing on how assessment
related to assessment.                                           can become both responsive and responsible for




                                          Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference 5
                                                            Socioculturally
                                                             Responsive




                                       Supports                                          Focuses on
                                                              Responsible                Most Critical
                                     Learning and
                                                              Assessment                Competencies
                                     Opportunity
                                                                                        for Education




                                                               Leverages
                                                             Responsible AI
                                                               for Better
                                                             Measurement




Figure 1: Responsible Assessment



meeting evolving educational goals. This approach                           Videos of key sessions at the event can be watched
promotes responsible AI by supporting valid and fair                        here: https://acceleratelearning.stanford.edu/
interpretations and uses of assessment (See Fig. 1).                        conference/responsible-assessment-in-the-ai-era/

In this paper, we use the term “responsible assessment
in the AI era” to refer to assessment that is grounded
in individuals’ sociocultural contexts and designed
to generate valid, trustworthy, and context-
specific inferences from accumulated evidence to
support learning and development. Responsible
assessment is a call that represents a shift toward
continuous, context-rich, and developmentally
oriented assessment practices that leverage AI in
responsible ways, whose principles and practices are
detailed in Johnson (2025).1 It emphasizes a focus
on complex critical competencies, accumulation of
multiple sources of evidence through socially situated
performance over time, and the consideration of
sociocultural context. It also seeks to support learning
and growth while prioritizing transparency and
interpretability in how evidence is used, elicited, and
accumulated.




6 Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
2. Why Traditional Assessment is Evolving

A
        ssessment is used to measure knowledge
        and skills, monitor growth over time, support
        learning, inform accountability, and guide
decisions about admissions and program access.2 In                     The promise of AI depends on how
practice, traditional assessments frequently operate
at the intersection of improvement and accountability,
                                                                       assessment helps educators see and
where the same measures are expected to support                        respond to learners’ next steps.
learning while also serving evaluative and comparative
functions.3 This section discusses three interconnected
factors shaping the evolution of assessment amid
multiple, sometimes competing demands.

                                                                   Temple Lovelace, executive director of assessment for
2.1 Misalignment with How Learning                                 good at Advanced Education Research & Development
Actually Occurs                                                    Fund (AERDF) described the value of providing space
                                                                   for learners to show up as their full selves in a learning
“Traditional assessments may tell us if a student
                                                                   environment. What students bring to learning follows
arrived at a correct answer,” stated Millie García,
                                                                   them into assessment. “Full self includes family
Chancellor of the California State University. “But
                                                                   history, home context, and societal context,” explained
they often tell us little about how [students] got
                                                                   Kadriye Ercikan, senior vice president of global
there. In today’s collaborative digital and increasingly
                                                                   research at ETS. “This affects how they relate to the
AI-supported communities, that gap matters.” While
                                                                   assessment…[and] how they engage and make sense
learning unfolds as a process, assessment often
                                                                   of the assessment.” Although standardization is often
remains largely event-based. This dynamic creates a
                                                                   intended to even the playing field, Ercikan suggested
mismatch between the realities of learning and the
                                                                   that equivalent assessment conditions are lacking.
way assessment is commonly designed.
                                                                   Many current systems still emphasize end products
This mismatch is most evident when learning and                    over the pathways that produce them.
assessment create varied conditions for students.
                                                                   AI-supported learning environments may create new
Cindy Mazow, director of learning technology and
                                                                   opportunities to address this misalignment, but only
design at the Stanford Graduate School of Business,
                                                                   if they are designed around sound learning principles.
emphasized that effective learning requires room
                                                                   Maureen Heymans, vice president of learning at
for students to make mistakes or take risks because
                                                                   Google, suggested that “learning is most effective
those actions can expose misconceptions or gaps
                                                                   when it is active, engaging, and personalized.” When
in understanding which can then be corrected. In
                                                                   connecting learning to Vygotsky’s Zone of Proximal
assessment mode, however, Mazow indicated that
                                                                   Development, Josh Arnold, co-founder of Impacter
students “don’t want to make a mistake because they
                                                                   Pathway, suggested that learners can do “hard and
know that there’s a consequence to that assessment
                                                                   gritty things” when those things are relevant and right-
whether it’s high-stakes or low-stakes.” The conditions
                                                                   sized to learners’ current abilities.4
that support learning, then, are not always present
when students are being assessed.




                                            Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference 7
In practice, the promise of AI depends on how             understanding”6 As Thille of the Stanford Accelerator
assessment helps educators see and respond to             for Learning said, when assessment no longer
learners’ next steps. “What do we change on the           measures what it was meant to measure, the validity
ground when it comes to assessment and learning?”         of inferences based on assessment outcomes may
Lovelace asked, connecting the broader discussion to      be compromised.
her experiences as a special education teacher. Her
question points toward a future where AI-supported        AI is also shaping how assessment evidence is
assessment is more integrated with learning. The          scored and interpreted. Ercikan of ETS specified
challenge is to strengthen learner support without        that scoring is the most widely used application
turning every moment of learning into another layer       of AI in assessment. This creates validity concerns
of assessment.                                            when AI-based scoring, such as automated
                                                          scoring systems, captures features of performance
                                                          unrelated to the intended construct. Ercikan offered
2.2 The Increasing Influence of AI                        an example of a science assessment in which
AI affects multiple points of the assessment process,     AI scoring rewards language competency rather
including how student work is produced and how that       than scientific reasoning to describe construct-
work is evaluated. As the use of AI increases among       irrelevant variance.
learners, there are potential disconnects between
what assessments are designed to measure and what         When such irrelevant variance systematically
they actually capture. The technology also impacts the    advantages or disadvantages particular learners, it
skills and competencies deemed important at schools       becomes a fairness concern and, as Ercikan noted, “is
and in the workforce.                                     at the core of bias in assessment.”7 Stanford Graduate
                                                          School of Education Professor Guillermo Salano-
2.2.a Validity                                            Flores has identified three areas where bias can be
                                                          introduced: language elements (e.g. words), visual
Generative AI (GenAI) has exposed the limitations
                                                          components (e.g. charts), and contextual scenarios
of traditional assessments based primarily on final
                                                          (e.g. word problems).8 Bias threatens validity when
outputs.5 Learners can now use GenAI to generate
                                                          parts of the assessment are not central to what is
high quality products without fully engaging in the
                                                          being measured.9
learning processes those products are meant to
represent. As a result, outputs alone can no longer       Contributors shared additional AI-influenced
be assumed to serve as reliable indicators of human       assessment practices that impact validity, including:
capability. In these cases, assessment “risks measuring
technological proficiency rather than human skill or
• Construct underrepresentation: Ercikan explained                     As AI systems take on more routine
  this can occur when an AI scoring system only
  partially captures what the assessment is intended                   cognitive tasks, assessment will
  to measure.10                                                        need to focus more deliberately on
• Generalization: Andrew McEachin, senior research                     capabilities that are uniquely human.
  director at ETS, described how AI systems can fail
  to generalize across contexts and how a human’s
  ability to recognize when AI systems generalize is
  not well-developed.11
                                                                   Together, these examples show that assessment
                                                                   priorities are changing. AI literacy is not sufficient by
• Training data: McEachin also expressed concern
                                                                   itself. As AI systems take on more routine cognitive
  about AI models that are not trained on data
                                                                   tasks, assessment will need to focus more deliberately
  representing authentic child learning, particularly in
                                                                   on capabilities that are uniquely human. They are
  K–12 education.12
                                                                   often described as “durable skills” and include critical
                                                                   thinking, creativity, curiosity, collaboration, agency,
• Calibration: Matthew S. Johnson, managing
                                                                   and adaptability.15 Emotional intelligence, relational
  director of innovative research at ETS, used the
                                                                   intelligence, ethical judgment, and challenging
  example of rubrics to highlight how what humans
                                                                   assumptions have also been identified as uniquely
  and AI systems take into account can differ on the
                                                                   human skills.16 The Stanford Accelerator for Learning
  same assessment.13
                                                                   examined these shifting priorities in Future-Ready
This list of validity threats is not exhaustive, but it            Voices, a November 2025 interview series featuring
shows how AI can influence assessment at multiple                  researchers, educators, and thought leaders.17
points, from the production of student work and the
scoring and interpretation of that work.
                                                                   2.2.c AI and the Workplace
                                                                   Job availability and workforce displacement remain
2.2.b AI literacy and essential competencies                       pressing challenges in the AI era. Citing a McKinsey
During a panel discussion on measuring essential                   report, Darienne Hudson, president and CEO of United
competencies, Jesse R. Sparks, research director                   Way for Southeastern Michigan, predicted that 30%
at ETS Research Institute, pointed to AI literacy                  of jobs will become automated within four years.18
as an emerging construct. Its rise reflects AI’s                   According to the World Economic Forum, employers
increasing influence across education and work.                    expect 39% of workers’ core skills to change by 2030.19
Sparks emphasized that redesigned assessments                      Regarding specific professions, Caldwell said, “It’s
are needed to capture key abilities associated with                much less about that an engineer won’t exist, but
AI literacy, including collaborative interaction and               what aspects of that role are going to fundamentally
critical evaluation.14                                             change.” In other words, assessment should focus on
                                                                   which skills workers will need to succeed in changing
The conversation also moved beyond AI literacy                     work environments.
to consider additional competencies learners
will need in the age of AI. Sparks, Sara Caldwell,                 Sparks envisioned a workplace where “we interact
head of GTM Readiness at OpenAI, and Victor Lee,                   with AI tools in ways that really enhance the strengths
associate professor at the Stanford Graduate School                that we’re bringing to the situation.” Caldwell used
of Education and faculty lead for AI+Education at                  as an example the ability to orchestrate a set of AI
the Stanford Accelerator for Learning, identified                  agents. Lydia Liu, associate vice president of ETS,
competencies such as systems thinking and deep                     highlighted a collaboration between ETS and OECD
domain expertise as especially important. Caldwell                 named PISA VET (Vocational Education and Training).20
suggested that people who combine their own                        The international initiative aims to benchmark
competencies with AI will be better positioned in the              performance across countries to identify factors that
workforce than those who rely on AI alone.                         are associated with positive outcomes in vocational
                                                                   education. Some of the AI-related innovations Liu
                                                                   described include:



                                            Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference 9
• VR glasses to observe if automotive technician                      assessment must broaden to show how capabilities
  students can perform critical tasks related to                      develop and appear in practice.
  the job such as inspections and diagnoses, and
  maintenance and repair, and                                         2.3.a Advances in Data Availability
                                                                      AI is expanding what can be observed and measured
• Pressure sensing gloves to determine if trainees
                                                                      about learning and performance. Throughout the
  can apply the right amount of pressure to complete
                                                                      panels and breakout sessions, stealth assessment
  a task.
                                                                      was discussed as one example of how AI is advancing
                                                                      existing approaches to gathering evidence.
Liu’s descriptions of immersive technologies reveal
                                                                      Sparks described stealth assessment as a process
a gap between what many traditional assessments
                                                                      methodologically grounded in evidence-centered
measure and the types of skills increasingly valued
                                                                      design (ECD) where educators and researchers
at work. Applications of new technologies and tools
                                                                      “continuously extract signals from a learner’s everyday
opened the door for more comprehensive and
                                                                      learning environment.” Contributors also discussed
authentic assessment. As expectations change, so do
                                                                      how systems can draw on process data, ambient data,
signals of readiness, as evidenced by the proliferation
                                                                      longitudinal data, and linked systems that connect
of alternative certifications, skill verifications, and
                                                                      assessment results with broader patterns of learning
micro-credentials.21 These signals reflect rising
                                                                      and performance.
demand for evidence that is more closely tied to
demonstrable capability.22
                                                                      Patrick Kyllonen, distinguished presidential appointee
                                                                      at ETS, highlighted continuous assessment,
2.3 A Growing Gap Between What We                                     which collects data to measure student progress
Measure and What Matters                                              continuously over time. He suggested that this
                                                                      approach may reduce assessment fatigue and test
“Traditional assessment is misleading,” said Daniel L.                anxiety. Advanced analytic methods are making it
Schwartz, dean of the Stanford Graduate School of                     possible to interpret patterns from those data in
Education and the Halper Family Faculty Director of                   new ways.
the Stanford Accelerator for Learning. He described
the implications of a study on how students adapt                     However, simply increasing the amount of data is not
knowledge to a novel situation.23 “We need instruction                sufficient. To paraphrase Julia Wilkowski, who directs
that produces adaptive learners and assessment that                   learning sciences at Google, more evidence does not
can tell.”                                                            automatically produce better insight.

Beneath the immediate concerns about AI and                           Assessments such as stealth assessment and
assessment is a more enduring issue expressed by                      continuous assessment raise concerns about how
OpenAI’s Caldwell: Measurement signals what systems                   learner data are collected, interpreted, and protected.
value. That relationship is not neutral. Assessment                   Vikas Wadhwani, a leader of the learning and
shapes not only how learning is measured, but also                    certifications team at Meta, identified privacy as one
what educational systems prioritize.24 A large-scale                  challenge. Another is learner awareness, since students
assessment used for comparability and accountability                  may change their behavior when they know they are
can influence curriculum, instruction, and public                     being assessed. Synthetic data and simulated learner
judgments of quality.25                                               models, sometimes described as “simulated students,”
                                                                      could offer a way to test assessment designs, study
Over time, systems establish priorities, assessments                  possible response patterns, and examine item
operationalize them, and repeated use reinforces what                 performance while reducing reliance on identifiable
counts as evidence. While these approaches enable                     student data.26 Simulated students may support
common metrics across large populations, they can                     faster experimentation, but validation would still be
also narrow how and what learning is represented. If                  needed to ensure simulated patterns do not distort
educational systems place more value on skills and                    assumptions about real learners.
learning processes, then the data gathered through




10 Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
2.3.b Measuring Human-Centered Skills
Human-centered skills, such as durable skills, were
described briefly earlier in this paper. There is
some consensus that human-centered skills can be
assessed. However, methods for measuring them
remain less established, particularly because they are
“interpersonal, intrapersonal, or behavioral
in nature.”27

2.3.b.i. Underdefined meanings for AI literacy and human-
centered skills
Stanford Accelerator for Learning’s Faculty Lead for
AI+Education Victor Lee recognized that AI literacy
is one of the core competencies for the AI era, but                    Wide applications of live demonstrations will require
stresses that “we don’t really have a clear articulation               operational efficiency for feasibility.
for what that is.” The AILit Framework, a joint initiative
from OECD and the European Commission, includes                     • Role-playing: ETS’s Patrick Kyllonen outlined an
a comprehensive definition of AI literacy and will                    initiative where students interact with AI as a role-
contribute to the 2029 PISA Media and Artificial                      playing partner to measure social skills.33
Intelligence Literacy assessment.28 Even so, definitions
                                                                    • Scenario-based assessments: Kyllonen also
of AI literacy continue to vary across recognized
                                                                      explained how AI can be used in assessment
institutions and organizations.29
                                                                      environments to generate scenarios that reflect
A similar issue appears with adaptability. Although                   non-traditional, real-world situations.34
widely named as essential in future-focused, AI-rich
                                                                    • Portfolio-based assessments: Contributors
environments, adaptability lacks a standard definition
                                                                      discussed evaluations that allow learners to
with operational indicators across AI-use scenarios.30
                                                                      demonstrate growth and real-world application of
These constructs must be carefully defined in
                                                                      knowledge through a curated collection of work
operational terms if assessments are to collect
                                                                      over time, rather than through a single test or exam.
meaningful data to measure.31
                                                                    • Conversation-based assessments: Diego Zapata-
2.3.b.ii. Existing and emerging tools to measure human-
                                                                      Rivera, distinguished president appointee at the ETS
centered skills
                                                                      Research Institute, described an evidence-centered
Contributors characterized the measurement of                         design (ECD) based approach that uses AI agents
human-centered skills as part of a larger shift from                  to help learners demonstrate knowledge and skills
controlled assessment environments toward more                        through dialogue.35
authentic, open-ended interactions. Some of the
tools discussed are already established in practice,                • Virtual robotics programming tasks: Sparks, of
while others are newer or remain in varying stages of                 ETS, detailed emerging work that examines the
research and development. Here is sampling of the                     potential for identifying signals of persistence and
tools referenced during the convening:                                resilience from learners’ interactions with virtual
                                                                      robots in block-based coding environments.36
• Cognitive interviews: Lydia Liu of ETS suggested
  this format to gain a better understanding of                     • Game-based assessment in naturalistic settings:
  students’ cognitive process when interacting with                   Guess What?, a smartphone app created by Stanford
  assessment or learning tasks, particularly when                     Professor and Stanford Accelerator for Learning
  considering cultural or behavioral cues.32                          faculty affiliate Dennis Wall, uses game-based
                                                                      interactive computer vision to identify early signs of
• Live demonstrations: In addition, Liu described                     autism in naturalistic settings.37
  this format as the most authentic way to collect
  evidence about what learners know and can do.


                                            Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference 11
3. What System Change Requires: Rethinking Assessment


S
       ystem change presents a challenge for how                      Johnson, also of ETS, said motivation is another
       assessment can support the complexities                        limiting factor because learners question the reasons
       of modern learning. Learning is continuous,                    for summative assessments for which they may never
adaptive, contextual, and increasingly shaped by                      see the results.
human and technological interaction. That complexity
makes it difficult to rely on assessment models                       A different orientation places learner development
designed primarily to capture performance at a                        at the center of assessment. This orientation focuses
single point in time. There are multiple pathways to                  on formative assessment that provides ongoing
rethinking assessment, and AI will play an important                  feedback rather than a snapshot of performance and
role in expanding the evidence available to guide                     draws on evidence that can illuminate “the course
interpretations, noted Karin Forssell, senior lecturer                of development and the status of skills.”38 Once the
with the Stanford Graduate School of Education and                    attention on assessment shifts from summative
director of the AI Tinkery at the Stanford Accelerator                assessment to learner development, the process of
for Learning.                                                         gathering and utilizing evidence must be reconsidered.

                                                                      AI may increase the feasibility of formative assessment
3.1 Testing Events to Systems of                                      by synthesizing evidence and identifying patterns
Inference                                                             within large and complex data streams, which in
Summative assessment is designed to evaluate                          turn can contribute to more refined inferences about
learning at the culmination of an instruction period.                 learning. Zapata-Rivera used the principles of ECD and
Despite its importance, ETS’ Sparks highlighted a                     the Autotutor dialogue framework to describe how
limitation: the results of some summative assessments                 AI agents could support this type of assessment. “In
arrive too late to be useful to learners and educators.               evidence centered design, each piece of each task is




12 Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
designed for a particular purpose,” explained Zapata-            in AI-mediated environments. For example, large
Rivera. “With AI, we could create expert agents to               language model (LLM) scoring of constructed
generate questions or create opportunities even if we            responses may require more extensive validity
are not there.”                                                  evidence than traditional approaches due to concerns
                                                                 about consistency and interpretability.40 In addition,
                                                                 using LLMs as opposed to traditional scoring methods
3.2 Building Infrastructure and Validity                         could be less reliable and more costly due, in part, to
for AI-Powered Assessment                                        the expense involved in curating validity evidence.41
Christian Pinedo, vice president of external affairs
and advocacy at aiEDU, emphasized that building                  3.3 Expanding What Counts as Evidence
comprehensive infrastructure “requires a lot of
human investment.” Systems that incorporate AI,                  The transformation of assessment requires a
including data architectures and interoperable                   broader evidentiary base. For example, relational
platforms, must have the infrastructure to support the           learning builds on Vygotsky’s social constructivism
collection, integration, interpretation, and storage of          by emphasizing strong teacher-student relationships
large amounts of data. In response, LeeAnn Lindsey,              in classroom settings.42 In a panel on rethinking
director of edtech and innovation at Northern                    measurement, Liu of ETS indicated that traditional
Arizona University’s Arizona Institute for Education             assessment systems were not designed to capture
and the Economy, added that many state and                       these dimensions of learning well.43 Social and
district budgets are stretched while acknowledging               relational learning often involve complex skills
the “lack of basic infrastructure for implementing               that have cognitive, affective, and behavioral
good decision-making through assessment.” Their                  components, but also depend on capacities for
exchange points to the need for sustained investment             dialogue, interactions, building trust, and collective
in assessment infrastructure, even in constrained                problem-solving. A future-oriented taxonomy
fiscal environments. Without such infrastructure, even           from Liu and her colleagues illustrate this breadth
the most advanced AI-supported assessment cannot                 by outlining 30 skills ranging from adaptability to
function effectively.                                            transformative competencies.44 Liu also emphasized
                                                                 the importance of gathering evidence from multiple
The value of robust digital data systems depends on              sources and creating an inclusive system for students
whether the claims, interpretations, and decisions               to demonstrate their skills. Leadership, for instance,
they support remain valid.39 Therefore, investment               may be shown not only from being a captain on varsity
in assessment infrastructure should be paired with               sports teams, but also from taking care of younger
investment in collecting ongoing validity evidence               siblings or leading community activities.

                                                                 Emerging research suggests that a broader evidentiary
                                                                 base would be beneficial in measuring social learning,
                                                                 even if the work remains preliminary. Positive
                                                                 outcomes were observed when students engage in
    Without infrastructure,                                      inquiry and constructive collaboration suggest that
                                                                 social processes can be documented and interpreted
    even the most advanced                                       in analytically useful ways.45 Contributors also
    AI-supported assessment                                      discussed developing storylines about future-focused
                                                                 scenarios, such as Participatory Scenario Planning,
    cannot function effectively.                                 which can be used to enable the social learning
                                                                 process.46 Researchers acknowledge limitations in
                                                                 scope, but they point toward a potentially important
                                                                 direction for assessment by expanding the evidentiary
                                                                 base to include social and relational learning.




                                         Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference 13
3.4 Extending the Role of Formative                       Even researchers advocating for formative assessment
Assessment                                                acknowledge the risks of AI use. AI can exhibit bias
                                                          resulting in outcomes that minimize or exclude
Formative assessment embeds assessment in                 learners from the learning process.51 The technology
instruction to guide next steps. Teaching and learning    still generates inaccurate information, which can skew
must be interactive, with teachers responding to          evidence for learning.52 Ben Domingue, associate
student progress and difficulties as they adapt           professor at Stanford Graduate School of Education
instruction.47 Contributors noted that this feedback      and faculty affiliate at the Stanford Accelerator for
loop can also include peer formative assessment, in       Learning, cautioned that passive data collection
which students exchange feedback and compare it           in some types of formative assessment raises
with AI-generated feedback.48                             increasingly difficult ethical, privacy, and transparency
                                                          questions. “The more data we’re able to collect
In a panel discussion on the potential of AI to help or   passively,” Domingue said, “the harder those questions
harm socioculturally responsible assessment, Bryan        are going to become.”
A. Brown, professor of science education at Stanford
Graduate School of Education, emphasized that             A stronger role for formative assessment does not
formative assessment should meet students where           diminish the need for summative assessment.
they are. To support that goal, Brown is working on       The goal is to ensure that each serves a clear and
Multicultural Oriented Science Adaptive Intelligence      complementary purpose.53 Instead of one assessment
for Classrooms (MOSAIC), a culturally responsive          having the final judgment on learners’ success,
AI-powered app that connects science to students’         multiple assessments can be designed to collect
lived experiences.49 Contributors also discussed a        and synthesize accumulated evidence of learning,
five-step framework for integrating AI into formative     development, and capability over time.54
assessment that builds on earlier work related to
digital technologies.50
4. Advancing a Focus on Responsible Assessment

T
      he work ahead includes reflecting on what
      assessment should make possible for learners.
      Despite advances in AI, the foundations of
assessment still rest on why it exists and who it is
                                                                       Responsible assessment should
meant to serve. This section invites readers to explore                be co-designed with the people
assessment design and implementation in the AI era.
                                                                       it affects and interpreted with
4.1 Reflecting on Responsible                                          careful attention to context.
Assessment
In keeping with the theme of the convening, panelists
were asked to share their thoughts on the concept of              students know what they know as a source of
responsible assessment. Their responses suggested                 empowerment. There is potential for AI-powered
that responsibility begins with the people assessment             personalized assessments that align tasks with an
is meant to serve. Several considerations emerged                 individual’s cultural, linguistic, and social context as
from their discussions.                          