Research Article
Towards Scalable Measurement of Durable Skills
Synthesis: Collaborative problem solving, Creativity and Critical Thinking are the skills employers most want and educational systems least measure, because valid assessment demands either naturalistic group interaction or psychometric control, and the two pull apart. This paper's proposal is an "Executive LLM" that plays every AI teammate in a group task with Gemini, holds the scoring rubric, and deliberately steers the conversation toward situations that force the participant to demonstrate the skill being assessed. Across 373 conversations from 188 participants, skill-matched Executive steering raised the share of conversations yielding ratable evidence to 92.4% for project management and 85% for conflict resolution, significantly above unconstrained independent agent teams, while LLM scoring agreed with expert annotators about as well as the two human experts agreed with each other (Kappa 0.45–0.64). The design reframes assessment as an adaptive test of complex behavior: a simulated partner engineered for evidence density rather than a standardized task assumed to elicit it.
Key Findings
- The Vantage protocol put a human participant in a 30-minute chat-based group task with three AI teammates, with Gemini 2.5 Pro generating all AI turns; 188 participants aged 18–25 recruited through Prolific produced 373 usable conversations (three discarded for technical faults).
- Inter-expert agreement between two trained pedagogical raters was only moderate, with Cohen's Kappa of 0.45–0.64 for conflict resolution and project management, covering both the "score or NA" decision and numerical score agreement; LLM–expert agreement fell in the same band.
- Conversation-level evidence — the share of conversations rated as carrying enough information for a skill rating rather than NA — reached 92.4% for project management and 85% for conflict resolution when the skill-matched Executive LLM was used.
- A skill-focused Executive LLM always elicited significantly more evidence for its target skill than unconstrained Independent Agents (Fisher exact test, p ≤ 0.05), and the steering showed a crossover effect: focusing on conflict resolution raised conflict-resolution evidence but lowered project-management evidence, and vice versa.
- Telling participants to pay attention to a skill before the conversation had no significant effect on informativeness in either the Independent Agents or the Executive LLM setting, and for both skills (all p > 0.6) — the gain comes from the system's steering, not from priming the person.
- Task topic did not drive informativeness: a logistic regression found no significant effect of a science versus debate task on either skill metric (p = 0.18 for conflict resolution, p = 0.9 for project management), supporting reuse of the framework across subject areas.
- Evidence rates were generally higher for project management than for conflict resolution in almost all conditions, which the authors read as project management behaviors being more abundant and needing less steering.
- Automatic scoring ran turn by turn with Gemini 3.0: each turn was rated 20 times, a turn was labeled NA if any single run returned NA, and the remaining labels were resolved by majority vote before a regression model produced the conversation-level score.
- In a Monte Carlo style recovery test, simulated participants were prompted to behave at a stated rubric level; each conversation ran 50 turns and each level was repeated 100 times, and the Executive LLM protocol reduced mean absolute error in recovering the known level relative to Independent Agents.
- For creativity, a Gemini-based autorater assessing complex tasks submitted by high-school students performed on a par with human expert raters, extending the claim beyond the adult Prolific sample.
Why Group Skills Resist Measurement
The paper situates itself between two established poles. PISA 2015's collaborative problem-solving assessment had subjects interact with scripted simulated teammates through multiple choice, maximizing control and losing authenticity; the ATC21S project ran human–human dyads acting on shared digital objects, gaining naturalism and surrendering standardization. Both sit far from the classroom interaction they aim to represent. Group work is simultaneously valued for individual learning and for developing durable skills, and the interdependence between members is exactly what breaks classical psychometric assumptions: a person's score depends on what their partners did.
The authors' claim is that large language models change the feasible point on that trade-off, because naturalistic conversational personas no longer require hard-coded rules. The risk is the mirror image: Sijtsma's observation that measurement is a compromise for efficiency, since waiting for a person to spontaneously exhibit a behavior in real life takes too long to gather evidence. The Executive LLM is the answer to that objection — an adaptive test of complex behavior, in which the system's goal is to manufacture the occasions on which evidence can appear while keeping the conversation plausible.
Rubrics, Steering and the Executive LLM
Rubric construction followed a documented sequence: review the literature to build a conceptual model of each skill, derive an initial rubric, have human experts score sample conversations with it, then refine the dimensions where agreement was low or raters found them ambiguous. Each dimension was scored 1–4 with an explicit NA option, and the same rubrics were given both to the steering model and to the evaluator — the mechanism by which the conversation is steered toward the rubric rather than away from it. The paper notes this aligns with earlier findings that LLMs code conversations reliably when the rubric is theory-derived and then refined through expert use on real data.
Operationally, the Executive LLM is a single model generating the responses of all teammates rather than several independent agents each playing a role. It has access to the rubrics and is prompted to maximize the information and accuracy of the assessment, which can mean having a teammate initiate a conflict and sustain it until evidence of conflict resolution has been observed. That single-model choice is not just an engineering simplification: the paper's comparison shows that a group of unconstrained LLMs produces insufficiently informative interactions, because teammates that happen to collaborate smoothly never surface the behaviors being measured. Feedback to the participant is quantitative, organized as a skills map with per-axis breakdowns and drill-down excerpts from the conversation that substantiate each rating.
Agreement, Evidence and the Simulation Sandbox
The evaluation's honesty lies in how it treats its own reliability ceiling. Human coding of conflict resolution and project management was difficult even after several calibration rounds, producing moderate inter-rater Kappa; the LLM evaluator's agreement with humans sits in the same range. The authors use that as a license to run the larger validity analyses with autoraters only, and they state plainly that human-rating results are qualitatively similar. That is a defensible move, and also a limitation: the protocol cannot currently claim more precision than two trained experts can achieve on the same transcripts.
Because human data collection is expensive, the framework includes a simulated subject — a Gemini model prompted to act like a student at a specific rubric level — used to test recovery, to probe evidence levels, and to develop the protocol before deployment. This is where the work connects to the wider Simulating Students literature: the simulator is not a convenience sample but an instrument for estimating whether the assessment can recover known proficiency. The evidence-rate results themselves are the cleanest demonstration of the core mechanism: matched steering produces high, rubrics-aligned evidence density, mismatched steering trades one skill's evidence for another's, and instructing the human to try harder does nothing measurable.
What It Contributes and What It Does Not Yet Show
For Assessment practice, three contributions stand out. First, a construct-validity argument for conversational AI teammates: authenticity and control are treated as jointly optimisable rather than as a fixed trade-off. Second, an operational lever — steering toward evidence — that measurably changes whether an assessment can score a person at all, which is the practical difference between an unratable and a ratable session. Third, an evaluation pipeline with a worked agreement analysis, which is what Educational Measurement generally demands and what most AI assessment proposals omit.
The boundaries are equally clear. Participants were US-based English-native adults aged 18–25 recruited on Prolific, so the classroom claim is indirect; the high-school creativity result is the exception and it is a separate analysis. The collaboration results rest on two sub-skills of one skill, the strongest evidence is at conversation level rather than the finer turn level, and the rubric dimensions scored 1–4 are coarse. The paper also acknowledges that the executive steering trades evidence breadth for depth — over-steering one skill reduces evidence for others, so a single conversation is a poor vehicle for a whole-profile assessment. What it establishes is that the amount of ratable evidence is a design variable under the system's control, which is a different and more tractable starting point than assuming a well-designed group task will elicit what it is meant to measure.
What this means for practice
- Assessment designers. Treat ratable evidence as a design variable: skill-matched Executive steering produced ratable evidence in 92.4% of project-management conversations and 85% of conflict-resolution conversations, significantly more than unconstrained independent agents.
- Assessment designers. Do not count on priming or task redesign to raise evidence — telling participants to attend to a skill had no significant effect (all p > 0.6), and a science versus debate task made no difference (p = 0.18 for conflict resolution, p = 0.9 for project management).
- Designers. Build one steering model that holds the rubric and manufactures occasions for evidence — for example a teammate who initiates a conflict and sustains it until resolution has been observed — instead of several unconstrained agents that collaborate smoothly and surface nothing to score.
- Edtech designers. Steer per skill and plan for multiple conversations per learner: matched collaboration steering raised evidence for its target skill but lowered the other's, so one conversation cannot carry a whole-profile judgment.
- Researchers. Report the reliability ceiling next to any automated score: two trained raters agreed at Cohen's Kappa of 0.45–0.64, and the LLM evaluator falls in the same band.
Limitations
- The human-rated conversations came from 188 Prolific participants aged 18–25 who were US-based English native speakers, so classroom claims are indirect; the high-school creativity result is a separate analysis.
- Human expert validation covered collaboration only, and within it two sub-skills (project management and conflict resolution); creativity and critical thinking were analyzed largely on simulated conversations rather than human participants.
- Rubric dimensions were coarse 1–4 scales with an NA option, and inter-expert agreement stayed at Kappa 0.45–0.64 after several calibration rounds, so the protocol cannot claim more precision than two trained experts achieve on the same transcripts.
- Evidence was measured at conversation level (each turn rated 20 times, with any single NA returning an NA label), and the steering trades breadth for depth, so a single conversation is a poor vehicle for a whole-profile assessment.
Connected Concepts
- Assessment
- Educational Measurement
- Collaborative Learning
- Critical Thinking
- Creativity
- Group Work
- Simulating Students
- Agentic AI
- Human AI Collaboration
- Learner Modeling and Adaptive Instruction
- Psychometrically Aware AI
- Authentic Assessment
- Generative AI
- Large Language Models (LLMs)
- Assessment Validity
Connected Articles
- Assessment in Team Problem-Solving Exercises in Computing Education — Assessing Team Problem Solving in Computing Education
- Causal Modelling of Support Interventions for Student Competency Assessment — Causal Modeling for Competency Assessment
- CLARA: An AI-Augmented Analytics Dashboard for Collaboration Literacy — CLARA: A Collaboration Literacy Dashboard
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents — Simulating Students at Diverse Cognitive Levels
- Simulating Students with Large Language Models: A Review of Architecture, Mechanisms, and Role Modelling in Education with Generative AI — Simulating Students with LLMs: A Review
- Modelling Individual Participants as LLM Agents in Collaborative Problem Solving Simulations — LLM Agents in Collaborative Problem-Solving Simulation
- AI as Teammate: Rethinking Task Distribution in Medical Training — AI Teammates and Task Distribution in Medical Training
- AgentSchool: An LLM-Powered Multi-Agent Simulation for Education — AgentSchool: Multi-Agent Simulation in Education
- AI Coaching for Accelerating Human Skill Development with Reinforcement Learning — AI Coaching and Skill Development
- Towards Valid Student Simulation with Large Language Models — Valid Student Simulation with LLMs
Citation
Globerson, A., Keeling, A., Choudhury, A., Iurchenko, A., Segal, A., Hassidim, A., et al. (2026). Towards Scalable Measurement of Durable Skills. arXiv preprint.