Research Article
Rubric-Guided Generative AI for Scalable Creativity Assessment in Educational Games
Synthesis: Scoring Creativity at scale is difficult for a structural reason: the construct is defined as producing something both novel and appropriate, and the established instruments either decontextualise it into divergent-thinking tests or price themselves out of routine use, since the Consensual Assessment Technique needs several expert raters per artifact. Rahimi, Li, Esmaeiligoujar and Ercan ask whether a rubric-guided large language model can close that gap inside a real game. GPT-4o rated 421 student-designed levels from Physics Playground against the same seven-dimension rubric two educational experts had used, under three prompt conditions that varied only the input modality. Given both the level image and its structured JSON data, the model reached r = .74 against the averaged human scores, compared with only r = .61 when the rubric was withheld; a single call was about as consistent as averaging three independent runs (ρ = .74 versus .71, RMSE 1.9 versus 2.1). The takeaway is that the leverage sits in rubric design and multimodal input rather than model choice: a validated rubric turns hundreds of scored game levels into a scalable, unobtrusive measure of student creativity, while dropping the image specifically destroys the model's grip on aesthetics and humor.
Key Findings
- A rubric-guided GPT-4o tracked expert judgments on 421 authentic artifacts. In Physics Playground, a 2D game where players build physics machines to guide a ball to a balloon, college students used the built-in level editor to design levels and then selected their most creative designs; each was saved as both an image and a structured JSON file. GPT-4o scored every level against the same seven-item rubric as the human raters, reaching a total-score correlation of r = .74 (95% CI [.69, .79]).
- The rubric, not the model, did most of the work. Prompting with the validated creativity rubric produced strong agreement (r = .74); scores generated without it showed only moderate agreement (r = .61). Rubric text and output format were held constant across conditions, so the comparison isolates the rubric's contribution to validity.
- Multimodal input was the strongest condition. Supplying both the level image and its structured JSON data produced the closest correspondence with human ratings. Removing the image reduced the model's ability to detect nuanced creative details such as aesthetics and humor, indicating that creativity in visual-spatial artifacts cannot be captured through structured data alone.
- Agreement was very uneven across rubric dimensions. Object elaboration was strongest (.73 [.68, .78]), followed by line elaboration and aesthetics (.64 each) and humor and surprise (.57). The weakest dimensions were title creativity (.35 [.26, .44]) and solvability, where correlations were low and sometimes non-significant (.24, .35, and −.01 in the three conditions; the last marked p > .05).
- A single call was about as reliable as averaging three. Comparing one model run with the average of three independent runs gave ρ = .74 versus .71 with RMSE 1.9 versus 2.1 on the rubric total, so repeated sampling bought little over a single automated pass.
- Larger architectures scored better. Comparing related models on the same task gave GPT-4o > GPT-4o-mini > GPT-4.1-nano, with overall correlations of approximately .71, .62 and .41 respectively — the smallest model tracked human judgment only weakly.
- Human ground truth was an averaged expert judgement. Two educational experts independently rated each level on a 0–2 scale per dimension, for a maximum total of 13; their per-dimension scores were averaged to form the reference standard against which the model was compared.
The rubric and how it makes creativity scoreable
The instrument is a seven-item rubric adapted from Rahimi and Shute (2021), covering originality, elaboration of lines and objects, aesthetics, humor and surprise, and title creativity, each rated 0–2 for a maximum total of 13. The dimensions map onto the definition of creativity the paper adopts in its introduction — ideas or artifacts that are both novel and appropriate, that is, original, functional and valuable — with solvability standing in for appropriateness (a level must be playable), elaboration for the investment of effort that distinguishes a designed artifact from a doodle, and aesthetics, humor and title creativity for the expressive layer that a science teacher might otherwise overlook.
That operationalization is what makes the task LLM-legible. Because the rubric fixes the construct, the dimensions and the scale, the model is not asked to invent a standard of creative quality; it is asked to apply an existing one and emit a structured JSON object with dimension-level and total scores. The paper positions this as the most cost-effective form of AI-supported Assessment for educational contexts, against fine-tuning or bespoke model training, and it is also what allows a machine judgment to be argued with: a score can be traced to a named dimension and the rubric text that produced it.
How the study was run
The dataset is 421 levels designed by college students in a previous study (Rahimi, 2020) after a brief tutorial, with up to two hours to build playable levels and then a self-selection step where participants nominated their most creative designs. Each level was exported as an image plus a JSON file carrying full structural data, which is what made modality comparisons possible. Level design in an open-ended environment is the point: the artifacts are authentic student productions rather than responses to a standardized test item, which addresses the ecological-validity complaint against instruments such as the Torrance Tests of Creative Thinking.
Validity was assessed with Spearman's rank-order correlation (ρ) between model and human scores, with confidence intervals obtained by bootstrapping (1,000 resamples), and all correlations reported as significant at p < .001 after FDR correction except where marked otherwise. Reliability was probed by comparing a single run with the average of three independent runs, and the three prompting conditions differed only in input modality while rubric text and output format stayed fixed — a design that isolates the information available to the model from the instructions it was given. This is a quantitative instrument-validation study rather than an experiment with a control group; no student learning outcome was measured, and the criterion is human rating agreement rather than an external standard of creative quality.
Where the model agreed and where it diverged
The headline is that a general-purpose model, given a rubric and the right inputs, can approximate expert ranking of creative quality across hundreds of human-generated levels. The dimension-level picture is more instructive than the total. Structural, visually inspectable properties — object elaboration, line elaboration, aesthetics — were recovered well, and the paper reads the multimodal result as evidence that visual context is indispensable: the model's performance declined when images were excluded, and specifically lost the nuance in aesthetics and humor, dimensions that have little representation in a JSON description of a level.
Where the model diverged is equally clear. Agreement was weakest on title creativity (.35 across all three conditions, remarkably stable at that low level) and on solvability, which produced correlations of .24, .35 and −.01, the last not statistically significant. Both dimensions arguably sit outside what an image-plus-structure representation carries: solvability depends on whether a physics machine actually works when run, not on how it looks, and title creativity is a verbal and cultural judgment. The paper does not report whether the model was systematically generous or harsh relative to humans — no level-bias analysis is presented — so the evidence supports rank agreement, not calibrated absolute scores; the residual spread, an RMSE of 1.9 to 2.1 points on a 13-point scale, is the nearest available indication of how far individual scores could sit from the expert value.
Implications: the rubric is the lever, and correlation is not agreement
For assessment designers the transferable finding is that rubric-guided prompting, not model selection, produced the validity gain, and that the model's ceiling on any dimension is set by whether that dimension is visible in the supplied evidence. A Multimodal AI pipeline is therefore a design requirement for visual-spatial artifacts, not an optimization. The reliability result — one call performing close to the average of three — matters for cost: repeated sampling is largely unnecessary, which keeps automated scoring cheap enough to run over an entire cohort's worth of student work.
The paper argues such raters can function as scalable, unobtrusive assessors that sit inside an authentic game activity rather than interrupting it, and connects to the broader program of evaluating creative products with language models where human raters would be prohibitively expensive. The caution is that a strong correlation is evidence for rank ordering, not for replacing the expert. Whether a model that ranks levels correctly should also be trusted to assign a grade, give Feedback, or feed a human-reviewed report is a separate question that agreement statistics do not settle, and the low correlations on solvability and title creativity show that the model's coverage of a rubric is always partial.
What this means for practice
- Assessment designers. Give the rater the validated seven-dimension rubric instead of a bare instruction: agreement with averaged expert scores was r = .74 with the rubric and only r = .61 without it, with rubric text and output format held constant.
- Assessment designers. Score visual-spatial artifacts with the image plus the structured level data, not structure alone, because dropping the image destroyed the model's grip on aesthetics and humor specifically.
- Assessment designers. Run a single pass rather than three: one call matched the average of three independent runs (ρ = .74 vs .71; RMSE 1.9 vs 2.1), which keeps scoring a whole cohort affordable.
- Educators. Reserve expert judgment for the dimensions the supplied evidence cannot carry — solvability (r = .24, .35 and −.01 across conditions) and title creativity (.35) — and treat model scores as a ranking and screening aid.
- Learning designers. Keep a human decision on grades and Feedback: this study established rank agreement only, with no score-level bias, calibration or absolute-error analysis.
Limitations
The study is bounded by its context: one 2D physics game, one level-editor task, and a dataset of 421 levels drawn from an earlier study of college students, so transferability to other domains, artifact types and age groups is untested, and the paper's own framing is physics and science education's game contexts rather than a general claim. The results are tied to GPT-4o and its specific version: prompting behavior and multimodal handling change across releases, and the model-comparison result shows that performance is not stable across a family, so the .74 figure should be read as a snapshot rather than a property of LLM-based scoring. Correlation is not agreement, and no score-level bias, calibration or absolute-error analysis is reported, so a model could rank levels in human order while displacing every score. The human criterion is two raters averaged, so their own variance is folded into the ground truth, and small differences in rater interpretation of dimensions such as humor or title creativity would cap what any model can achieve. Finally, dimension-level correlations are reported without per-dimension reliability for the human raters, and no qualitative error analysis examines the levels where model and experts diverged most. These are ordinary limitations of AIED research on measurement: the evidence supports rubric-guided LLMs as a screening and ranking aid at scale, and leaves their standing as a substitute for expert judgment open.
Connected Concepts
- Creativity — the construct under measurement, operationalized as novelty plus appropriateness in a designed artifact
- Automated Assessment — rubric-guided LLM scoring of student products as a scalable alternative to expert panels
- Assessment Validity — Spearman correlations with bootstrapped confidence intervals as the validity criterion
- Educational Measurement — averaged expert ratings as ground truth, run-to-run reliability and RMSE
- Game-Based Learning — Physics Playground's level editor as the authentic, unobtrusive assessment context
- Multimodal AI — image plus structured JSON as the condition that best matched human ratings
- Large Language Models (LLMs) — GPT-4o as the rater, with smaller models compared alongside it
- Generative AI — the capability class being tested as an assessor rather than a generator
- AI Feedback Quality — the accuracy question behind machine judgment of creative work
- Prompt Engineering — rubric text and structured JSON output as the highest-leverage intervention
- Psychometrically Aware AI — correlation, reliability across runs and the correlation-is-not-agreement caveat
- Limitations in AIEd Research — single game, model-version dependence and unresolved rater variance
Connected Articles
- Using LLMs to Detect Growth in Computational Thinking in Introductory Physics — LLMs mirroring human coders on physics problem-solving constructs
- Comparing GPT and human raters in essay assessment: Variability, bias, and the potential of LLM-based scoring — Variability and bias when GPT and human raters score the same essays
- Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring — Cutting the cost of LLM scoring through prompt selection
- AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models — Open-source LLM writing evaluation with adapted instruction-tuned models
- Can AI Evaluate Assessment? A Study of Large Language Model Meta-Assessment Performance — LLMs turned on assessment artifacts themselves
- Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education — Improving the reliability and validity of the human criterion
- Towards Scalable Measurement of Durable Skills — Rubric-based autorating of creativity in AI-mediated group tasks
- Cultivating Design Creativity of Vocational Students: A Model of Project-Based Learning in AI-Enabled Immersive Virtual Environments — Design creativity as an outcome in AI-enabled project-based environments
- ProductiveMath: A Generative-AI-Powered App to Support Productive Failure Teaching — The same group's generative-AI app for productive-failure teaching
Citation
Rahimi, S., Li, H., Esmaeiligoujar, S., & Ercan, D. (2026). Rubric-Guided Generative AI for Scalable Creativity Assessment in Educational Games.