On this page

Synthesis: Alam SM et al. (2026) present a multimodal collection of scanned examination answers paired with expert-designed Outcome-Based Education (OBE) grading metadata, intended to support criterion-level assessment research. The dataset contains 485 answer submissions from 415 consenting students at four academic institutions, with eight faculty contributors supplying examination materials across nine subjects and 12 question templates. Each answer-level item links a scanned PDF to a randomized identifier, subject label, question, model answer, criterion definitions, performance-level descriptions, criterion marks, and a total mark; the 12 rubrics contain 47 criteria in total. The scans deliberately preserve realistic visual variability — handwriting, crossed-out work, equations, code, figures, and varied capture pipelines — to enable robustness and generalization studies in vision-language document understanding and grading.

Key Findings

  1. The dataset links visual examination evidence with question context, model answers, expert-authored OBE criteria, performance descriptions, criterion marks, and total marks at the answer level — enabling criterion-level (not just holistic) assessment.
  2. Nine-subject coverage with mixed visual content (handwriting, printed text, equations, tables, code, figures, sketches, diagrams) supports research on handwritten and printed document understanding, OCR, vision-language modeling, and cross-subject generalization.
  3. Deliberate variation in illumination, contrast, orientation, compression, resolution, and handwriting (including crossed-out work and inserted corrections) provides the robustness testing ground typical of real assessment pipelines.
  4. An answer-level audit confirmed 485 unique identifiers, 485 unique PDF filenames, and agreement between criterion counts and rubric definitions — ensuring data integrity for Benchmark use.

Discussion

This is a data-article contribution to the emerging field of automated grading and criterion-level assessment. By pairing raw visual evidence with expert-designed OBE rubrics, it provides the kind of ground-truth-labeled multimodal corpus needed to benchmark AI-based evaluation fairly — addressing a recurring weakness in the literature where grading models are trained and evaluated on narrowly homogeneous data. The explicit OBE structure ties the technical corpus to contemporary higher education assessment practice, and the emphasis on realistic visual degradation supports claims about robustness that lab-clean datasets cannot. For the knowledge base, it connects the Assessment and Educational Measurement agendas with the multimodal/vision-language frontier of automated evaluation.

What this means for practice

  • Designers. Train and evaluate at the question level or with grouped controls: a random answer-level split can place visually related responses to the same prompt in every partition and overstate cross-question generalization.
  • Researchers. Normalize criterion marks or build question-aware models, and reflect the imbalance in sampling plans — question rubrics use different maxima and criterion counts, and subjects range from 5 answers (E-commerce) to 88 (Machine Learning).
  • Designers. Keep a human in the loop: outputs from models trained on this corpus are research results and must never be presented as an official grade or the sole basis for a decision affecting a student.
  • Researchers. Exploit the deliberate visual variability — illumination, contrast, skew, compression, handwriting style, correction density — to test whether a model stays stable outside a single capture pipeline.

Limitations

  • The collection covers nine subjects but only 12 question templates, and its subject distribution is uneven (five answers in E-commerce versus 88 in Machine Learning), which caps cross-question generalization claims.
  • Question-specific rubrics use different maxima, criterion counts, and textual scales, so comparisons across questions require normalization or question-aware modeling.
  • Access is restricted because some scanned pages retain student names, IDs, university names, and logos, limiting redistribution and requiring re-identification protection.
  • The 485 answer submissions come from only 415 students at the four contributing institutions, so conclusions should not be extended beyond them.

Connected Concepts

Connected Articles

Citation

Alam, J. S., Syfullah, M. K., Ahmed, S., Mou, M. A., Rahman, A. K. Z. R., Rahman, A. K. M. M., & Ali, M. S. (2026). Multimodal Examination Answer Data with Expert-Designed Outcome-Based Education Rubrics for Criterion-Level Assessment.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.