Research Article
AI Personas: Can LLMs Replace Fieldwork in Experiential Learning?
Synthesis: Elhajj et al. (2026) report an emergency pedagogical substitution: when the 2024 war in Lebanon made community fieldwork impossible for the HEHI 303 Experiential Learning course at the American University of Beirut, student groups interviewed ChatGPT-generated AI personas instead of real stakeholders in needs assessment interviews and focus group discussions. Two HEI raters scored all 10 group prompts on five indicators, and AI dialogue came out strongest where schooling normally wins anyway. Educational alignment (mean 5.00) and diversity of perspectives (4.90) topped the rubric, while group dynamics (3.80) and limitations (3.20) were weakest, where flat emotion, over-polished phrasing, absent disagreement and thin cultural specificity ran through every context. The authors conclude that AI personas are a useful complement and a crisis stopgap, not a replacement for real fieldwork.
Key Findings
- The study analyses all 10 AI persona prompts developed by 10 student groups out of the 63 students enrolled in the graduate HEHI 303 Experiential Learning course, each group working a distinct humanitarian case, with all personas built using ChatGPT 4.0.
- Prompts were rated on a five-indicator rubric: Authenticity and Realism, Diversity of Perspectives, Alignment with Educational Goals, Group Dynamics & Coherence, and Limitations and Gaps, on a 1 to 5 scale, each scored independently by two HEI research assistants and averaged.
- Educational Alignment was the strongest indicator, reaching a perfect mean of 5.00 across all ten prompts: regardless of context, personas produced usable material for stakeholder mapping, needs assessment, thematic coding and systems thinking.
- Diversity of Perspectives scored a mean of 4.90, with nine of ten prompts generating broad demographic, occupational and geographic variation. The exception was Prompt 2, whose women's focus group produced near-identical answers, a pattern the authors warn risks reinforcing stereotypes about displaced women.
- Authenticity and Realism averaged 4.38 and was highest where prompts were most locally specific (WASH professionals in Lebanon, Maasai women in Kenya, climate professionals in Ethiopia) and lower for broad or complex settings where responses drifted toward polished generality.
- Persona register mattered: authority figures such as a Ministry of Education official returned bullet-point policy prose that "reads as a policy document rather than spoken dialogue", while youth, elderly and other vulnerable personas produced the most conversational, emotionally expressive exchanges.
- Group Dynamics & Coherence was the most variable indicator (mean 3.80, range 2.5 to 4). Its lowest score came from Prompt 4, where 30 simultaneous Syrian women personas exceeded ChatGPT's capacity for group interaction and the output became a set of sequential interviews rather than a discussion.
- Limitations and Gaps recorded the lowest scores everywhere (mean 3.20), a cross-cutting weakness the authors read as a structural property of current LLMs rather than a prompt-design failure: emotional flatness, repetition, no contradiction or hesitation, and little cultural specificity in every one of the ten contexts.
- Inter-rater reliability was high overall: 84% exact agreement with a mean absolute difference of 0.15 points. Educational Alignment and Diversity of Perspectives reached 100% exact agreement, Authenticity and Realism 80% exact and 90% within ±0.5, Group Dynamics & Coherence 80%, and Limitations and Gaps only 60%, a gap the authors say means that indicator should be used alongside instructor reflection and student debriefing, not as a standalone metric.
- Student groups were trained on a seven-step persona routine: specify the task, define role and context, request detailed background, state the role-play, state the aim of the interview, provide an interview guide, and ask for a conversational tone.
- Specificity drove quality. More contextually detailed prompts produced richer, more differentiated responses, while vague prompts produced generic, culturally thin output, consistent with prior findings on prompt sensitivity in simulated interview training.
- The cross-case conclusion is blunt: AI personas perform best against educational and diversity objectives and worst at reproducing the dynamic, unpredictable, emotionally layered nature of real fieldwork.
Study Design & Method
This is a qualitative case study of a single bounded context: the HEHI 303 Experiential Learning course of the Humanitarian Engineering Initiative at AUB during the 2024 conflict, when students could not reach communities for needs assessments. The design is descriptive rather than comparative. Each group wrote one prompt to generate personas for simulated stakeholder interviews and focus group discussions, and the unit of analysis is the resulting prompt-and-dialogue set, not the students. The paper frames the work as analyzing student reflections alongside AI-generated dialogue, but the reported Results are the prompt-by-prompt and aggregated ratings; no separate student-perception survey or reflection scores are presented. Data collection is dated to December 2024, anonymized before analysis and granted a retrospective IRB exemption because the activity was originally course assessment, not research.
What this means for practice
- Instructors. Pair every AI persona session with a real interview or community contact so students can compare emotional texture and disagreement against what the personas omit, and debrief that gap explicitly: Limitations and Gaps was the weakest indicator across all ten contexts (mean 3.20). Train students to interrogate AI output for inconsistency, bias and cultural generalization rather than accept it as evidence, and keep instructor oversight on that output, because over-generalized claims (for instance, assuming all students had online access) can pass into student analysis unchecked.
- Instructors. Require the seven-step persona routine — specify the task, define role and context, request detailed background, state the role-play, state the interview aim, provide an interview guide, ask for a conversational tone — and insist on locally specific detail, since authenticity was highest where prompts named concrete populations and settings (mean 4.38).
- Curriculum designers. Reserve AI personas for situations where community access is unsafe or impossible, and for low-stakes rehearsal of interview guides and question sequencing; do not let them displace the interpersonal objectives of an experiential learning course.
- Instructors. Cap how many personas a student group tries to run inside one dialogue: 30 simultaneous Syrian women personas exceeded the model's capacity and the discussion collapsed into sequential interviews, landing Group Dynamics & Coherence at a mean of 3.80.
- Researchers. Never treat the Limitations and Gaps rating as a standalone quality metric: the two raters agreed exactly only 60% of the time on that indicator, against 100% exact agreement for Educational Alignment and Diversity of Perspectives.
Limitations
- The study is a single descriptive case inside one course at one institution during one conflict, and it reports no learning-outcome data, so it cannot show that persona practice produced better competencies than fieldwork would have.
- The evaluation rubric was applied only to AI-generated dialogue; traditional FGDs and interviews were used informally during prompt development rather than scored as a matched comparison dataset.
- The two raters were HEI research assistants actively involved in the course, and their agreement was weakest on Limitations and Gaps at 60% exact agreement; the rubric also measures response quality rather than student learning, with emotion, contradiction and cultural texture consistently underplayed regardless of geographic setting.
- For wider adoption the authors list technical accuracy, training-data bias, privacy exposure through personalization, resource intensity, and the risk that over-dependence on AI erodes the interpersonal skills that social and emotional learning depends on.
Connected Concepts
- Experiential Learning — the course model AI personas were tested against
- Simulation — AI role-play standing in for community interviews and FGDs
- Conversational AI — ChatGPT 4.0 as the persona engine
- Prompt Engineering — the seven-step persona prompt routine taught to students
- Pedagogical Agent — personas cast as stakeholders rather than tutors
- Qualitative Research — needs assessment, thematic coding and focus group method being taught
- Human-in-the-Loop — instructor guidance required to contextualize inaccurate output
- Human AI Collaboration — the complement-not-replace framing the paper ends on
- Culturally Relevant Pedagogy — the cultural specificity gap in AI dialogue
- Social-Emotional Learning — emotional nuance as the consistently missing element
- Global South — fragile, low-resource and conflict-affected educational settings
- Limitations in AIEd Research — single-case, rater-based evaluation and its reliability ceiling
Connected Articles
- Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation — Stress-testing role-playing LLM agents
- Designing Conversational Agents for Adaptive Instructional Support in Business Simulation Gaming — Conversational agents in business simulation
- Using Generative AI to Simulate Patient History-Taking in a Problem-Based Learning Tutorial: A Mixed-Methods Study — Simulated patient histories for problem-based learning
- AI-Powered Simulation for Nursing Education: Mixed Methods Systematic Review — AI-powered simulation in nursing education
- How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding — LLMs for qualitative coding and consensus
- Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis — Human versus LLM coding of qualitative data
- Invisible Impact of Empathy on Behavioral Change: Isolating the Effect of Empathy in Long-term Physical Activity Coaching Chatbot Interactions — Chatbot coaching of empathy
- Culturally-Aware AI for Cross-Boundary Community Learning: Undergraduate Innovation at the Intersection of Computation — Culturally aware AI in community learning
- An Experimental Study Exploring Human–AI Complementarity in Early Social-Emotional Learning — Human-AI complementarity in social-emotional learning
- Prompting for Teachability: Designing Novice Personas in LLMs for Learning by Teaching Contexts — Prompting teachability in novice personas
Citation
Elhajj, I., Germani, A., Farah, J. M., Sabbagh, D., & Hajj Hassan, J. (2026). AI Personas: Can LLMs Replace Fieldwork in Experiential Learning?. AI in Education, 2(3), 30.