On this page

Synthesis: A pedagogical model where human mentors and AI tools jointly support student learning in project-based contexts. Human mentors provide conceptual guidance, debugging, and problem formulation support; AI tools accelerate execution, code generation, and rapid iteration. Demonstrated by Chawla et al. (2026) in a ten-week financial forecasting project with high-school and early-undergraduate students, the model produced accelerated progress, instructive failure modes, and a clear division of labor: AI handled tactical execution while human mentors retained strategic judgment.

Definition

A pedagogical model where human mentors and AI tools jointly support student learning in project-based contexts. Human mentors provide conceptual guidance, debugging, and problem formulation support; AI tools accelerate execution, code generation, and rapid iteration. Demonstrated by Chawla et al. (2026) in a financial forecasting project with high-school students, the model positions AI as a complement to—rather than a replacement for—human mentorship, with each side covering the other's limits.

Key Findings

  1. AI tools functioned as effective tactical co-mentors—generating code scaffolding, suggesting hyperparameter ranges, and offering just-in-time explanations—while human mentors retained strategic guidance, problem framing, and validation of architectural decisions.
  2. A workflow-driven approach (students identifying the sequence of steps needed and executing each with AI) let novices with limited AI and finance backgrounds build meaningful predictive models without prerequisite classroom instruction.
  3. Three recurring failure modes emerged: plausible-but-incorrect code, missing specialized or niche knowledge, and degraded awareness of long-term project objectives across context windows.
  4. LLM-generated sentiment scores did not improve ETF forecast accuracy—likely already priced into movements or too noisy—a negative result mentors reframed as a valuable lesson in hypothesis testing.
  5. Students followed a clear arc of critical evaluation: from accepting LLM outputs with little scrutiny to independently hypothesizing data leakage, suggesting that repeated exposure to imperfect AI under supervision accelerates evaluation skills.

Workflow-Driven Learning over Traditional Instruction

Instead of "learn theory first, apply later," students identified the sequence of steps needed to solve their problem and used AI tools to execute each step. This just-in-time workflow-driven approach enabled students with limited backgrounds in AI and finance to build meaningful predictive models. Ten weeks were structured as a series of daily stand-ups (around 30 minutes each) focused on problems to be solved and tasks to be accomplished, with students pursuing specific learning exercises—including dialogues with the AI tools—between sessions. AI tools served as scalable accelerators that kept students engaged with higher-level concepts rather than mired in implementation details.

The Human-AI Complementarity

Within the study, LLM-based tools excelled at tactical support: generating code scaffolding, suggesting hyperparameter ranges, and offering quick explanations of unfamiliar concepts, allowing students to engage immediately rather than waiting for prerequisite instruction. However, AI could not replace the strategic guidance of human mentors, who framed problems, validated architectural decisions, and helped students recognize when outputs were incorrect. The most effective configuration was complementary: AI handled routine implementation tasks, freeing human mentors to focus on higher-level reasoning, debugging, and conceptual explanation. This division reflects a broader human-in-the-loop model in which routine work is offloaded to the machine while judgment and oversight stay human.

Failure Modes and Productive Failure

The study identified three recurring categories of AI failure. First, plausible-but-incorrect code: AI-generated code often ran without errors yet contained subtle methodological flaws in modeling assumptions or feature construction. Second, a lack of specialized knowledge: models produced generic templates but missed niche libraries such as curl-impersonate that proved critical for scraping. Third, limited long-term context: the AI system lost awareness of overall research objectives over time, risking incorrect advice. Notably, each failure mode became a teaching moment—an instance of instructive failure that deepened hands-on engagement with the underlying concepts and, per the mentors, increased the depth of the scholar–student relationship.

The Sentiment Experiment: A Negative Result

Students built a pipeline that scraped news for 29 ETFs and used gpt-5-mini to assign sentiment scores on a −10 to +10 scale, which then fed forecasting models spanning classical statistics (ARIMA, SVR, XGBoost) and deep learning (LSTM). Contrary to expectations, models that included sentiment underperformed those without it, likely because sentiment was already priced into price movements by the time it appeared in news, or because the explicit scores introduced noise. Though initially disappointing to students who had invested heavily in the sentiment pipeline, mentors reframed the outcome as a lesson in hypothesis testing: the purpose of the experiment was to discover what the data revealed, shifting students' understanding of research from "proving ideas right" to "finding out what's true."

Connections to Agentic Workflows

This model bridges Evolution of AI in Education: Agentic Workflows and practical classroom implementation. Where agentic workflows describe AI agent paradigms (reflection, planning, tool use, multi-agent collaboration), co-mentorship shows how those paradigms manifest in human-AI educational partnerships. The daily stand-up structure echoes the reflection loop in Agentic Education with AI Coding Assistants. The complementary division of labor—AI tactical execution paired with human strategic judgment—resonates with how agentic AI is imagined in educational contexts, but grounds it in observed classroom practice.

What this means for practice

  • Instructors. Keep strategic work — problem framing, validation of modeling choices, interpretation of results — with humans, and delegate tactical execution (code scaffolding, hyperparameter suggestions, just-in-time explanations) to AI; that division of labor is what the ten-week project actually used.
  • Curriculum designers. Replace theory-first sequencing with a workflow-driven arc in which students define the steps their problem requires and execute each with AI support, which let students with limited AI and finance backgrounds reach a working predictive model.
  • Instructors. Treat imperfect AI output as curriculum rather than a hazard: each of the project's three recurring failure modes (plausible-but-incorrect code, missing niche knowledge, and loss of long-term project objectives) became a teaching moment that pushed students toward independent verification.
  • Curriculum designers. Keep an explicit hypothesis-testing lesson in the syllabus: the pipeline that scored 29 ETFs with gpt-5-mini sentiment underperformed models without sentiment, and mentors used that negative result to shift students from "proving ideas right" to "finding out what's true."

Limitations

  • This is a single case study of one ten-week financial forecasting project with high-school and early-undergraduate students; no control group or comparison cohort is reported.
  • All evidence comes from one project in one domain (ETF forecasting with news sentiment), so the reported division of labor between human mentors and AI tools is not shown to transfer to other subjects or project types.
  • The three failure modes and the students' arc toward skeptical evaluation rest on mentor observation during the project rather than on a measured instrument of verification skill.
  • The authors state that broader validation across more diverse cohorts and longitudinal outcomes is needed before the model can be generalized.

Citation

Chawla, F., Chawla, A., Singh, R., Germino, J., & Khvatskii, G. (2026). Human-AI Co-Mentorship in Project-Based Learning: A Case Study in Financial Forecasting.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.