FAQ
How Should I Handle AI in Group and Collaborative Assignments?
You set a group assignment, and two failure modes arrive with the submissions: an artifact that is polished and coherent but that no member can explain, and a team where one student ran the work through generative AI while three others coasted on the shared grade. Nothing in the file tells you which.
Both are grading problems before they are integrity problems, and design problems before either: a group assignment makes two claims at once — what the team produced, and what each member learned — and AI presses on both, changing how easily work can be partitioned, how smoothly contributions fuse, and how hard it is to tell whose thinking is in the submission.
The bottom line: decide AI's place per task, not per course; make the group's negotiation about AI use something they submit and you grade; design the work so it cannot be split into parallel chatbot sessions; and keep at least one component each member answers for alone. None of that requires policing tools — it requires designing the task.
The short version
- Permit deliberately, and say what the permission covers. Name which parts AI may touch — idea generation, language polishing, layout, scripting — and which are meant to be worked out without it.
- Protect the segments where the group builds shared understanding — the framing, the disagreement, the reconciliation — because AI-mediated efficiency compresses exactly those.
- Require a short, jointly authored AI-use statement, and grade the reasoning. This turns a peer norm into reviewable judgment.
- Make the process visible: intermediate deliverables, shared planning documents, per-member reflections, peer assessment of contribution.
- Keep individual accountability beside the group mark — an oral defense, or an individual explanation of any section a member did not draft.
- Choose the access configuration on purpose. A shared interface produces negotiated, visible AI use; private prompting produces opaque flows and quietly de-labeled outputs.
- Read a group's restraint carefully — declining AI can be deliberate self-regulation or simple unfamiliarity, and the two call for different responses.
Decide what AI is for in this assignment
The corpus does not settle the permission question as a permission question, and it argues against treating it as one. Chen and Zou (2026) studied groups of three to five delivering a presentation worth 30% of the grade under an institutional "use only with explicit acknowledgment" policy, with GenAI explicitly permitted for idea generation, language polishing, visual layout and scripting. Inside that one permissive policy, students reached opposite conclusions with defensible reasoning on each side, and the design shaped which conclusion a group reached. A course-level rule does not determine what a group does; the task does.
So write the permission at the level of the artifact: say which phases are AI-open and which are not, and tell students why. The scoping review's design logic points the same way — if AI-mediated communication compresses negotiation and collective sensemaking, the segments where a group builds shared understanding are the ones to protect. Be honest about the warrant, though: neither review claims that withholding AI improves learning outcomes. That case rests on the structure of the trade-off, not a measured comparison.
Build the task so the group cannot just divide it
Wei and Perkins (2026) — a PRISMA-guided scoping review of 18 English-language studies published between January 2023 and March 2025, analyzed with reflexive thematic analysis and mapped across eight themes — gives both sides of the ledger. The benefits: group knowledge development, idea generation, support for reflective thinking, communication efficiency, task coordination, and feedback. The risks: reduced peer interaction and engagement under over-reliance, plus privacy, transparency and accuracy concerns. Their assessment is blunt — benefits are empirically supported, while risks around privacy, transparency, bias and accuracy are largely discussed conceptually — and much of the research is short-term or conceptual, with little longitudinal work on group dynamics or cognitive development.
The trade-off governs design. The review reports that group work became more efficient with reduced demand for communication, negotiation and collective sensemaking — Lin et al.'s finding, cited within the review — so the efficiency gain and the interaction loss can be the same mechanism. You cannot take the speed without the loss unless you build the negotiation back in as required work. The review concludes that group-based assessment should shift focus from product to process.
Chen and Zou show how hard that is to do by accident: their three patterns of Learner Agency ran in different directions at once. Cooperation-oriented agency (five groups) intensified use to hold the work together, often feeding peers' contributions into a chatbot to decode them and align their own part, with one group rebuilding its workflow as "discussion → externalisation to GenAI → collective review → re-discussion." Normative agency (seven groups) deliberately restrained use, judging the task to demand situated knowledge AI could not reach — "AI only knows that moment when you type" — or treating AI-generated sections as unfair to the groupmates who would carry them. Non-enacted agency (three groups) changed nothing, partitioning work into discrete subtasks on "individual platforms" while individual students used GenAI in ways they never contributed to the team.
The rubric already required integration and coherence under a multicultural theme, each member's reflective insight tied to their own classroom contribution, and demonstrated group collaboration — and three groups still divided and conquered. A coherence criterion does not produce collaboration on its own.
The access configuration is part of this. Xu et al. (2026) observed 18 students in six groups of three on a single shared ChatGPT-4 interface for roughly 45 minutes, and interviewed nine further students about asynchronous teamwork. With a shared, visible interface, teams co-constructed "collective prompts," ran a recurring surface–evaluate–embed cycle, treated the chat as shared external memory, and openly negotiated AI's role. In asynchronous work, private prompting and output "de-labeling" produced opaque, privatized information flows that raised the cost of maintaining a shared cognitive model. GenAI's role also shifted, from subordinate assistant to contested teammate, and the system is cognitively involved but contextually unaware of the team's shifting focus.
Make each member's contribution visible
Korchak, Costley and Fanguy (2026) add the detail that should worry an assessor. In a scientific writing course, from interviews with 10 postgraduate students aged 23 to 39 (mean 27.7), groups used GenAI through pre-planned strategies or through open, individually driven interactions coordinated in shared documents; a distinctive use was GenAI as an editorial integrator, merging parallel sections into coherent text. Their most pointed observation: students' accounts sometimes contradicted their own written reflections — one student denied having a group strategy while describing one in writing. Group strategies are not always equally visible to every member, so they are not reliably visible to you either.
Capable individual AI use that never becomes team practice is the same phenomenon from the other direction. Both studies imply one countermeasure: engineer visibility into the deliverables, with a decision log, shared documents you can open, and a short note from each member on what they contributed and which parts they still cannot explain.
Assess the group without losing the individual
Chen and Zou's central design argument is that the negotiation of acceptable GenAI use should itself become an explicit, assessable learning outcome: teams should collectively justify and document how GenAI will and will not be used, rather than leaving the norm to peer pressure or perceived risk. That is a gradeable artifact that exists only if the group actually talked, and it converts an unenforceable rule into reviewable student judgment.
Pair it with an individual component — an oral defense of their own understanding, or a written explanation of any section they did not draft. That is also the honest answer to the free-rider problem: it grades comprehension rather than authorship, and comprehension is what a group mark was always a proxy for.
Martínez-Peláez et al. (2025) screened 203 studies for 2023–2024 in the Web of Science Core Collection and included 22 under a PRISMA 2020 protocol registered in PROSPERO. They report LLMs acting as catalysts for collaboration — idea generation, organization, peer feedback, simulated rubric-based evaluations and expert reviews, lower participation barriers — and, for problem solving, help exploring alternative solutions, interdisciplinary perspectives and authentic scenarios. They also note a useful inversion: LLMs often produce incomplete or incorrect responses, which prompts students to question, verify and improve the information — the AI use to look for inside an individual checkpoint.
Their limitations travel with any citation: a single database, journal articles only, no focus on ethics or privacy, only the first two years of the technology, and no risk-of-bias instrument applied. Treat "catalysts" as a description of observed practices, not a measured outcome.
Two things not to do. Do not reach for detection to sort individual contributions: none of these six studies tests detectors in group settings, and the individual-level problems with detection are documented elsewhere in this knowledge base. And do not read a group's restraint as either endorsement or a problem without checking which it is — Chen and Zou caution that restraint can reflect unfamiliarity or risk avoidance rather than deliberate self-regulation.
Handle disagreement about AI inside the team
Disclosure is not only a question you ask at the end; it is a norm the group has to settle, and the settling is where the interesting failures live. Chen and Zou identified two rationales for intensified use. Performance: students read the criteria and peer benchmarks and used AI to protect their part of a shared grade — a rational response to a shared mark. Perceived safety through norms: a permissive collective climate lowered the felt risk of misuse, so groups used the tool more openly and, in the cooperation-oriented case, more dependently. If you want transparency, a stated classroom norm is more effective than a warning.
A group that brings you a disagreement about AI use has usually done the hard part. Give them a structure: require the AI-use statement to record what the group decided not to do as well as what it did, and grade the justification. This is where Chen and Zou's point about collective agency lands: individual capability does not become collective capability on its own, which is why the negotiation has to be required rather than hoped for.
The three objections you will actually hear
"Group work is already unfair — AI just makes it worse." AI changes the shape of the problem, not its existence. The fairness complaint is already inside the group: Chen and Zou's restrained students treated generating their own section through AI as unfair to the groupmates who would carry it, and the performance rationale shows students protecting their share of a shared grade. No study here shows that AI has made free-riding worse, or that any intervention reduces it. Individual accountability and visible process are defensible on the design argument, not as proven fixes.
"AI makes collaboration better, so I should encourage it." Partly, and the partly matters. Martínez-Peláez et al. describe genuine collaboration supports — idea generation, organization, peer feedback, simulated rubric evaluations, lower participation barriers — but that is a review of observed practices with the limitations above, and Wei and Perkins found the benefits empirically supported while the risk side is largely conceptual. The randomized evidence they cite is the sharpest caution: AI produced more innovative suggestions without significantly boosting participants' overall innovativeness, and could homogenize ideas and invite cognitive offloading. Simulation work does not close the gap. Fang (2026) fine-tuned LLM agents on real participant dialogue (LoRA/QLoRA adapters on LLaMA 3.2–3B) to model 48 participants across 3,824 turns and six thematic codes, then compared real and simulated discourse with Epistemic Network Analysis; the simulated network reached a distance of 0.17 from the empirical network, below the 0.30 threshold, with a permutation p-value of 0.65, statistically indistinguishable, though the simulation slightly overemphasized Technical Constraints–Design links and under-represented Data and Performance Parameters. This is a framework paper with medium confidence: it supports one claim only — participant-specific agents can reproduce realistic collaborative discourse — and does not show that students learn more from simulated teammates.
"I cannot grade process." You are already grading a proxy for it and calling it product quality. The move is not to grade more; it is to grade the artifact that exists only if the process happened — the negotiated AI-use statement, the decision log, the individual defense. All three are bounded and checkable. What you should not claim is that any of this is validated: the agency patterns come from one qualitative study, 52 pre-service teachers in one course, and the authors say to read them as possible responses to design, not a distribution you can expect in your class. The writing study's 10 participants were AI-expert postgraduates, and its authors call the findings context-specific rather than a taxonomy.
What the evidence does not support
GenAI does not improve unaided collaboration: the scoping review's benefits concern the process and the product, it explicitly calls for longitudinal work on group dynamics and cognitive development, and gains measured during AI-supported group work have not been shown to transfer to collaboration without the tool. Detection does not sort group contributions. And no study here compares a group task with AI against the same task without it, so "AI-free group work produces more learning" is a design argument, not a finding you can cite.
Do this week
In ten minutes, before the next assignment goes out: require a jointly authored half-page stating how this team will and will not use GenAI, and why, and add an individual component to the rubric.
In one class period: run the task by Xu et al.'s test — can the work be cut into independent pieces and reassembled? If it can, change the sequencing so at least one stage requires the whole team, and put a shared interface or document at the center of it. Tell the class which segments are meant to be AI-free, and why.
This term: decide AI permission per task, protect the framing and reconciliation stages, and read any group's restraint as a signal to interpret rather than a position to correct.
A practical checklist for group assignments
- Write the negotiation into the task. Require each team to submit a short, jointly authored statement of how GenAI will and will not be used, and grade the reasoning.
- Make the process visible, not just the product. Intermediate deliverables, peer assessment of contribution, shared planning documents and reflective contributions create the interactions through which norms get negotiated.
- Protect individual accountability alongside the group mark. Keep a component each member must answer for alone — an oral defense of their own understanding, or an individual explanation of any section they did not draft.
- Choose the access configuration deliberately. One shared interface produces collective prompts, shared memory and negotiated roles; private prompting produces opaque flows and de-labeled outputs.
- Design the task so it cannot be partitioned. A coherence rubric line does not produce collaboration on its own; the workflow has to require joint work.
- Say which parts are AI-free. If negotiation and collective sensemaking are the point, name the segments where the group works without the tool, and tell students why.
- Assign roles, especially for learners who need them. Explicit role definitions, structured assignments and small consistent teams are requirements AI collaboration tools routinely fail to accommodate.
- Read restraint carefully. A group that declines GenAI may be exercising deliberate self-regulation or guarding against unfamiliarity and risk; the two call for different responses.
- Keep the evidence standard honest. Engagement, enjoyment and satisfaction are weak indicators; group GenAI use under a permissive policy with reduced peer interaction is not evidence of learning.
The underlying evidence sits in the Group Work and Collaborative Learning concept pages.