FAQ
How Do I Design Faculty Development for AI That Actually Changes Practice?
If you run faculty development, the pressure you are under is probably one of these: your provost wants an AI program this year, attendance is fine but nothing changes in classrooms, or the people who most need it will not come. All three point at the same design fault. What changes a teacher's practice is not learning what a tool can do; it is spending supported, structured time on a problem in their own course, long enough that a new habit replaces an old one, with the institution explicitly saying yes to the experiment instead of merely permitting it.
The short version. Stop running tool demonstrations. Standardize a session cycle in which every participant brings a real course artifact and leaves with a revised one, wrap verification, disclosure and risk control into the working template rather than an ethics add-on, stretch the program past a single day and pair it with mentoring, run discipline-specific cohorts, and measure the artifacts people produced at a delay rather than how confident they felt on the last afternoon. The rest of this page is the evidence for each of those choices, the institutional conditions that decide whether they hold, and what to tell a dean who asks whether it worked.
Why the workshop model underperforms
The most direct comparison in this literature is Bi et al. (2026), who surveyed 568 faculty across five universities and then ran a quasi-experiment with 160 of them. Both arms got four 90-minute workshops over consecutive weeks with identical contact time, the same platform and the same facilitator support. The difference was the content: one arm had conventional GenAI faculty development, the other worked a structured cycle of prompt formulation, output review, course-material revision and written reflection against a seven-component template covering objectives, task constraints, assessment criteria, verification, disclosure and risk control. The structured arm scored higher on all seven readiness dimensions, with its largest advantage in prompt design at roughly two standard deviations, and it was the only arm whose prompt-design gain survived delayed testing — the conventional arm's smaller gain had fallen back below baseline by follow-up.
Two more results explain why demonstrations fail to travel. In Bai and Hsieh (2026), who surveyed 898 university teachers on AI-TPACK scales, teachers who understood how generative tools work, or could operate them competently, were no more likely to report inventing and implementing new teaching practices than those who could not. The integrated AI-TPACK construct had no direct path to innovative behavior at all: its influence ran entirely through AI Literacy, teaching Self-Efficacy and professional identity. In other words, capability is mediated by how the teacher sees themselves, which is not something a tool walkthrough addresses.
There is also a ceiling. In a within-subjects study with 13 higher-education teachers, chatbot access raised the share of activities reaching two or more higher-order Bloom tasks from 41.7% to 91.7% (p = .01) and cut perceived cognitive effort from a median of 6.33 to 3.33 (p = .004) — but the pedagogical prompting training layered on top did not improve design quality further and slightly increased effort. And in a workshop with 60 middle-school science teachers, the productive element was collaborative review of ChatGPT-generated assessment questions, which is what surfaced conceptual-precision problems and the risk of reinforcing Misconceptions about AI. What those two studies share is that the learning came from critiquing and revising material, not from being shown material.
What to build instead
The design architecture that most closely fits the evidence above is Dogan's (2026) i-TPACK framework, built from a PRISMA review of AI professional development plus a search that surfaced 17 further studies. It aligns five knowledge domains — intelligent technological knowledge, technological content knowledge, technological pedagogical knowledge, integrated i-TPACK and AI ethics — with four evidence-based pathways: Active Learning, models and examples, coaching and expert support, and Feedback and reflection, governed by five principles running from Collective Synergy Over Silos to Flexibility within Structure. Its criticism of the field is worth taking seriously before you copy anyone's program: in many existing offerings, ethics is addressed superficially or absent entirely. The framework's limit is that it is conceptual, with no program-level outcome data.
Translated into a program you could run next term, that means:
- Every session ends with a revised artifact. Participants bring an objective, a rubric or an assessment task and leave with it changed. Bi et al. (2026) measured exactly this cycle, and it is the version that held up at follow-up.
- The template carries the hard parts. Verification, disclosure and risk control belong in the working document participants use all week, not in a closing segment. It is also how nervousness gets somewhere useful to go, which matters below.
- Treat assessment design as the highest-yield target. In Talebzadeh's (2026) eight-hour program across four two-hour sessions with 163 teachers and pre-service teachers, total AI-PCK rose from 151.65 to 220.14 (d = 2.36, p < .001) and the largest component effect was Assessment Rubrics (d = 2.19) — also the weakest area at pretest. The gap was narrower in teaching method than in assessment.
- Budget more than a day, and add mentoring. Pre-service teachers in that same program gained significantly more than experienced ones (p = .033), mostly in scenario-based task design and rubric design. The experienced teachers said eight hours was too short to shift deep-seated habits and asked for long-term, mentor-based support.
- Segment your audience instead of running one universal workshop. Sutedjo, Chowdhury and Liu (2026) found among 127 respondents that self-perceived pedagogical content knowledge was high (M = 4.70) and content knowledge near ceiling (M = 5.15), while technological pedagogical knowledge (M = 2.62), technological content knowledge (M = 2.75) and holistic TPACK (M = 2.55) were markedly lower. Content knowledge showed no significant correlation with any technology-integrated domain (r = .11–.15), so subject-matter mastery does not carry over, and those three low domains correlated r = .81–.91 with each other — plausibly one capacity to teach as a shared foundation rather than three separate modules.
- Decide whether relationships are in scope. Aponte et al. (2026) argue AI contributes to teachers' socio-emotional development only when it acts as relational infrastructure — protected time for mentoring, peer dialogue, communities of practice — and propose relational densification as the standard, counting trust, continuity of support and co-regulation rather than contact hours. Two of their claims translate directly: freed time is reabsorbed as extra demand unless the institution formally requires it to be reinvested, and conversational agents can quietly become a low-friction substitute for the difficult human negotiations that produce growth. Their framework is a heuristic, not a tested causal claim.
The institutional conditions decide more than the curriculum does
This is the finding to bring to a leadership meeting. In the same study, Kazakhstani faculty reported higher baseline readiness than their Chinese counterparts on all seven dimensions, with the widest gaps in disciplinary transfer, prompt design and basic understanding. Yet in hierarchical regression, the Kazakhstan coefficient — clearly significant on its own — fell to non-significance, a reduction of about four fifths, once prior GenAI use, recent AI training, institutional support, perceived permission, multilingual resource access, policy clarity and risk sensitivity were entered. Institutional support and perceived permission carried the most positive weight. Risk sensitivity was the only predictor working against readiness. More than twice as many Kazakhstani as Chinese faculty reported AI-related training in the previous six months.
For a program lead the levers are therefore local and proximate: hands-on practice, explicit permission to experiment, materials in the languages your faculty actually work in, and visible backing from the unit. National policy framing does not carry readiness on its own. Where risk sensitivity bites, the answer is not to talk caution down but to give it a place to be exercised — rubric design and verification routines are how nervousness turns into judgment.
Two practical frictions belong in the same conversation. Continued-use intention was the highest-rated satisfaction dimension in Bi et al. (2026) while tool usability was the lowest, so some of what looks like reluctance is platform friction. And Aponte et al. (2026) locate the hard limit: AI cannot compensate for chronic overload, punitive accountability cultures, weak leadership, or the absence of protected time. If those are your conditions, a better workshop will not fix them.
The skeptics and the anxious are the point, not the obstacle
Laidlaw (2026) argues from an autoethnographic account that resistance is not a skills deficit. Faculty asking what the point of teaching is in a GenAI world are not asking how to use a tool safely; the question is ontological. Reading GenAI as a threshold concept — transformative, troublesome, irreversible, integrative, bounded — she concludes that anxiety, resistance and confusion are necessary parts of crossing it rather than problems to be corrected, and her recommendations invert the standard workshop: open with identity questions rather than demonstrations, run discipline-specific cohorts, allow different timelines, and treat principled non-adoption grounded in disciplinary values as legitimate. She also warns that communicating rules can slide into an "enforcement illusion" that substitutes compliance messaging for support.
There is a design argument for adaptation here as well. In Bai and Hsieh (2026), professional identity predicted innovative behavior nearly twice as strongly among infrequent AI users as among daily users — a borderline result, but it suggests identity work pays off most with the people who barely touch the tools. Sutedjo and colleagues argue for deliberate discipline-specific programming because expertise does not transfer; Laidlaw wants discipline-specific cohorts as the setting for identity conversations; and Bai and Hsieh suggest differentiating by experience, with limited users needing foundational guidance and frequent users getting more from interdisciplinary projects.
For the anxious participant, Aponte et al. (2026) add a governance condition rather than a motivational one: separate well-being support from managerial evaluation, minimize data, limit purpose, keep participation voluntary, and guarantee human oversight. Those are what make it safe to admit uncertainty in an AI-mediated space.
The affective layer has its own instrument now. Vassallo (2026) surveyed 109 academics and identified an "AI guilt complex": 35% worried AI use undermines their credibility and 26% reported feeling they are "cheating", with anticipatory guilt about credibility exceeding remorse after use, so the distress is socio-professional rather than private. Four profiles emerged — Comfortable Adopters (27%), Guilty Non-Users (29%), Cautious Users (28%) and Morally Distressed Avoiders (16%) — which means a single programme is addressing four different problems, and the 29% who feel guilt while not using the tools need something other than a demonstration. The scale of the institutional task is visible in the AAC&U/Elon survey of 1,057 faculty (Watson and Rainie 2026): 68% said their schools have not prepared faculty to use generative AI for teaching and mentoring, and a similar share for scholarship, while 26% of respondents do not use the tools at all — including 40% of arts and humanities faculty — and 82% name colleagues' resistance as a barrier to departmental adoption.
Proving it worked to someone who funds it
The institutional scaffold worth aligning to is Crompton, Burke and Nickel (2026), whose design-based research across two macro cycles and 114 participants — faculty from 28 U.S. institutions and 11 countries, plus centers-for-teaching directors and accreditation coordinators — produced six standards spanning the whole academic role rather than the teaching slice of it: Instructor, Coordinator, Leader, Researcher, Learner and Contributor. The gap is real: existing frameworks and educator standards target K-12 teachers or address only teaching. Standards of this kind let a development center align its offer with appraisal and accreditation instead of running disconnected workshops, which is also the argument that survives a budget review.
Be careful how you evaluate, because most of this literature measures self-report. Bi et al. (2026) measured self-reported readiness, not enacted teaching or student learning, and its groups were formed by voluntary sign-up and institutional scheduling rather than randomization, so its results are associations rather than causal estimates. Bai and Hsieh (2026) is cross-sectional and self-selected, drew on eight universities in one country, and measured every construct from the same respondents at one time; their innovation scale also kept the wording of a general innovation measure. Sutedjo and colleagues cannot establish direction of causation, which matters because their strongest claim — that technological knowledge acts as a gateway — is exactly what a cross-sectional design cannot settle. Talebzadeh (2026) reports very large effects without randomization, follow-up or discipline-level breakdown, so d = 2.36 should read as short-term rather than as proof of durable practice. Laidlaw (2026) is an autoethnographic reflection, and Aponte et al. (2026) state that their framework is conceptual.
A defensible evaluation therefore measures artifacts rather than opinions: the course materials a participant revised, the rubrics written, the tasks designed. It tests at a delay, because the gap in Bi et al. (2026) appeared only at follow-up, where the conventional group had decayed and the structured group had not. It treats satisfaction and confidence as weak signals — Pishtari and colleagues (2026) note that quality gains alone do not prove internalization, since rising design quality can equally reflect delegating routine work or accepting AI-generated structure. And it adds at least one relational indicator, following Aponte et al. (2026), because improving artifacts while isolating participants is not what the program claimed.
Three objections, answered
"We do not have the hours." Then spend them differently rather than adding them. The comparison in Bi et al. (2026) held contact time constant at four 90-minute sessions; the structured arm outperformed on the same budget. Six hours arranged as an artifact cycle beats six hours of demonstrations, and a shorter format spent on assessment rubrics targets the weakest area in Talebzadeh (2026).
"Our faculty will not come, or the skeptics will derail it." Do not recruit the skeptics into a demonstration. Open with identity questions, run discipline-specific cohorts, allow different timelines, and say out loud that principled non-adoption is a legitimate position — that is the recommendation from Laidlaw (2026) and it removes the argument about compliance from the room. Expect identity work to matter most with infrequent users, per Bai and Hsieh (2026).
"We cannot make people act on it." You mostly cannot, but you can change what the institution signals. Perceived permission and institutional support carried the most positive weight in Bi et al. (2026) after the national-background effect collapsed, so the useful moves are the cheap ones: state that experimentation is sanctioned, protect the time formally so it is not reabsorbed, provide materials in faculty's working languages, and remove platform friction, which was the lowest-rated dimension in that same study.
Checklist for program designers
- Anchor every session in a real course artifact — objectives, rubric or assessment task — that participants bring with them and revise before they leave.
- Put verification, disclosure and risk control inside the working template rather than adding ethics as a closing unit.
- Run long enough, and pair the program with mentor-based follow-up; experienced staff report that short formats cannot shift deep-seated habits.
- Build discipline-specific cohorts and differentiate by experience level instead of running one universal workshop.
- Open with identity questions before demonstrations, and treat principled non-adoption as a legitimate professional position.
- Secure institutional support, perceived permission, protected time and access in faculty's working languages; these carry more weight than national policy framing.
- Evaluate with artifacts, delayed measurement and at least one relational indicator; report effect sizes honestly against self-report and cross-sectional designs.
For the competency target itself see What Competencies Do Faculty Need in Regard to AI?; for structuring AI into a course rather than into a workshop, see How Should AI Be Designed Into the Learning Experience?; and for the wider institutional picture, see What Are the Top 10 Findings from AI in Education Research That Instructors Should Know About?.