AI Enabled SAFe
39 MIN READ

Mollick’s Leadership-Lab-Crowd: AI Blueprint for SAFe Portfolios

2-3x individual AI gains never reach enterprise level. Mollick's Leadership-Lab-Crowd framework for SAFe Portfolios maps a fix through LACE and PI cadence.

Your teams are quietly achieving 2–3x productivity gains with AI tools. Individually, they cut hours-long tasks to minutes. But when you zoom out to the organizational level, the needle barely moves; 10–20% improvement at best. This disconnect isn’t a technology gap; it’s an organizational design problem that Wharton professor Ethan Mollick has documented extensively. Mollick’s Leadership-Lab-Crowd: An AI Adoption Blueprint for SAFe Portfolios names the paradox and the path: individual gains do not automatically become enterprise gains without a structured system to capture, scale, and govern them.

Table of Contents


Leadership, Lab, and Crowd: The Three-Part Structure That Makes AI Adoption Work

Mollick’s framework distributes AI adoption across three distinct functions, Leadership sets direction and removes fear, the Lab builds and tests organizational AI capability, and the Crowd experiments with AI in daily work, and the evidence shows that organizations deploying all three see adoption rates that shadow-use-only organizations miss by a wide margin.

The Three Parts and Their Functions

Leadership: Direction, Urgency, and Incentives

Leadership owns vision, urgency, and incentives. Its job is not picking AI tools: that belongs to implementation. Leadership names what an AI-augmented operating model looks like for the organization, creates urgency that moves teams beyond curiosity-driven experimentation, and establishes incentive structures that reward transparent adoption over hidden use. The mechanism here is psychological safety backed by resource commitment: when leadership publicly treats AI capability as a strategic priority rather than an IT pilot, the entire organization reorients how it evaluates tooling decisions. Mollick describes this as creating “explicit experimentation zones” where workers know the boundaries are safe rather than discovering them through policy violations (Mollick, Making AI Work).

This Direction-Urgency-Incentives triad is what distinguishes genuine leadership commitment from generic AI strategy documents. Direction answers “where are we going”: the specific AI-enabled future the organization is building toward. Urgency answers “why now”: the competitive or operational pressure that makes waiting costly. Incentives answer “what’s in it for me”: the structures that make individual adoption rational rather than risky. Leaders who articulate all three produce organizational alignment; leaders who articulate only direction produce memos.

The Lab: Centralized Capability Building

The Lab is a dedicated group, subject matter experts plus technologists and non-technologists, whose full-time job is building organizational AI capability. This is not a governance body or a policy committee. The Lab benchmarks AI performance on the organization’s real tasks, prototypes new workflows that the Crowd surfaces, and builds scalable solutions from individual discoveries. Mollick emphasizes that the Lab must be “ambidextrous”; simultaneously exploring what will be possible in six months while exploiting what works today Danish and US (thekb.eu). The most visible failure mode for the Lab is treating itself as a research project generator rather than a product development function: a lab that produces reports instead of deployable workflow patterns is not doing its job.

The Lab’s composition matters as much as its charter. Mollick’s guidance is that the Lab should include subject matter experts who understand the workflows being augmented, technologists who can build the integrations, and non-technologists who prevent the group from becoming engineering-dominated. At least some Lab members should come from the Crowd, the enthusiasts who discovered AI use on their own and started sharing, because they bring real patterns rather than theoretical tool knowledge. A Lab staffed entirely by external AI specialists or internal IT staff who do not know the workflows will build solutions that miss the actual pain points.

The Crowd: Distributed Discovery

The Crowd is the entire workforce using AI tools inside their daily work context: the people closest to the work who know which tasks actually benefit from AI augmentation and which do not. McKinsey’s 2024 survey found that 91% of respondents use generative AI for work, yet only 13% of companies have implemented multiple gen-AI use cases at scale (McKinsey, The State of AI). That gap, 91% individual use versus 13% organizational implementation, is the exact paradox Mollick’s framework exists to resolve. Experienced workers with deep domain knowledge can assess AI’s utility in ways junior staff cannot, because they understand which parts of their workflow carry real judgment and which are pattern-matching the AI can handle.

The Crowd is the primary discovery engine, not a passive audience for centrally developed tools. This distinction is what separates Mollick’s model from traditional top-down AI adoption programs. A centrally developed AI tool that does not match how experienced workers actually do their jobs will be ignored regardless of how well it was built. The Crowd discovers what works through trial and error in their actual context: the underwriter testing AI on one step of a 14-step process, the developer using it for test scaffolding but not code review, the marketer using it for first drafts but not final copy. These granular, context-specific discoveries are the raw material the Lab exists to scale.

Why All Three Must Operate as a System

The framework’s analytical spine is interdependence: each part fails differently when the other two are missing. Leadership without a Crowd produces mandates that nobody executes; policies written in isolation that employees ignore or work around. The organization gets paperwork instead of adoption. A Crowd without a Lab produces scattered hacks; individual productivity gains that never scale because nobody is responsible for turning one person’s clever prompt into a repeatable workflow. The organization gets pockets of excellence and an aggregate flatline. A Lab without Leadership produces unfunded demos; prototypes that prove something works but never get resourcing or strategic cover to deploy. The organization gets proof-of-concept graveyard (Sidecar AI).

The practical consequence for SAFe portfolios is that launching any single element, standing up an AI governance board (Leadership only), funding a research team (Lab only), or distributing tool licenses (Crowd only), produces visible activity without organizational impact. Mollick’s central claim is that AI adoption is an organizational learning challenge, not a technical one, and learning systems require feedback loops between all three parts (One Useful Thing). The organization that builds all three and wires them together creates a flywheel: Leadership funds the Lab, the Lab equips the Crowd, the Crowd feeds discoveries back to the Lab, and the Lab reports capability gains to Leadership. Break any link and the system stalls.

The Model’s Research Foundation

The framework originates from Mollick’s “Making AI Work: Leadership, Lab, and Crowd” essay, grounded in Wharton research on AI adoption patterns. The empirical hook is the adoption paradox: Danish occupational studies show 65% of marketers, 64% of journalists, and 30% of lawyers had already used AI at work by 2024, while US worker surveys show adoption rising from 30% to 40% between December 2024 and April 2025 (thekb.eu). These figures track bottom-up diffusion happening independently of organizational programs. McKinsey confirms that nearly two-thirds of organizations remain in experiment-or-pilot stages despite 88% reporting AI use in at least one business function (McKinsey). The gap between individual use and organizational capability is not closing on its own: it requires intentional structure, which is what the three-part model supplies.

This research base distinguishes the framework from the many AI adoption checklists circulating in the enterprise market. It is descriptive of what actually happens in organizations (secret cyborgs, scattered hacks, unfunded demos) before it is prescriptive about what to do about them. The prescriptions emerge from observed failure modes rather from theoretical ideal states. The Danish and US survey data also establish the baseline that later PIs measure against: an organization that knows 30-40% of its workforce already uses AI has a more honest starting point than one that assumes adoption begins when the sanctioned tool is deployed.

How This Differs from a Center of Excellence

The critical distinction is where innovation originates. A traditional Center of Excellence positions a central team as the source of expertise, best practices, and governance: the CoE discovers, codifies, and distributes. In Mollick’s model, the Crowd, distributed workforce experimentation, is the primary discovery engine, while the Lab is the accelerator and productization function. This is not a semantic difference: it changes where leadership invests discovery resources.

CoEs typically appoint subject matter experts to define standards, which works well for compliance-heavy domains where consistency matters more than innovation. But in AI, where no instruction manual exists and the capability frontier shifts every quarter, the most valuable use cases are the ones embedded in job-specific context that no central team can replicate (True Platform). A CoE can discover that AI helps with email drafting; it cannot discover that a specific underwriter’s 14-step risk assessment process has a specific bottleneck that an AI classification model can eliminate in step 9. Only the Crowd, the person doing the work, can surface that pattern. The Lab’s job is then to productize and scale the pattern, not to find it in the first place. This role inversion, the crowd discovers, the central team scales, is the structural innovation that distinguishes Mollick’s framework from the CoE pattern most organizations default to.

SAFe Already Hosts the Scaffolding

The article’s central operational claim is this: SAFe portfolios already contain the structures that map onto all three parts of Mollick’s framework, which means teams do not need to invent new organizational charts to adopt it. The Lean-Agile Center of Excellence (LACE) is the natural home for the Lab: it already owns transformation practice and has cross-ART reach. Strategic Themes provide Leadership’s directional mechanism: they articulate strategic intent, which is exactly the “clear vision” Mollick requires from Leadership. The Portfolio Kanban serves as the experiment intake system, routing crowd-discovered ideas through a governance process that feeds back to strategic decision-making. The PI cadence, particularly the Innovation and Planning (IP) iteration and Inspect and Adapt events, provides the Lab’s operating rhythm without requiring a parallel schedule.

The ICON Agility AI adoption roadmap for SAFe organizations validates this alignment: their phased model places the initial pilot inside a single ART across one PI, with Inspect and Adapt as the go/no-go checkpoint, exactly matching the Lab’s experiment-container structure Inspect and Adapt (ICON Agility). Organizations that recognize these existing structures save months of design time that would otherwise go into building a parallel AI adoption operating model. The mapping is not theoretical, each structure already has the charter, cadence, and cross-organizational reach that Mollick’s three parts require, which means the adoption timeline compresses from “stand up a new function” to “extend an existing charter.”


The Secret Cyborg Problem: Why Your Crowd Is Already Using AI and Hiding It

Mollick’s research reveals that the majority of AI use in organizations is invisible, employees hide their AI tool usage because disclosure risks looking replaceable, surrendering the personal productivity surplus they have earned, or violating unclear policies, which means official adoption metrics measure a fiction while real adoption runs in the shadows.

Why Cyborgs Stay Secret: Replaceability, Surplus, Policy Fear

The term “secret cyborg” describes an employee who uses AI tools regularly at work but conceals that use from managers and colleagues. Mollick identifies three distinct fear vectors that drive secrecy, and each requires a different organizational intervention.

How Replaceability Fear Blocks Disclosure

Employees worry that if managers discover AI does large portions of their job, they will look redundant. This is not abstract anxiety: it tracks the actual information asymmetry of performance reviews. An employee who produces 60% more output with AI than without cannot safely attribute that output to AI assistance without inviting the question “then what do we need you for?” The Wharton research on incentive design identifies this as the most structurally consequential fear, because it is rational even in organizations with stated no-layoff policies; employees have watched restructurings bypass those commitments before Wharton (Wharton Blueprint for AI Agent Adoption).

The consequence is a productivity paradox that operates at individual level: the employee produces more, the organization captures none of the gain, and both parties know the arrangement is unstable. The employee cannot safely claim the productivity they actually deliver, so they calibrate output to what looks reasonable without AI. The organizational productivity gain that could fuel reinvestment never materializes because it is invisible.

Why the Productivity Surplus Stays Hidden

Employees who have gained back two hours per day through AI use are reluctant to disclose that surplus because they have already reinvested it; into higher-quality work, better research, or simply sustainable working hours. Disclosure means the organization may claim that surplus through increased expectations or redistributed workload. The employee captures zero benefit from honesty and faces tangible downside. This dynamic operates silently and accumulates across teams, producing an aggregate productivity gain in official metrics that disguises the real gain underneath.

The surplus-hiding mechanism is economically rational in any organization where work is implicitly bid-based; whoever seems least busy gets the next assignment. An employee who discloses a two-hour daily surplus signals available capacity, which attracts more work without compensation adjustment. The employee who stays quiet keeps the surplus and delivers output at a sustainable pace. Organizations that cannot align individual and organizational incentives on this point will never surface their true AI-driven capacity.

How Policy Ambiguity Drives Use Underground

Many organizations have AI usage policies that are either restrictive (“do not input company data into public AI tools”) or unclear (“use AI responsibly”), creating a situation where virtually all actual AI use is technically non-compliant even when it is beneficial. Employees cannot safely distinguish between policy violations that matter and ones that do not, so they conceal all of it (Mollick, Detecting the Secret Cyborgs). The result is a policy regime that drives adoption underground rather than governing it.

Dashboards Measuring a Fiction

The governance consequence of secret cyborgs is direct and corrosive: official adoption metrics measure sanctioned tool usage, while real adoption happens in unauthorized tools, personal accounts, and shadow workflows. An organization whose dashboard shows 20% AI adoption while 70% of its knowledge workers regularly use AI tools has built its adoption program around a data fiction. Every decision tuned to those metrics, training investment, tool licensing, policy refinement, optimizes for the fiction rather than the reality.

This is not merely a measurement problem; it is a resource allocation problem. The Danish and US worker surveys show that real adoption is 30–40 percentage points higher than most enterprise tracking captures (thekb.eu). Organizations spending on sanctioned tools and formal training programs based on official metrics are systematically under-investing in the support their workforce actually needs. McKinsey’s employee survey found that only 29% of US workers currently report feeling “fully supported” in their AI use, while 44% report moderate-to-significant support and 22% report minimal or no support (McKinsey). The gap between actual use and perceived support widens as secret usage grows. The portfolio implication is that adoption budgets should be reallocated from tool licensing to support infrastructure, training, safe experimentation zones, and disclosure incentives, because the tools are already in use regardless of whether the organization pays for them.

Paying for Disclosure: The Incentive Redesign

The Wharton incentivization analysis proposes a direct remedy: reward disclosure rather than penalizing it. The mechanism has three components that must operate simultaneously to be effective.

First, guaranteed no-replacement commitments from leadership, publicly stated and credibly backed, that disclosure of AI use will not be used to reduce headcount or redistribute work. This addresses the replaceability fear at its source. Second, recognition structures that explicitly credit workflow innovation: “AI Contributor” designations, featured showcases in all-hands meetings, and time allocation for employees to develop and document their discovered patterns. Third, time credit; employees who contribute validated workflows receive explicit time back for capability building, often through Communities of Practice or dedicated innovation time (Knowledge at Wharton).

The logic is calibrated to the psychology it counters. Secret cyborgs hide because the costs of disclosure (replaceability risk, surplus loss, policy exposure) outweigh the benefits. The incentive redesign does not require employees to be altruistic: it shifts the cost-benefit calculation so that disclosure serves the employee’s rational self-interest. The valence.co implementation recommendation captures the operational version: “Require all internal reports to include a short note on how AI was or wasn’t used; and what value it added. In every leadership meeting, dedicate 5 minutes to share new uses of AI; live demos encouraged” (Valence).

The Psychological Precondition: Making Experimentation Safe

Incentive redesign works only when the psychological precondition is met: leadership must have made experimentation explicitly safe. This is not achieved through policy language. It is demonstrated through observable behavior; leaders sharing their own AI use, acknowledging failures without repercussions, and visibly protecting employees who surface controversial use cases. Fear converts cyborgs to secrecy faster than any incentive program can convert them back. The True Platform analysis frames this as psychological safety being the foundational layer: “Leaders must express a clear and inspiring vision for an AI-driven future, including addressing employee concerns about job security and fostering the psychological safety required for them to embrace new technology” (True Platform). Until that safety is demonstrated, the crowd stays underground regardless of what the incentive policy says.

The observable test for psychological safety is simple: do employees share AI discoveries unprompted, in public channels, without checking with their manager first? If yes, safety exists. If sharing happens only anonymously, in private messages, or after explicit permission, safety has not been established regardless of what the intranet policy page states. Leaders who read silence as compliance are misreading the signal; silence in AI adoption contexts almost always means fear, not satisfaction.

What an Open Crowd Produces

When the crowd does surface its innovations, the output is qualitatively different from what a central team discovers. An open crowd produces workflow-level patterns; specific task sequences where AI augmentation changes the time, quality, or cost equation. An underwriter discovers that AI summarization of claims histories cuts assessment time by 40% while improving accuracy on standard cases. A product manager finds that structured prompt chains for competitive analysis reduce research time from three hours to twenty minutes while surfacing competitor moves the manual process missed. A developer realizes that AI-generated test scaffolding catches edge cases the manual test design overlooked.

These are not generic productivity gains. They are job-context innovations that no central team would discover because the innovations live inside the specific intersection of domain knowledge and tool capability that only the practitioner occupies. The Lab’s mission is to receive these patterns, validate them against real baselines, productize them into repeatable workflows, and promote them across the organization. Without the crowd, the Lab has nothing to scale. The volume and variety of crowd-sourced patterns is itself a health metric for the overall framework: a portfolio where the Lab is productizing more patterns per PI than the previous PI is one where the feedback loop is working, and the ratio of crowd-discovered to lab-originated patterns should trend toward crowd-dominant over successive PIs.


The Lab Lives in the LACE: Running AI Experiments on SAFe Cadence

The LACE already owns transformation practice and holds cross-ART reach; renaming it an “AI Lab” adds nothing that extending its charter does not accomplish more cleanly, and the SAFe 6.0 AI integration guidance provides the experiment container: one ART, one PI, one workflow, with explicit hypothesis and Inspect and Adapt checkpoint.

Extend the LACE Charter, Don’t Rename It

The Lean-Agile Center of Excellence exists to drive and sustain agile transformation. It already has the mandate to change how work happens across ARTs, the organizational positioning to reach portfolio decision-makers, and the practitioner credibility that teams trust. The shortest path to a functioning Lab is not creating a new organizational entity: it is extending the LACE’s charter to include AI capability building as a transformation practice area. SAFe 6.0 explicitly positions AI integration as an extension of existing lean-agile practices rather than a separate initiative (SAFe 6.0 AI Framework).

This avoids the organizational friction that new entities generate: budget competition with existing transformation work, unclear reporting lines, and the startup phase where a new team builds legitimacy instead of delivering capability. The LACE already has the relationships with RTEs, the rhythm of PI Planning access, and the data from retrospectives and Inspect and Adapt. A LACE with an AI capability remit can start running experiments in the next PI instead of spending months establishing itself.

The composition of the Lab within the LACE follows Mollick’s guidance: subject matter experts from different ARTs plus technologists, with at least some members who emerged from the Crowd: the enthusiasts who figured out AI use on their own and started sharing it (Mollick). These practitioners bring real workflow patterns, not abstract AI knowledge, which is exactly what the Lab needs to productize.

One ART, One PI, One Workflow

The SAFe 6.0 AI integration guidance provides the experiment container pattern: one ART, one Program Increment, one workflow. The unit of experimentation is not “try AI on this team”: it is “test whether AI augmentation of this specific workflow produces a measurable improvement in flow efficiency, quality, or cycle time for this ART over this PI boundary.” Each experiment carries an explicit hypothesis, a measured baseline, and an Inspect and Adapt checkpoint where the experiment lives or dies.

The Hypothesis Baseline

Before the experiment runs, the Lab establishes the current state of the target workflow; cycle time, defect rate, preparation hours, or whatever metric the hypothesis names. This baseline is the reference point for go/no-go decisions at the PI boundary. Common starter workflows from the SAFe guidance include backlog refinement (measuring time from draft to ready) and dependency detection (measuring cross-team dependency identification lag) Portfolio Kanban (ICON Agility).

The baseline measurement serves a second function beyond the experiment itself: it establishes the measurement discipline that the organization will need when AI-augmented workflows become standard practice. Teams that learn to baseline before experimenting build the habit of evidence-based workflow decisions. Teams that skip baseline and rely on anecdotal before-after comparisons cannot distinguish real improvement from Hawthorne-effect noise.

The Inspect and Adapt Checkpoint

At PI end, the Lab presents the experiment results against the baseline. The I&A event, already on every ART’s calendar, becomes the natural review cadence. If the workflow shows measurable improvement, the Lab productizes it and promotes it to the Crowd. If it does not, the experiment is killed. There is no extension without re-hypothesis. This cadence-bound kill mechanism is what separates contained experimentation from open-ended pilots that drift into indefinite side projects.

The checkpoint produces explicit documentation regardless of outcome. A passed experiment yields a promoted-pattern package; workflow definition, baseline comparison, conditions-for-success map. A failed experiment yields a kill record; hypothesis, baseline, result, and the analysis of why the hypothesis failed. Both are equally valuable portfolio artifacts because both inform the next experiment’s hypothesis design. A portfolio that only documents passes is biasing its learning toward confirmation.

Promote or Kill: The Lab’s Two Outputs

The Lab produces exactly two outputs: validated workflow patterns promoted to the crowd, and kill decisions on hypotheses that failed their baseline. Everything else, research reports, tool evaluations, vendor comparisons, training materials, is intermediate work that serves those two outputs but is not itself an output.

Promoted Patterns as Organizational Assets

A promoted pattern is more than a documentation page. It includes the workflow definition, the measured improvement over baseline, the conditions under which the improvement holds (team size, domain, tooling context), and the adoption guide that lets other teams replicate the result without starting from scratch. The ICON Agility roadmap phases map this promotion across PIs: Phase 1 contains the pilot to one ART, Phase 2 refines the pattern based on I&A feedback, Phase 3 promotes to additional ARTs only after the pattern is validated (ICON Agility).

Each promoted pattern also carries a conditions-for-success map: the specific organizational context in which the improvement was measured. This map is what prevents blind replication: a pattern that delivered a 40% cycle-time improvement on a mature ART with stable velocity is marked as “validated for stable teams with established baselines,” not as “proven to work everywhere.” Later adopters read the conditions first and assess their own fit before adopting.

Kill Decisions as Learning Outputs

A kill decision is equally valuable. An honest kill, “the hypothesis was wrong, the workflow showed no improvement, we are not deploying this”, saves the organization from propagating an ineffective pattern across multiple ARTs. Labs that never kill are not labs; they are pet projects. The discipline of killing at the PI boundary is what keeps the Lab honest and maintains Leadership’s trust for the next experiment.

The portfolio should track kill decisions as explicitly as it tracks promoted patterns. A PI with three kills and one promotion has produced more organizational learning than a PI with four promotions and no kills, because the three kills tell the portfolio where NOT to invest next, which is often more valuable information than where to invest. The kill rate also serves as a check on the hypothesis quality: if every experiment passes, the hypotheses are probably too conservative, and the Lab needs to take bigger swings.

IP Iterations as Structured Slack for Capability Work

The Innovation and Planning (IP) iteration is the Lab’s primary cadence surface for capability building that does not raid feature capacity. SAFe already dedicates the IP iteration for innovation, training, and planning: the Lab uses this time to build prototypes, run benchmarks, and prepare the next PI’s experiment slate. This is not additional time; it is the existing slack structure applied to AI capability. Organizations that skip the IP iteration for AI capability work end up raiding feature capacity during the PI, which creates the very tension between delivery and capability building that the IP iteration was designed to prevent.

Communities of Practice provide a second surface. CoPs already bring practitioners together across ARTs to share knowledge on specific practices. A CoP focused on AI-augmented workflows, seeded by the Lab, driven by Crowd members who have discovered useful patterns, becomes a distribution channel for validated patterns and a source of new experiment ideas. The CoP structure is especially valuable for the feedback loop: practitioners using promoted patterns report back with modifications, refinements, and new use cases that feed the next experiment cycle. A well-functioning AI CoP should produce at least one new experiment hypothesis per PI, sourced from practitioners who adapted a promoted pattern to their own context and discovered an extension the original Lab work did not anticipate.

The Anti-Pattern: Framework-Wide AI Mandates Without Pilot Data

The anti-pattern the SAFe 6.0 guidance warns against is clear: framework-wide AI mandates issued before a single contained experiment has run. A mandate that every team must use AI for daily standup summaries sounds decisive but produces resistance, cynicism, and metric gaming; teams comply in form while continuing their real work unchanged. The contained experiment exists precisely to prevent this.

The discipline is simple: no pattern is promoted beyond the pilot ART until it has survived at least one I&A checkpoint. The first PI proves or disproves the hypothesis for that ART. Only a validated pattern, proven improvement against baseline, earns promotion. This constraint protects the portfolio from scaling noise instead of leverage. The mandate anti-pattern is most tempting when leadership feels pressure to “show progress” on AI adoption: a quarterly all-hands announcement about AI being mandatory across all ARTs looks decisive but usually represents the opposite of progress. The one-ART constraint is the portfolio’s defense against this pressure: it lets the organization demonstrate AI progress visibly while keeping the scaling decision tied to evidence rather than executive enthusiasm. An unproven pattern propagated across four ARTs creates four times the remediation surface when the pattern fails at I&A.


Portfolio Wiring: AI as a Managed Portfolio of Value Hypotheses

Every AI initiative enters the portfolio as a value hypothesis under a strategic theme, an epic with explicit leading indicators rather than an open-ended pilot, and this reframe converts AI adoption from a technology program into a portfolio management discipline.

Epics as Value Hypotheses, Not Pilots

The most common mistake organizations make when structuring AI initiatives is treating them as pilots; time-boxed explorations with vague success criteria and no governance mechanism for continuation or termination. The portfolio reframe changes this. Every AI initiative is an epic: it lives under a strategic theme, carries a Lean Business Case with explicit leading indicators, and moves through the Portfolio Kanban from funnel to implementation to completion under the same governance rules as any other epic.

The difference is the nature of the leading indicators. A traditional epic uses leading indicators tied to delivery: features shipped, stories completed. An AI epic uses indicators tied to capability validation: baseline established, hypothesis tested, workflow improvement measured, pattern ready for promotion. The HBR portfolio management analysis calls this “viewing the connected portfolio through a dual lens; as an advancement pipeline with clear gates through which projects must pass, and as a whole-portfolio dashboard that shows balances across risk and return, time horizon, capability areas, and mission alignment” (HBR).

An AI epic does not need to produce production features in its first PI. It needs to produce a validated or invalidated hypothesis. That is a legitimate epic outcome; just as a feature epic’s outcome is working software. The portfolio governance system should treat a well-run AI experiment that produces a clean kill with the same respect as one that produces a deployable pattern, because the organization learned something it can apply to the next experiment.

Nine Steps to Wave-Sequenced Epics

The PwC nine-step AI adoption sequencing playbook, adapted to SAFe portfolio mechanics, provides a structured approach to converting the blueprint market into actionable epic sequencing. The full sequence runs: readiness assessment, capability baseline, use-case identification, use-case ranking, wave grouping, epic definition, epic sequencing onto the Portfolio Kanban, execution with I&A checkpoints, and feedback into the next wave’s ranking.

Wave Grouping as the Critical Portfolio Step

Wave grouping is the critical portfolio step. Individual use cases from the ranking are grouped into implementation waves based on dependency, shared capability requirements, and domain clustering. Each wave becomes a portfolio epic or epic set sequenced onto the Portfolio Kanban. Wave 1 (the pilot wave) targets the highest-ranked, lowest-dependency use case: the one that can run on one ART, one PI, one workflow. Wave 2 expands to related use cases that share the Wave 1 capability pattern. Wave 3 addresses cross-ART patterns that require coordinated deployment (ICON Agility).

This wave sequencing prevents the portfolio from getting clogged with unconnected AI experiments that compete for the same LACE capacity. Each wave builds on the last, creating a capability stack rather than a collection of isolated projects. The Portfolio Kanban provides the visualization; at any PI boundary, leadership can see which waves are in progress, which are validated, and which are blocked. The wave structure also simplifies the go/no-go decision: if Wave 1 does not validate, Waves 2 and 3 are automatically postponed, preventing the portfolio from scaling unproven patterns.

Trust First: The Ordering That Decides Success

KPMG’s trust-first blueprint for AI adoption establishes the ordering principle: trust, then leadership, then technology (KPMG AI Blueprint). Organizations that start with technology selection, picking a platform and then building an adoption program around it, see adoption stall because they skipped the foundation layers. The Utah AI Leadership Blueprint follows the same ordering: foundation and governance before implementation; readiness assessment before tool selection (Utah AI Leadership Blueprint). The consistency across these independent frameworks suggests a structural pattern rather than a preference: trust enables leadership to set direction credibly, leadership enables technology selection to serve strategic intent rather than vendor convenience.

For SAFe portfolios, the ordering implication is that the first funded activity under an AI epic should not be vendor evaluation or platform procurement. It should be a readiness assessment that answers three questions: Is the ART’s current flow efficiency baseline stable enough to measure AI impact against? Have Leadership and the LACE established the safety conditions for crowd disclosure? Does the existing governance structure (Portfolio Kanban, Lean Budgeting, I&A) have the capacity to handle AI experiment intake without disrupting current delivery? An organization that answers all three affirmatively before writing the first tooling check has a fundamentally higher probability of adoption success than one that procures the platform first and searches for a problem to solve with it.

Guardrails as Leadership’s Primary Instrument

Leadership’s direct portfolio instrument is guardrails; constraints that govern how AI experimentation happens without micromanaging it. The two essential guardrails are WIP limits on concurrent AI experiments and risk gates that follow NIST-aligned decision criteria.

WIP Limits on Concurrent AI Experiments

WIP limits on concurrent experiments prevent the portfolio from overloading the LACE with simultaneous AI experiments. If the LACE has capacity to run three experiments per PI, the portfolio limits active AI epics to three. Additional epics wait in the Portfolio Kanban funnel. This seems obvious but is routinely violated because AI experiments feel less risky than feature work; they are small, they are learning-oriented, and each one individually seems harmless. Accumulated, they consume the same LACE capacity that transformation work depends on.

The WIP limit applies per value stream, not per portfolio. An organization with three value streams can run three experiments per PI per value stream, nine total, as long as no value stream exceeds its own limit. This prevents one value stream’s AI experimentation from blocking another’s, while preventing any single value stream from consuming the LACE’s full capacity. The LACE communicates its capacity at PI Planning, and the portfolio uses that capacity figure to set the WIP limit for the upcoming PI.

Risk Gates Following NIST-Aligned Decision Rails

Risk gates provide the compliance boundary for AI use cases, following the Utah AI Leadership Blueprint’s NIST-aligned decision rails: data sensitivity classification, output verification requirement, human-in-the-loop threshold, and escalation path for classification failures (Utah AI Leadership Blueprint). Every AI epic must pass these gates before entering the Portfolio Kanban’s implementing state. An epic that proposes using AI on customer-facing decisions without a human-in-the-loop mechanism does not progress until the mechanism is defined.

The risk gates are not static; they tighten or relax based on the organization’s demonstrated AI capability maturity. A portfolio running its first AI experiment uses conservative gates; human-in-the-loop on all outputs, strict data boundaries, mandatory output verification. After six PIs of clean operation with no compliance incidents, the gates relax for low-risk workflows. This dynamic gating prevents the most restrictive setting from becoming the permanent default: a trap that causes teams to bypass the gates entirely.

The Five-Phase Roadmap Spine

Underneath the epic structure runs a five-phase roadmap spine that aligns AI capability building with PI cadence: maturity baseline, readiness check, ROI framework, governance alignment, and change integration. Each phase corresponds to a PI or PI set and produces a specific portfolio-level output:

  • PI 0; Maturity Baseline: The LACE establishes current state: which workflows are candidates, what the baselines are, what the current adoption landscape actually looks like (including secret use).
  • PI 1; Readiness Check: The portfolio runs the readiness assessment, establishes guardrails, and selects the Wave 1 epic.
  • PI 2–3; ROI Framework: The first experiments produce data. The portfolio calibrates its ROI expectations based on actual measured improvements rather than vendor claims.
  • PI 3–4; Governance Alignment: Portfolio governance updates to incorporate AI epics as a standard epic category with standard policies, no longer requiring special handling.
  • PI 4+; Change Integration: AI capability building becomes embedded in the LACE’s normal transformation practice. No separate AI initiative exists: it is just how work gets done.

This spine ensures that portfolio-level AI adoption proceeds at a pace the organization can absorb while the PI boundaries provide natural forcing functions for decisions. The portfolio becomes a capability builder rather than a project allocator.


Blueprint Fatigue: Choosing One Operating Model in a Market of Eighteen

The Enterprise AI Executive compilation counts eighteen distinct AI adoption playbooks from universities, consultancies, and technology vendors; blueprint choice is now itself a decision problem, and for SAFe shops the selection argument comes down to structural mapping rather than theoretical elegance.

Eighteen Playbooks, One Decision

The landscape of AI adoption guidance is saturated. Wharton has its Blueprint for AI Agent Adoption, Utah has the AI Leadership Blueprint, KPMG has a trust-first blueprint, Microsoft has its Cloud Adoption Framework for AI, PwC has nine-step sequencing, Vector Institute has a health AI toolkit, ISED Canada has a small-business blueprint: the Enterprise AI Executive documentation catalogs at least eighteen distinct offerings (Enterprise AI Executive).

For an enterprise transformation lead, the problem is not lack of guidance but excess of it. Each blueprint makes its own case on its own terms, using its own maturity model and success metrics. Choosing between them requires criteria that no single blueprint supplies. The natural organizational response is to adopt none; to keep reading, keep evaluating, and never commit to an operating model. That response is itself a decision: it keeps the organization in the experiment-or-pilot stage where McKinsey reports nearly two-thirds of organizations remain (McKinsey). Blueprint fatigue is a real cost, not an abstract complaint. The cost is measured in months that pass while the evaluation continues and the gap between individual AI use and organizational capability widens. Every PI that starts without an adopted operating model is a PI where secret cyborgs accumulate and promoted patterns remain at zero.

Why Mapping Beats Elegance for SAFe Shops

Leadership-Lab-Crowd wins the selection decision for SAFe organizations not because the framework is more elegant than the alternatives but because its three parts map directly onto structures the organization already runs. Leadership → Strategic Themes and LPM governance. Lab → LACE. Crowd → ARTs and teams. The framework requires no parallel operating model, no new governance bodies, and no special funding mechanisms: it extends existing structures rather than replacing them.

Alternative blueprints often require organizational changes that compete with active transformation work. Some position AI adoption as a separate initiative with its own governance, which creates the shadow-process risk. Others assume a centralized AI team that reports outside the agile delivery structure, which creates coordination overhead the LACE would not face. For a SAFe organization mid-transformation, the cost of parallel structures is not just financial: it is attention and cadence coordination that the portfolio cannot afford to split.

The ICON Agility roadmap demonstrates this mapping at the level of specific SAFe events: PI Planning becomes the venue for AI experiment selection, ART Sync provides the weekly coordination for Lab activities, Inspect and Adapt provides the experiment review cadence, and the Portfolio Kanban provides the intake and WIP mechanism (ICON Agility). No event is added, none is replaced. The existing operating model absorbs AI capability building without structural modification.

Where the Model Strains: Regulated Value Streams

The model’s transparent admission is that regulated value streams, healthcare, financial services, aerospace, any domain where compliance boundaries constrain experimentation, strain the framework. The Crowd’s primary mechanism (try it and see) collides with regulatory requirements for validation, audit trails, and approved tooling. An underwriter cannot informally experiment with AI on a rate-setting workflow the way a product manager can experiment with AI on competitive analysis. This constraint is structural, not cultural: no amount of incentive redesign can override a regulatory requirement for approved tooling.

The adaptation is that the Lab must pre-clear the sandbox before the Crowd enters. In regulated contexts, the Lab identifies permissible experimentation zones, specific workflows where AI use is explicitly allowed under existing compliance frameworks, and communicates those zones clearly. The Crowd experiments only within those zones; everything else follows existing compliance processes. The Vector Institute’s health AI toolkit provides this pattern for healthcare settings: the “Lab” validates AI safety and efficacy within a compliance framework before the “Crowd” (clinicians) deploys it in patient care (Vector Institute). This reduces the crowd’s experimentation surface area but preserves the discovery mechanism where it matters most. The portfolio managing regulated streams should expect a 50-70% reduction in experiment volume compared to unregulated streams and plan LACE capacity accordingly.

The Adoption-Gap Honesty Check

Leaders systematically overestimate how close their organization is to crossing the adoption gap. The dashboard shows sanctioned tool uptake, the LACE reports experiment progress, and the training completion numbers look healthy. The secret cyborg data tells a different story: real adoption is happening at 2-3x the tracked rate, and none of it is flowing into organizational capability because the crowd is not surfacing its discoveries. Structure, LACE, guardrails, epic framework, is necessary but insufficient. The crowd’s trust is the binding constraint.

The honest check is the ratio of promoted patterns to secret users. An organization with 200 engineering staff, 140 of whom use AI tools regularly (the 70% from the US surveys), but zero crowd-sourced workflows promoted to validated patterns in the last two PIs, has a trust problem, not a structure problem. The portfolio governance is firing on empty. The fix is not a better epic template or tighter guardrails. It is the incentive redesign and psychological safety work described in the second section: the human-side work that no blueprint can automate. Leaders who find this conclusion uncomfortable should sit with that discomfort rather than commissioning another tool evaluation: the tool is not the bottleneck, and another vendor briefing will not surface a single hidden workflow.

The Two-PI Test: Has the Crowd Gone Public?

Mollick’s framework supplies a falsifiable test for organizational AI adoption: if two PIs after the Lab is operational and the incentive program is in place, the Lab has promoted no crowd-sourced workflow into a validated pattern, the crowd is still secret. The prescription is unambiguous: stop investing in structure and return to incentives. Extending the LACE charter further, adding more guardrails, or selecting a different AI platform will not help. The bottleneck is trust, and only leadership’s demonstrated behavior, not policy documents, can repair it (Mollick).

This test is the practical equivalent of a PI predictability measure. It gives the portfolio a concrete signal to track: count of validated patterns promoted per PI, sourced from crowd-discovered workflows. A PI where this number is zero is a PI where the system is not working. The portfolio does not need to invent a complex AI adoption maturity model. It needs to answer one question: did the crowd share something useful this quarter? The two-PI window accounts for the natural lag between incentive implementation and crowd behavior change: one PI to establish safety, one PI for the first patterns to surface and be validated. After two PIs with zero promoted patterns, the lag explanation no longer holds and the portfolio must conclude that either the incentives are wrong or safety has not been credibly established.


How Do You Sequence AI Adoption Across Multiple ARTs?

Scaling from a single-ART pilot to multi-ART adoption requires a sequencing decision at each PI boundary: the pattern that worked on one train may fail on another, not because the tool is different but because the team’s workflow baseline, domain context, or trust level is.

The first ART selects the pilot workflow and the pilot PI. The second ART does not adopt the same workflow; it adopts the Lab’s validation process. Each ART runs its own baseline measurement, applies the pattern, and measures its own delta. A pattern that produced a 30% cycle-time improvement on a mature ART with stable velocity may produce 5% on an ART that is still stabilizing its baseline; and that is valuable information, not a failure. The Lab records the conditions under which the pattern delivers, building a conditions-for-success map that later ARTs use to assess their own fit.

Cross-ART synchronization requires coordinating the experiment slates so that adjacent ARTs do not run conflicting or overlapping experiments in the same PI. The ART Sync event provides the coordination mechanism: each ART’s Lab liaison reports the current PI’s experiment, its status, and the next PI’s proposed experiment. The portfolio’s role is to ensure no more than one ART per value stream changes its workflow in any single PI: a WIP limit on concurrent workflow changes that prevents coordination debt from accumulating faster than the LACE can productize patterns.


What Metrics Should a Portfolio Track for AI Adoption?

The ICON Agility framework distinguishes leading indicators (tracked during the adoption PI) from lagging indicators (measured at Inspect and Adapt). This distinction prevents the vanity-metric trap that most adoption dashboards fall into; counting tool logins and training completions while the real adoption story stays hidden.

Leading indicators measure process change during the experiment PI: time spent on ceremony preparation (PI Planning prep, ART Sync inputs), story defect escape rate, and team confidence vote scores. These tell you whether the AI-augmented workflow is changing how work actually happens. A team using AI for backlog refinement that shows reduced time to “ready” status and stable or improving defect rates has a positive leading signal regardless of whether anyone has completed the formal AI training module.

Lagging indicators measured at I&A include PI Predictability (did the AI-augmented ART deliver its committed objectives?), flow efficiency (is the ratio of active to wait time improving?), and qualitative team sentiment from retrospectives. McNinsey’s state of AI data shows that only 39% of organizations report any EBIT impact from AI and most report less than 5%, which suggests that lagging indicators are the ones that matter for resource allocation decisions (McKinsey). The portfolio’s job is to look at the leading indicators for early signal and the lagging indicators for investment reallocation.

The metric that does not appear on either list is tool usage count. Organizations measuring AI adoption by how many employees logged into an approved AI platform this week are measuring an input, not an outcome. The relevant metric is validated workflow patterns promoted per PI: the number of crowd-discovered, lab-validated, organizationally promoted workflow changes that demonstrably improved flow efficiency or quality. A PI with zero promoted patterns is a PI where the system is not producing value, regardless of how many tool licenses were activated.


How Do You Handle AI Adoption in Regulated Environments?

Regulated environments, financial services compliance, healthcare patient data, aerospace certifications, constrain the Crowd’s experimentation freedom and require the Lab to operate as a compliance boundary, not just a capability builder.

The Lab’s first regulated-environment function is sandbox definition: identifying which workflows, data types, and AI tool categories are explicitly permitted for experimentation under existing regulatory frameworks. This is not a permissive default (everything not forbidden is allowed) but a bounded one (only these specified zones are open). The Vector Institute’s health AI implementation toolkit demonstrates the pattern: the Lab validates AI tool safety and efficacy within a defined compliance envelope before frontline practitioners are authorized to deploy it (Vector Institute).

The compliance sandbox has four structural elements. Workflow scope; exactly which tasks are eligible for AI augmentation. Data boundary; what data may enter AI tools and what may not. Verification requirement; when human review of AI output is mandatory versus when it is discretionary. Audit trail; how each AI-assisted decision is recorded for regulatory review. The Lab pre-defines these for each experiment zone, the Crowd operates within them, and I&A evaluates whether the zone should expand, contract, or terminate based on the evidence.

The strain on the framework in regulated contexts is real: bounded experimentation produces fewer discoveries because the exploration surface is smaller. The mitigation is that the discoveries it does produce are deployment-ready; they come with pre-clearance, so the time from experiment to promoted pattern is shorter than in unregulated contexts where each discovery must go through a separate compliance review. The portfolio should adjust its expectations accordingly: fewer experiments per PI, higher conversion rate per experiment.


When Should a Portfolio Kill or Pivot an AI Initiative?

The kill decision at portfolio level follows different criteria than the kill decision at experiment level: a portfolio initiative dies or pivots when the conditions that made its hypothesis plausible have fundamentally changed, not simply when a single experiment fails against its baseline measurement. An experiment dies when its hypothesis fails against baseline. A portfolio initiative dies, or pivots, when the conditions that made the hypothesis plausible have changed.

Three signals trigger a portfolio-level kill review. First, the experiment has failed two consecutive PI checkpoints on the same hypothesis. A single failure is learning; a second failure without a revised hypothesis is the portfolio funding an open-ended exploration. Second, the workforce conditions have shifted: a restructuring, a reorganization, or a policy change that makes crowd disclosure less safe. Third, the technology baseline has shifted: a new AI capability has made the experiment’s hypothesis either obsolete or trivially achievable.

The pivot decision is distinct from the kill decision. An initiative pivots when the experiment produced a validated negative: the hypothesis was wrong, but the data reveals a different opportunity. The Lab found that AI-assisted code review did not reduce defect rate (hypothesis failed), but discovered that AI-generated test scaffolding caught edge cases the manual process missed (unexpected discovery). The portfolio can terminate the original epic and open a new epic for the discovered pattern without wasting the PI that produced the data. This pivot pathway is valuable enough that the portfolio should budget for it: not every AI epic will hit its original target, but every well-run AI epic should produce actionable data that feeds the next epic’s hypothesis.


Summary

Mollick’s Leadership-Lab-Crowd framework gives SAFe portfolios an operating model for AI adoption that maps onto structures they already run, Strategic Themes, LACE, and ART cadence, while making the crowd’s trust, not the technology choice, the binding constraint on adoption speed.

The framework works because it distributes discovery and production across three load-bearing roles

Leadership’s job is not to pick tools or define workflows. It is to set direction, remove fear, and provide guardrails that govern experimentation without suppressing it. The LACE, extended to include AI capability building, runs the experiment container: one ART, one PI, one workflow, with baseline measurement and I&A checkpoint. The Crowd, employees doing their actual jobs, is the primary discovery engine, producing workflow-level innovations no central team could find. These three roles form a learning system: the Crowd feeds discoveries to the Lab, the Lab productizes and promotes validated patterns back to the Crowd, and Leadership funds the cycle based on I&A evidence.

For a portfolio manager or LACE lead designing next PI’s experiment slate, the design question is not “which AI tool should we adopt?” It is “which of our three parts is weakest?” If the Crowd is not surfacing discoveries, the fix is incentives and safety, not tooling. If the Lab cannot keep pace with Crowd discoveries, the fix is LACE capacity, not more experiments. If Leadership is not demonstrating visible commitment, the fix is behavioral change at the executive level, not a policy document. The framework’s elegance is that it transforms the AI adoption problem from an infinite surface of tooling decisions into a finite diagnostic of three functions: the portfolio always knows where to look when adoption stalls, and the answer is never “buy a different platform.”

The falsifiable test surfaces when the system is not working

The two-PI test is the portfolio’s most practical diagnostic tool: if the Lab has promoted no crowd-sourced workflow into a validated pattern after two PIs, the system is not producing value and the portfolio should stop adding structure and return to incentives. Leading indicators (ceremony prep time, defect escape rate) catch this earlier, but the promoted-pattern count is the signal that cannot be gamed. A PI with zero promoted patterns is the portfolio’s canary. The organization’s willingness to hear that signal and act on it, to stop investing in tooling and start investing in trust, is ultimately what separates organizations that cross the adoption gap from those that stay stuck in the experiment stage.

This test works because it bypasses every vanity metric that adoption dashboards rely on. Tool login counts, training completion rates, and policy acknowledgment signatures all correlate weakly with actual organizational AI capability. The promoted-pattern count correlates directly: it measures whether crowd discoveries are being validated and deployed across the organization. A portfolio that tracks this single metric and acts on it when it hits zero has a more honest AI adoption program than one tracking thirty metrics that all show green while the crowd stays silent.

Privacy Preference Center