AI Enabled SAFe
26 MIN READ

The Cybernetic Teammate: What Mollick’s P&G Study Means for SAFe Team Composition

Mollick's P&G cybernetic teammate study (NBER 33641) proves one person plus AI matches a two-person team—changing ART composition from headcount to capacity. How SAFe leaders can pilot, measure, and restructure ARTs for human-plus-AI teams.

Most SAFe leaders still size teams by counting chairs. The cybernetic teammate research suggests they have been using the wrong metric entirely; and the difference matters more in July 2026 than it did when the paper dropped. A single rigorously designed field experiment now licenses a fundamental rethinking of how ARTs should be composed, and the 2026 adoption wave makes that rethinking urgent rather than optional.


Where this article sits

Journey stage 1 of 7: Readiness

readiness use-cases roi pilots kpis operationalize scale

this articlelinkedjourney stagepillar

Your trail so far

The articles you visit light up on this map.

The Cybernetic Teammate Study: How Mollick and Colleagues Tested AI as a Team Member at Procter & Gamble

The canonical evidence for AI-as-teammate claims comes from a single NBER working paper, The Cybernetic Teammate, a field experiment randomising 776 P&G professionals into a 2×2 design: individuals or cross-functional teams, with or without GPT-4 access, testing whether AI can function as a genuine collaborator (NBER. The experiment paired commercial and R&D specialists on real product-innovation challenges rather than artificial lab tasks, and expert judges evaluated output quality blind to condition.

Study Title and NBER Provenance

The full title, “The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise”, appears in the NBER working paper series as number 33641, published April 2025. NBER working papers carry weight in organisational decision-making because they undergo rigorous pre-publication review while remaining accessible to practitioners rather than locked behind paywalls. The term “cybernetic teammate” was chosen deliberately by the authors to signal that AI was evaluated not as a productivity tool but as a collaborative entity that could replicate core benefits of human teamwork: collective performance, expertise sharing, and social engagement. The paper’s placement in the NBER series signals that the research met the evidentiary standards of labour economics, not just the AI research community.

NBER Working Paper 33641

NBER Working Paper 33641 provides the primary-source foundation for every claim in the cybernetic teammate discussion. The working paper runs approximately sixty pages and includes the full experimental design, statistical methodology, robustness checks, and supplementary analyses that the practitioner-oriented summaries omit. The paper documents the pre-registration of the experimental design, the randomisation protocol, the expert-judge evaluation rubric, and the effect sizes for each outcome measure; details that matter when a skeptical LACE member asks “what did they actually measure and how?” Accessing the paper directly rather than through secondary summaries also reveals the study’s limitations that the authors themselves document, including the single-company scope, the bounded-task context, and the specific statistical power calculations that determine which comparisons reached significance. For any RTE or transformation lead who intends to cite this study in an ART restructuring proposal, reading the full working paper is the precondition for credible citation; second-hand summaries inevitably flatten the evidence into a more convenient shape than the paper itself supports.

Harvard-Wharton-P&G Author Team

Eleven authors spanning three institutional contexts, Harvard Business School, the Wharton School at the University of Pennsylvania, and Procter & Gamble’s internal innovation team, co-designed and executed the experiment.

Ethan Mollick

Ethan Mollick, associate professor at the Wharton School of the University of Pennsylvania, is the most publicly visible author of the study through his “One Useful Thing” Substack newsletter, which reaches hundreds of thousands of subscribers with research-informed analysis of AI’s organisational impact. Mollick’s role in the study was not merely that of a public communicator; he co-designed the experimental protocol, contributed to the statistical analysis, and co-authored the working paper alongside his Wharton colleague Lilach Mollick. His companion essay “The Cybernetic Teammate” on One Useful Thing (March 2025) is the most accessible entry point into the study’s findings, compressing the sixty-page working paper into a six-minute read aimed at practitioners rather than academics. For SAFe leaders building a brief for LACE or ART leadership, Mollick’s essay serves as the executive summary and the NBER paper as the evidence appendix; both are necessary for a credible citation chain. The academic leads (Dell’Acqua, Ayoubi, Lifshitz-Assaf, Sadun, Mollick, Mollick, and Lakhani) brought field-experiment methodology and organisational behaviour expertise. The P&G-based authors (Yi Han, Jeff Goldman, Hari Nair, Stewart Taub) brought deep knowledge of the company’s product-innovation workflow, ensuring the experimental tasks reflected actual professional challenges rather than artificial lab puzzles. This composition matters because it addresses a common objection to AI studies: that they measure toy problems performed by gig-economy workers. The cross-sector author team means the study was grounded in real organisational constraints from day one.

776 Professionals in Field Setting

The sample of 776 was drawn from P&G’s professional workforce; employees with a decade or more of tenure on average, working on genuine product-innovation tasks their division needed solved. This is not a study of MTurk workers solving captchas or undergraduates completing word puzzles. The participants were experienced commercial and R&D professionals whose daily work involves exactly the kind of cross-functional collaboration that SAFe teams perform. Field experiments at this scale are rare in organisational research because they require an industry partner willing to randomise access to technology across their workforce. The study’s ecological validity, the degree to which findings transfer to real organisational settings, is its primary methodological advantage over lab experiments and survey-based studies.

The 2×2 Individual-Team Design

Participants were randomised into four conditions: individuals working without AI, individuals working with GPT-4 access, two-person cross-functional teams working without AI, and two-person cross-functional teams working with GPT-4 access. The cross-functional pairing paired a commercial specialist with an R&D specialist, mirroring how SAFe teams combine business and technical roles. All participants in the AI condition received identical training and prompt sets so the intervention was standardised. The experiment measured outcomes on quality (evaluated by at least two expert judges per solution), time spent, and participant emotional experience. This design isolates the effect of AI from the effect of teamwork and tests their interaction: a cleaner distinction than most organisational research achieves (Mollick – One Useful Thing.

Field Experiments for Organizational Decisions

Field experiments beat surveys and benchmarks for organisational decisions because they randomise the intervention and measure actual outcomes rather than self-reported intentions. A survey asking “would AI make your team more productive?” captures optimism bias. A benchmark comparing teams with and without AI on toy problems lacks external validity. The P&G study’s design, real employees, real work, randomised assignment, blind expert evaluation, provides the kind of evidence that should inform ART restructuring decisions. The same design logic applies inside any SAFe enterprise: run your own randomised comparison before committing to structural change, using the P&G protocol as your template.

Accessing the Working Paper and Essay

The full working paper is available on the NBER website at w33641 and through SSRN. Ethan Mollick’s companion essay “The Cybernetic Teammate” on his One Useful Thing Substack (March 2025, updated with community discussion) provides a practitioner-oriented summary that covers the experimental design, headline results, and organisational implications in approximately 6 minutes of reading: an efficient brief for LACE members who need the gist before engaging the full paper. The Harvard Business School faculty page for Dell’Acqua, Sadun, and Lakhani also hosts the abstract and citation metadata. These three sources, the NBER paper, the One Useful Thing essay, and the HBS faculty page, form the complete primary-source kit for any SAFe leader who wants to cite the study from original evidence rather than from a secondary summary.


One Person Plus AI Equals a Small Team: The Three Findings That Rewrite Team Math

Three experimentally demonstrated findings from the P&G field trial, performance parity, expertise democratisation, and positive emotional effect, each act as a distinct design lever for ART composition, not AI capability headlines. Reading them as team-design levers rather than technology news changes what a SAFe leader should do next, and each lever points to a different decision: how many humans a team needs, what skill mix to staff for, and how to frame the change story for the works council.

Performance Parity: The One-Person-Plus-AI Result

Individuals working with GPT-4 produced solutions judged at quality levels comparable to two-person teams working without AI (Dell’Acqua et al. 2025, NBER. This is the headline finding, and it carries a specific meaning for ART composition: AI does not simply speed up individual work: it replaces the collaborative benefit of a second knowledgeable human on bounded innovation tasks. The mechanism appears to be that AI supplies the missing reasoning from the absent role’s knowledge domain. A commercial specialist with AI can produce solutions that incorporate technical feasibility constraints; an R&D specialist with AI can produce solutions that account for market positioning. Each participant gained access to reasoning they would otherwise have needed a teammate to provide. The parity finding does not say AI replaces all team functions: it says AI replaces the cross-functional knowledge gap that teams traditionally fill. That distinction matters for ART design: composition decisions start from the question “what function does this team member serve?” rather than “how many people do we need?”

The Substitution Mechanism

The substitution mechanism operates through knowledge bridging rather than productivity amplification. When the P&G study’s commercial professionals produced technically grounded solutions with AI, the output quality matched cross-functional teams because the AI supplied the technical reasoning the commercial professional lacked: not because the commercial professional became faster at their own work. This distinction is critical for ART composition: AI substitutes for a missing role’s knowledge domain on tasks where that knowledge is the binding constraint. A team designing a customer-facing feature can productively reduce its human cross-section if AI covers the technical feasibility dimension, but a team debugging a production incident cannot substitute AI for the system architect who understands the deployment topology. The substitution is domain-specific and task-dependent, not universal. Mapping which team decisions involve cross-functional knowledge gaps, and which involve unique human judgment, determines where parity licensing applies and where it does not.

R&D and Commercial Boundary Erosion

Without AI, the study’s participants showed clear functional silos: R&D professionals proposed technically oriented solutions, commercial professionals proposed market-focused solutions. Cross-functional teams without AI produced balanced solutions through human collaboration. With AI, individuals in both roles produced more balanced solutions: the AI effectively bridged the knowledge gap that separates functional specialties (HBS Working Knowledge. This expertise democratisation effect means AI reduces the performance penalty of working outside one’s home discipline. For SAFe teams, the implication is direct: a team that might have required both a Product Manager and a Technical Anchor to achieve balanced decision-making can now operate with one human who has effective AI access, provided governance structures catch what the AI misses. The boundary erosion is not complete, the paper shows AI narrows but does not eliminate functional gaps, but it is large enough to change staffing calculations.

Expertise Democratization

Expertise democratization describes the mechanism by which AI enables professionals to produce reasonable output outside their core training, a commercial specialist generates technically grounded proposals, and an R&D specialist produces market-aware recommendations, without requiring a cross-functional human partner. The cybernetic teammate study found that AI access narrowed the gap between commercial-only and R&D-only solution scores by a statistically significant margin, effectively compressing the penalty for operating outside one’s home knowledge domain.

For ART composition, this finding licenses a specific structural change: teams staffed with generalists who have AI access can produce more balanced output than the same generalists working alone, reducing the need for dedicated cross-functional pairing on every team. The boundary condition is that governance must verify the AI-generated cross-functional reasoning: a Product Owner using AI to consider architectural trade-offs still needs a system architect to validate the reasoning, just less frequently than without AI. Organizations operating at scale discover that expertise democratisation reduces the minimum team size for cross-functional capability from roughly six human roles to four humans plus AI, freeing two roles for other value streams.

The Wellbeing Finding Nobody Quotes

Participants working with AI reported a more positive emotional experience than those working without it. This finding is mentioned least often in coverage of the study, yet it is the finding that most directly answers standard engagement objections to AI adoption. Works councils and team leads who worry that AI will create monitoring pressure, reduce autonomy, or increase stress can point to experimental evidence for the opposite effect within the study’s conditions. The mechanism the study suggests: AI reduces the friction of working outside one’s expertise zone, lowering frustration and increasing the sense of progress. The wellbeing effect is small-to-moderate by the paper’s reported statistics, but its presence contradicts the assumption that AI augmentation is inherently dehumanising.

Cross-Functional Teams Plus AI

Teams with AI outperformed all other conditions; they produced higher-quality solutions than individuals with AI, teams without AI, and individuals without AI (Mollick – Cybernetic Teammate. The interaction effect is the study’s most actionable finding for ART design: AI does not replace teams, it amplifies them. The best configuration is a human-human-AI triad, not a human-AI pair. The practical implication for PI planning: an ART that distributes AI access to existing teams gains more than one that uses AI to reduce team sizes. The parity finding tells leaders they can shrink teams in a pinch; the interaction effect tells them they should grow teams with AI.

The Amplification Effect

The amplification effect describes what happens when two professionals with complementary expertise each have AI access: the team’s output quality exceeds the sum of its parts because AI accelerates the cross-functional integration that human teams already do well. In the P&G study, the AI-amplified teams did not simply produce more output; they produced more integrated output, with commercial and technical considerations woven together rather than patched at handoff points. The mechanism is that AI reduces the transaction cost of cross-functional reasoning: the commercial specialist can independently verify technical constraints through AI rather than waiting for the R&D specialist’s availability, and vice versa. This frees the team’s synchronous collaboration time for the highest-value integration work. For ART design, the amplification effect means that AI access should be treated as a team-level investment; giving AI to every member of a cross-functional team generates more value than giving AI to individuals working alone, even though the individual gains are themselves substantial.

Bounded Innovation Task Limits

The P&G study measured performance on bounded product-innovation tasks: the kind of work done in a single session with a defined problem statement and success criteria. It did not measure sustained multi-PI delivery, integration with existing production systems, maintenance of long-lived codebases, or cross-team dependency management. These are precisely the activities that consume the majority of an ART’s capacity. The parity finding licenses composition experiments and pilot investment; it does not license reducing team sizes across an entire ART based on a single session’s evidence. Holding this boundary protects the study’s credibility with skeptical engineers who will correctly point out that one-shot innovation tasks differ materially from eight-iteration delivery cycles with production incidents.


SAFe Team Composition After the Study: From Human Headcount to Human-Plus-AI Capacity

SAFe 6.0 already treats AI-augmented agile teams as a force multiplier for speed and collective intelligence, meaning the framework has accepted the premise of human-plus-AI team design and the practitioner’s job is now disciplined application rather than persuasion (Scaled Agile Framework. SAFe 6.0’s team composition base is Team Topologies-informed, stream-aligned, platform, enabling, and complicated-subsystem teams (Skelton and Pais), and the cybernetic teammate evidence attaches to this existing structure rather than requiring a restructuring from scratch.

SAFe 6.0 AI-Augmented Teams Position

Scaled Agile’s own guidance positions AI-augmented agile teams as an organisational design choice rather than a tooling decision. SAFe 6.0 describes these teams as combining human creativity with AI technology to deliver customer and business value, organising around four core capabilities: Product, Builder, Domain Expert, and AI (SAFe AI-Native Teams. The Product capability defines the vision; the Builder capability uses AI to generate outputs; the Domain Expert capability provides specialised context; and the AI capability supplies data and automation. This four-capability structure maps naturally onto the cybernetic teammate findings: the parity result validates the AI capability’s ability to substitute for the missing function; the expertise democratisation finding validates the Domain Expert capability’s transformation; the wellbeing finding validates the human-centric culture SAFe emphasises. The framework’s adoption of this language means LACE members and enterprise architects can cite SAFe 6.0 itself as authorising composition experiments, not just the underlying research.

Team Topologies-Informed Team Structure

SAFe 6.0’s standard team types, stream-aligned, platform, enabling, and complicated-subsystem, each absorb AI teammates differently. A stream-aligned team whose value-stream involves significant content generation or data analysis (marketing ART, data-platform ART) can reduce its human count toward the lower end of SAFe’s 5-11 range because AI covers the generation or analysis function. A platform team responsible for shared infrastructure, where reliability and consistency dominate, likely keeps its human count higher and uses AI for test generation, incident analysis, and documentation; augmentation rather than substitution. An enabling team whose job is cross-ART capability building gains the most from the expertise democratisation effect, because its members must operate across multiple knowledge domains they do not individually own. Complicated-subsystem teams, where deep specialisation is the product, see the smallest composition change; their value is precisely the expertise AI cannot yet replicate. Mapping each team type to its AI-augmentation pattern prevents the one-size-fits-all error of treating “AI team” as a single configuration.

DORA and MITRE Readiness Assessment

The AI Maturity for SAFe Enterprises framework, built at the intersection of SAFe, DORA 2025, Team Topologies, and MITRE’s AI Maturity Model, structures the readiness conversation around four dimensions. Delivery Pipeline Readiness (assessed through DORA 2025 metrics: deployment frequency, lead time for changes, change failure rate, time to restore service) determines whether the ART’s delivery infrastructure can absorb AI-generated contributions without introducing risk. A team whose deployment frequency is already monthly should not add AI-generated code that increases change volume by 40% until its CI/CD pipeline can handle the load. AI Infrastructure Readiness (assessed through MITRE’s AI MM: data governance, model lifecycle management, monitoring capability) determines whether the enterprise can manage the AI tools themselves; important for regulated industries where model explainability and bias auditing are compliance requirements.

Autonomy and Coordination Cost Trade-Off

The 2022 International Journal of Information Systems and Project Management multiple case study on SAFe implementations documents a structural tension: when autonomous teams coordinate toward a common goal, they sacrifice some level of autonomy (IJISPM. AI teammates increase throughput, which raises the coordination bill because more output crosses more team boundaries per PI. An ART whose AI-augmented teams produce 40% more features must invest in coordination mechanisms, dependency management, system-level integration, cross-team retrospectives, at the same rate or the autonomy sacrificed outweighs the throughput gained. The study recommends explicit coordination capacity planning during PI Planning: if teams expect AI-augmented throughput to increase by N%, allocate corresponding PI time for cross-ART synchronisation events. This prevents the familiar pattern where productivity gains from tooling are eaten by coordination overhead the organisation did not anticipate.

Composition as Capacity Function

After the parity evidence, composition stops being a headcount question and becomes a four-variable function. Human cognitive capacity sets the upper bound on what the team’s humans can process individually. Pipeline throughput, deployment frequency, automation coverage, test reliability, determines how fast AI-generated work can reach production. Automated quality gates determine how much human review each unit of AI work requires: a team with strong unit test coverage and automated security scanning can trust more AI output without human validation than a team with manual regression testing. Governance maturity, change approval processes, compliance auditing, risk acceptance criteria, determines the organisational speed limit. The four variables interact: high governance maturity with low pipeline throughput means AI-generated work queues at the deployment gate; high pipeline throughput with low governance maturity means AI-generated risk reaches production before anyone notices.

Pipeline Throughput and Quality Gates

Pipeline throughput and automated quality gates form the operational half of the capacity function, determining how much AI-generated output an ART can actually absorb. An ART whose deployment frequency is weekly and whose regression testing is manual cannot safely integrate AI contributions at the volume the parity finding suggests is feasible: the automation layer becomes the bottleneck. The DORA 2025 metrics provide the assessment framework: deployment frequency must be at least weekly, lead time for changes must be under one day, change failure rate must be below 15%, and time to restore service must be under 24 hours before an ART can productively absorb AI-generated contributions at scale. Quality gates, unit test coverage above 80%, automated security scanning in the pipeline, integration tests that run in under 30 minutes, determine how much human review each AI contribution needs. An ART below these thresholds should invest in pipeline maturity before scaling AI teammates; an ART at or above them can proceed with the pilot confidently.

Sizing and Cognitive Load Boundaries

SAFe’s recommended 5-11 person team size reflects cognitive load and communication complexity constraints that do not disappear with AI augmentation. Cognitive load, from Team Topologies theory, refers to how much a team can hold in working memory about its domain; and AI reduces but does not eliminate the load of cross-domain knowledge. The same 2024 biomimetics paper on team-size optimisation found that communication channels grow quadratically with team size, and AI teammates add their own coordination channels, AI-to-human, AI-to-AI, AI-to-CI/CD, that current sizing models do not account for (Biomimetics. Early patterns from SAFe 6.0 adopters suggest 5-7 humans plus one or more AI agents as an emerging configuration, but this number is descriptive of early practice, not prescriptive. The correct sizing for a given ART depends on its four-variable capacity function, not on a recommended range.


Running Your Own Cybernetic Teammate Experiment Inside One Program Increment

A SAFe enterprise can replicate the P&G study’s design within a single PI rather than importing its conclusions; running a structured comparison between teams or individuals with and without AI assistance on the same workflow provides local evidence that carries more weight with skeptical stakeholders than any external paper. The protocol compresses naturally into SAFe’s existing cadence: one ART, one PI, one workflow, with System Demo checkpoints and a principle-based go/no-go decision at the PI boundary.

Scoping: One ART, One PI, One Workflow

The experiment-first integration pattern recommended by the SAFe AI Integration Survival Guide starts with one Agile Release Train, one Program Increment, and one well-defined workflow. Backlog refinement and dependency detection are the most frequently named candidates because they occur every PI, involve cross-functional input, and produce measurable outputs that can be evaluated blind. The workflow must be one that the ART already performs without AI assistance, so the comparison has a real baseline. Scoping to a single workflow on a single ART limits the blast radius if the experiment reveals problems with AI-generated content quality or coordination overhead. It also limits the change management surface: one RTE, one set of teams, one System Demo to modify.

Program Increment

A Program Increment functions as the natural time boundary for an AI teammate experiment because it is long enough to generate meaningful data, typically eight to twelve weeks of delivery work, and short enough that the control group does not feel permanently disadvantaged by not having AI access. SAFe’s PI cadence provides built-in measurement points: iteration System Demos, the Inspect and Adapt workshop, and the PI-level retrospective all collect data that the pilot evaluation needs. The PI boundary also defines the decision point: at the end of the experiment PI, the ART leadership evaluates the evidence and decides whether to scale, adjust, or revert. This alignment with the existing PI cycle is what makes the P&G study’s design compressible into SAFe without introducing new ceremonies. An ART running the pilot should designate the PI’s Inspect and Adapt session as the formal evaluation event, giving the evidence review a slot in the existing calendar rather than requiring a separate meeting.

Recreating the Comparison Cells in PI Cadence

The P&G design translates directly into the PI structure. Identify two comparable teams or four comparable individuals within the same ART. Assign AI access to one team or two individuals; the others work with their current tooling. All participants work on the workflow selected in the scoping step. The PI provides the natural time boundary: one PI of work produces enough data to evaluate while being short enough that the control group does not feel permanently disadvantaged. The comparison is not a formal randomised controlled trial, the ART cannot randomise team composition for experimental purposes, but a matched-comparison design with baseline velocity and quality data is sufficient to detect effect sizes in the range the P&G study found.

Measuring at the System Demo

SAFe defines the System Demo as an integrated view of new features each iteration, providing an objective measure of progress. It is the natural measurement event for the pilot because it already collects the data the evaluation needs. Each iteration’s demo should include three dimensions from the P&G study: output quality (rated by a reviewer who does not know which condition produced the work), solution balance (does the work reflect cross-functional considerations or only the author’s home discipline?), and participant experience (a brief survey). The System Demo’s existing audience, RTEs, Product Management, System Architect, provides ready-made evaluators. No new ceremony is needed; the pilot uses the ART’s existing heartbeat.

The Principle-Based Go/No-Go

The 2024 International Journal of Managing Projects in Business study on principle-based decision-making in scaled agile environments argues that benefits realisation in iterative delivery fits principle-based gates better than rule-based thresholds (IJMPB. The go/no-go should therefore ask: “Has the pilot produced enough evidence that AI augmentation, on this workflow, for this ART, in this context, improves output quality without degrading coordination or participant experience?” rather than “Has velocity improved by N%?” The principle-based frame avoids the trap of statistical significance in small samples: the question is directional evidence, not proof. If three iterations show consistent direction on all three measurement dimensions, the ART can commit to scaling. If the evidence is mixed or negative, the ART returns to baseline and documents what did not work.

Backlog Refinement Outcome Metrics

Backlog refinement as the pilot workflow produces specific measurable outcomes. Quality of refined backlog items can be rated on clarity, testability, and acceptance-criteria completeness by a Product Manager blind to whether AI was used. Cycle time from raw input to refined backlog item, minutes rather than days, captures the efficiency dimension that the P&G study measured as task completion speed. Cross-functional balance; do refined items address both technical and business considerations?; captures the expertise-democratisation dimension. The ART probably already has the data for these metrics in some form; the pilot formalises collection. These metrics serve double duty: they evaluate the pilot and they improve backlog refinement for teams that continue with AI.

Orchestrating the Human-Plus-AI Pilot

John Winsor and Jen Stave of the Digital Data Design Institute at Harvard frame orchestrating human-plus-AI teams as an operating-model design task, not a tool deployment. The pilot needs an explicit orchestration owner whose responsibility is coordinating the human-AI interaction patterns, not just provisioning the AI tool. This person, typically the RTE or a senior Scrum Master, tracks who is using AI for which decisions, where handoffs between human and AI occur, what governance questions the pilot surfaces, and how the team’s interaction patterns evolve. Without an orchestration owner, the pilot becomes a tool rollout and the organisational learning is lost. The orchestration owner reports at each System Demo alongside the feature demos, making the pilot’s process visible as well as its outcomes.


What the Study Does Not Prove: Skeptics, Limits, and the 2026 Adoption Reality

A single field experiment on bounded innovation tasks licenses pilot investment and composition experiments, not ART-wide headcount cuts: that distinction is the section’s central payload. The cybernetic teammate study is the strongest evidence yet that AI functions as a genuine collaborator, but its boundaries determine which decisions it supports, and the 2026 adoption context defines the cost of waiting for stronger evidence.

Steelmanning the Zitron Skeptic Objections

The Blacklynx Brief review explicitly weighed Ethan Mollick’s evidence against Ed Zitron’s AI skepticism and sided with Mollick based on observed professional practice; but Zitron’s objections deserve a fair hearing. Zitron argues that AI productivity studies measure narrow, well-defined tasks that favour AI’s strengths while ignoring the broader context of knowledge work: integration of multiple information sources, maintenance of organisational relationships, navigation of ambiguity, and exercise of professional judgment in non-routine situations. These are exactly the dimensions the P&G study’s bounded-innovation-task design does not measure. The steelman version of the objection: the parity finding is real but applies to a fraction of what knowledge workers actually do, and extrapolating from that fraction to whole-team composition is the same error as concluding that a faster typing speed makes a better writer. On balance, the Mollick evidence carries more weight because the P&G study uses real professionals on real tasks rather than students on puzzles, but Zitron’s boundary objection is correct and must constrain the study’s application.

Enterprise Adoption Evidence Limits

The P&G study provides one field experiment at one company: a large, well-known, sophisticated organisation with resources most enterprises cannot match. P&G’s internal AI readiness, data infrastructure, and organisational culture are not representative of the median ART. The parity result may not replicate in organisations with weaker data practices, lower AI literacy, or more constrained IT environments. The responsible position is to treat the study as the strongest available directional evidence, sufficient for pilot investment, insufficient for enterprise-wide policy change, and to commission local replications before committing to structural changes. Organisations operating at lower AI maturity should expect smaller and more variable effects, and should design their pilots with lower baselines and longer evaluation periods.

The 2026 Adoption Wave Context

Marc Zao-Sanders’ Harvard Business Review analysis (June 2026) reports 900 million regular ChatGPT users and Gemini surpassing 750 million users, with agentic workflows and vibe coding now mainstream. These numbers change the composition question’s urgency: the workforce’s AI literacy is no longer hypothetical; hundreds of millions of professionals already use these tools. The composition question is no longer “should we restructure for AI?” but “is our current structure absorbing the AI capability that employees already bring?” An ART whose team composition assumes no AI augmentation is designing for a workforce that may already be using AI tools on their own, creating ungoverned variance in how work gets done. The adoption wave also means the talent market has shifted: professionals who have developed AI-augmented working methods at current employers will expect their next employer to support those methods; and ARTs that cannot provide structured AI integration will lose talent to those that can.

International AI Safety Governance

The International AI Safety Report 2026, with a writing group including Sayash Kapoor and Rishi Bommasani, frames capability and deployment risks that enterprise team-design decisions must respect. The report’s governance framework identifies four risk categories relevant to ART composition: loss of human judgment in critical decision pathways, erosion of skill through over-reliance on AI recommendations, concentration of risk in pipelines where AI contributions go unmonitored, and accountability gaps when AI-generated output causes production incidents. These risks do not invalidate the cybernetic teammate evidence; they define the governance conditions under which AI-augmented teams should operate. An ART implementing AI teammates must demonstrate that it has addressed each risk category before scaling beyond the pilot phase.

Pilot Investment and Team Experiments

The study licenses specific, bounded investments now. A single-PI pilot on one workflow, as described in the previous section, is proportionate to the evidence. A controlled comparison within one ART, with explicit success criteria and a defined decision date, is the correct response to the study’s findings. What the study does not license: reducing team sizes across an entire ART, replacing junior roles with AI assistants across the board, eliminating cross-functional pairing because AI covers the function gap, or using the parity finding as the sole justification for headcount reduction. Each of those decisions requires replicated evidence specific to the organisation’s context, workflows, and maturity level. The study enables the pilot; the pilot enables the structural decision.

Enterprise Readiness Assessment Next Step

The proportionate commercial next step is a structured pre-pilot evaluation with an accountable owner and a defined decision date: a bounded commitment rather than an open-ended transformation program. The evaluation assesses the ART’s four-variable capacity function: human cognitive capacity (current team sizes, skill distribution, turnover rate), pipeline throughput (DORA metrics), automated quality gates (test coverage, security scanning, compliance checks), and governance maturity (change approval times, audit trail completeness, risk acceptance framework). Each variable maps to a specific assessment method. The evaluation produces a readiness score per dimension and a recommendation for pilot scope. The decision date, typically six to eight weeks, gives the evaluation a deadline without the pressure of a transformation mandate. Commissioning this evaluation is the single action that converts this article from reading into practice.


Summary

The cybernetic teammate study provides experimentally grounded evidence that AI-augmented teams can produce higher-quality output with different composition than traditional teams; but the translation from experimental finding to ART design requires structured experimentation, not wholesale restructuring. Three principles govern that translation: replicate locally before committing structurally, assess readiness on four dimensions before scaling, and design for human-AI-teams as a triad rather than a replacement.

Replicate Before You Restructure

The P&G study’s design compresses cleanly into one PI: one ART, one workflow, matched comparison. Running that pilot generates local evidence that carries more organisational weight than any external paper. The protocol described in this article, scope, recreate, measure, decide, gives an RTE everything needed to execute within a single PI Planning cycle. The cost of not running the pilot is either making structural decisions without local evidence or deferring structural decisions until external evidence is stronger than it will ever be.

Capacity Replaces Headcount as the Composition Metric

After the parity evidence, composition decisions start from a four-variable capacity function, human cognitive capacity, pipeline throughput, automated quality gates, governance maturity, rather than headcount targets. An ART that assesses its four variables can design composition changes that match its specific readiness level, avoiding both the overreach of cutting teams based on a headline finding and the inaction of waiting for perfect evidence. The assessment is simple enough to complete in a single iteration sprint and specific enough to produce actionable recommendations for pilot scope, team sizing, and governance adjustments.

Morné Wiggins · Agility at Scale · Talk to me

Privacy Preference Center