AI Enabled SAFe
36 MIN READ

DORA 2025 for AI Teams: Seven Archetypes Every SAFe RTE Should Know

DORA 2025 swaps tiers for seven archetypes — and AI amplifies what your SAFe teams already run. Learn the RTE playbook to diagnose trains and target interventions.

DORA 2025 for AI teams delivers an uncomfortable finding for every SAFe RTE: seven archetypes, not one performance curve, describe how Agile Release Trains respond to AI adoption. Identical tooling across ten trains produces evidence for all seven; because AI magnifies whatever pattern, healthy or dysfunctional, a team already runs.

Table of Contents


DORA 2025: What the State of AI-Assisted Software Development Report Changed

DORA 2025 retitled itself the State of AI-Assisted Software Development report, and in one release replaced the program’s familiar elite-to-low performance tiers with seven team archetypes while introducing the inaugural AI Capabilities Model as a companion framework. That renaming carries a claim most SAFe leaders haven’t registered yet: the research no longer treats AI adoption as a variable worth measuring on its own, because the report assumes every team already has it, and moves the research question to which organizational conditions convert that adoption into performance.

The DORA program, run out of Google Cloud, has published an annual State of DevOps report for over a decade. The 2025 edition breaks with that history in two structural ways at once. First, it retires the elite/high/medium/low ranking readers have used for years in favor of Team Performance Profiles: the seven archetypes this guide walks through section by section. Second, it debuts the AI Capabilities Model in the same release, a separate but connected framework naming the organizational practices that determine whether AI benefits scale (DORA. Both changes point the same direction: differentiation among teams no longer comes from whether they’ve adopted AI, but from what their delivery system was already like before AI arrived (Google Cloud.

From Elite Tiers to Seven Profiles

DORA scrapped its long-standing elite, high, medium, and low performance tiers in the 2025 report and replaced them with seven team archetypes derived from cluster analysis across eight organizational factors. A single rank collapses a team into one number; a profile locates it in a multi-dimensional space, which is why two teams that would have shared the “high performer” label under the old system can now land in entirely different archetypes.

That shift matters for RTEs because a rank tells you where a team stands relative to others, while a profile tells you what’s actually happening inside the team; whether the constraint is throughput, instability, or the human cost of sustaining the pace. Google Cloud’s own framing of the change describes the seven-profile model as a way to “connect performance data to experience” rather than simply score it, explicitly acknowledging that raw delivery metrics “tell you what is happening but not why” (Google Cloud. Team Performance Profiles, in other words, replace a leaderboard with a diagnosis.

The practical consequence for a SAFe portfolio is that comparing trains by tier no longer means anything useful; comparing them by archetype does, because it tells a portfolio exactly which capability gap separates one train from the next.

AI as the New Developer Baseline

DORA 2025 treats AI usage as the default condition of software work rather than an emerging trend, reporting that roughly 90 percent of respondents already use AI daily and describing organizational conditions, not tool adoption, as the source of remaining differentiation. More than 80 percent of those respondents believe AI has increased their own productivity, while 30 percent still report little or no trust in AI-generated code: a trust gap that has persisted since the 2024 edition, where 39 percent reported the same distrust even as 81 percent of organizations had already shifted resources into AI development.

Two entities carry this shift structurally. Performance Tiers, the old ranking, no longer exist as a category; the report’s Developer Baseline assumption now starts from ubiquitous AI use rather than treating it as a variable to isolate. And Rework Rate, the report’s newest formal metric, gets elevated to first-class status in the same release specifically because near-universal adoption changes what “delivered work” means: a detail this guide’s measurement section owns in full.

That near-universal adoption is exactly why some teams turn AI into faster delivery while others turn it into faster, better-documented failure.


The Amplifier Thesis: AI Magnifies What Your System Already Is

DORA’s central 2025 finding is that AI produces an amplifier effect rather than uniform acceleration, magnifying the strengths of well-run engineering organizations and the dysfunction of struggling ones in equal measure. That single sentence resolves an anomaly the 2024 report couldn’t explain; teams with identical tool access were producing wildly divergent outcomes, and the divergence tracked existing system health, not AI sophistication.

Harvey’s Amplification Thesis

Nathen Harvey, DORA’s lead at Google Cloud, frames the 2025 finding as multiplication: AI takes whatever performance pattern a team already runs and scales it, strength and dysfunction alike. Google Cloud’s own engineering analysis of the survey states the thesis directly, “AI’s primary role is that of an amplifier: it magnifies the strengths of high-performing organizations and the dysfunctions of struggling ones” (Google Cloud engineering team, resolving what had looked, in the 2024 data, like contradictory evidence about AI’s value.

Harvey’s own summary of the report’s headline changes, posted alongside its release, frames this as the real surprise buried under the archetype news: a fifth metric, seven profiles, and a clustering method built on eight factors all exist to make one underlying claim measurable: that identical AI access produces divergent results because the access isn’t what’s actually driving the outcome (Nathen Harvey.

Teams find this counterintuitive because they expected AI to level the field between strong and weak delivery organizations. The opposite happens: AI accelerates whatever process already governs code from idea to production. Where review discipline, test coverage, and observability are strong, the acceleration compounds into faster clean delivery. Where those conditions are weak, the same acceleration compounds into more code shipped with more defects embedded in it, arriving faster than the team’s existing process can catch them.

IT Revolution’s Mirror Reading

IT Revolution’s analysis reads the 2025 report as a diagnostic instrument rather than a productivity study, arguing that an organization’s AI results produce a mirror effect that reflects its true underlying capabilities rather than creating new ones. The essay ties this back to research on high-performing enterprises more broadly, arguing that organizational wiring, not technology, resources, or talent in isolation, is what separates enterprises that turn AI into performance from those AI merely exposes (IT Revolution.

Because the report treats AI output as a signal rather than a goal in itself, leaders can use uneven results across an enterprise’s teams to locate exactly where system conditions are weakest, instead of guessing which team needs help based on complaints or gut feel. A train producing more defects after an AI rollout isn’t reporting an AI problem: it’s reporting a pre-existing quality-gate problem that AI made visible faster than a quarterly retrospective would have.

Leaders who treat that visibility as an AI failure and pull back on tooling miss the actual signal; leaders who read it as a system diagnosis get a free audit of exactly which capability to fund next.

Techstrong’s System-Failure Corollary

Trade coverage of the report, including analysis from outlets like Techstrong, supplies the negative case directly: struggling systems get faster at producing rework once AI enters the pipeline, without any corresponding repair to the process generating it. CircleCI’s own reading of the same survey, nearly 5,000 technology professionals, states the corollary in almost identical terms: “For teams with broken processes and dysfunctional systems, AI magnifies the chaos” (CircleCI.

This corollary matters most in budget conversations. A struggling train that buys AI licenses without first repairing its review, test, or deployment process isn’t accelerating its recovery: it’s accelerating the accumulation of the same defects it was already producing, just with a shorter interval between commit and incident. The corollary gives portfolio owners a diagnostic question to ask before approving any AI spend: is this train’s existing process good enough to amplify?

For trains that can’t answer yes, every dollar spent on tooling before fixing the underlying process buys faster failure, not faster delivery.

Gain Dissipation Beyond the Individual

Individual-level speed gains from AI dissipate somewhere between the developer’s keyboard and the organization’s delivery numbers unless the surrounding system has capacity to absorb them. A developer who completes a task 30 percent faster doesn’t produce 30 percent more shipped value if the code review queue, the test suite, or the deployment pipeline downstream can’t absorb the extra volume: the bottleneck simply relocates rather than disappears.

DORA’s own longitudinal data on adoption illustrates how much gets lost in that relocation: a 25 percent increase in AI adoption tracked with a 7.5 percent gain in documentation quality, a 3.4 percent gain in code quality, and a 3.1 percent gain in review speed, but only a 1.3 percent gain in approval speed and a 1.8 percent drop in code complexity; modest organizational numbers next to the productivity claims individual developers report (DORA. The gap between what an individual feels and what an organization measures is the dissipation this section names; the metric layer built to catch it belongs to this guide’s measurement section.

RTEs who measure AI impact only through individual developer telemetry are measuring the part of the gain that’s most likely to have already evaporated before it reaches anything the business can act on.

Reframing AI Investment for SAFe

SAFe portfolios that model AI adoption as a licensing decision are answering the wrong investment question, because the amplifier mechanism moves the return calculation from tool cost to system readiness. Two portfolios spending identically on AI seats can produce opposite ROI outcomes depending entirely on the System Conditions each train had in place before the spend: a variable no procurement conversation currently prices in.

Reframing the investment case means asking which capability gaps a portfolio needs to close before AI licensing generates Organizational Performance gains, rather than asking which vendor offers the best per-seat rate. AI Investment sized against system readiness produces a materially different budget line than AI Investment sized against headcount, and the difference compounds across every ART on a train that skips the reframing.

Portfolios that fund the tool before funding the system end up funding the same capability twice; once for the license, and again later for the rework the license didn’t prevent.


The Seven AI Capabilities: Storer’s Model of Conditions That Compound

DORA’s inaugural AI Capabilities Model, led by researcher Storer, names seven organizational practices, clear AI stance, healthy data ecosystems, AI-accessible internal data, strong version control, small batches, user-centric focus, and quality internal platforms, that determine whether AI gains reach beyond individual developers. None of the seven work in isolation, and treating the model as a checklist to tick off one capability at a time is exactly how enterprises under-deliver on what it promises.

Storer’s AI Capabilities Model

Storer leads the DORA research team’s construction of the inaugural AI Capabilities Model, published alongside the 2025 report as the prescriptive counterpart to the seven-archetype taxonomy. The model itself is licensed and distributed directly by Google as a standalone reference document, distinct from the report’s diagnostic chapters (DORA AI Capabilities Model.

The division of labor between the two frameworks is deliberate: archetypes describe where a team currently sits, while the capabilities model describes what to build to move it. An RTE who only reads the archetype chapter gets a diagnosis without a treatment plan; an RTE who only reads the capabilities model gets a treatment plan with no way to tell which teams need which parts of it.

Skip either framework and archetype coaching turns into guesswork about which of seven possible levers to pull first.

The Seven Named Capabilities

Seven named capabilities make up the model, and each earns its place because DORA’s data links it independently to stronger AI outcomes rather than to general engineering maturity alone. A DORA-affiliated researcher walkthrough of the model lists them directly:

  • Clear AI stance, an explicit, communicated policy on when and how AI may be used, reducing ambiguity that otherwise slows adoption
  • Healthy data ecosystems, internal data practices that keep AI usage from being blocked by organizational silos
  • AI-accessible internal data, policies and documentation written to be usable by both humans and AI systems
  • Strong version control, discipline that provides rollback safety when AI-generated code diverges from expectations
  • Small batches, work broken into units small enough that AI-assisted changes stay reviewable
  • User-centric focus, thinking that keeps AI effort tied to real human outcomes rather than internal metrics alone
  • Quality internal platforms, infrastructure that lets improvements scale across teams instead of staying local (DORA researcher walkthrough

None of these seven is exotic or AI-specific on its own, each is a familiar engineering discipline that DORA’s data now links explicitly to whether AI investment converts into measurable performance rather than noise. What makes the list a model rather than a checklist is that the seven items measurably interact, which the next section addresses directly.

Platform Capabilities That Scale Team Gains

Quality internal platforms convert individual AI wins into organization-wide ones by giving every team the same self-service paths, golden paths, and toolchains instead of forcing each to solve integration problems alone. DORA’s supporting analysis is explicit about the mechanism: platforms “provide a streamlined developer experience with self-service capabilities, clear golden paths, and integrated toolchains,” and platform teams “can use AI to solve real problems across entire organisations by introducing capabilities into centralised internal platforms” rather than leaving each product team to reinvent the same integration (GetDX.

For a SAFe enterprise running multiple ARTs, this is the difference between one AI integration effort and a dozen redundant ones. A platform team that builds a single, well-governed path for AI-assisted code review, testing, and deployment lets every team on every train inherit that capability immediately, rather than each team negotiating its own AI tooling stack with its own gaps.

Without a functioning platform layer underneath them, the model’s organization-wide scaling promise never leaves the individual teams that happened to build it well.

Cross-Capability Compounding

The seven capabilities interact rather than sit side by side, and version control is the clearest case: strong version discipline is what makes AI-generated code safe to review, which is what makes quality platforms safe to deploy through, which is what makes small-batch delivery viable at organizational scale rather than just at the level of one disciplined team.

Healthy data ecosystems work the same way from a different angle; they ground an AI system’s outputs in real organizational context, and without that grounding, a user-centric focus has nothing accurate to point the AI’s effort toward. A clear AI stance compounds with all of the above by removing the ambiguity that otherwise makes every other capability slower to adopt, since teams unsure what’s permitted move cautiously regardless of what tooling and process discipline they already have.

Building capabilities in a sequence that ignores these dependencies wastes the investment on capabilities the organization can’t yet actually use.

Cost of Building Capabilities Alone

A single capability built without its supporting conditions delivers a fraction of its modeled value, because the model’s benefit case assumes all seven are moving together rather than one at a time. An enterprise that mandates strong version control while leaving its AI stance ambiguous still faces the same adoption friction the stance capability was meant to resolve, regardless of how disciplined its version-control practice already is.

This is the sequencing question every SAFe portfolio business case for AI capability investment has to answer before funding anything: which capability gaps, closed in which order, actually unlock the value the other capabilities are already positioned to deliver. Getting that order wrong doesn’t just slow the rollout: it can make an enterprise’s later, correctly-sequenced investment look like it underperformed, when the real cause was the capability built alone the quarter before.

That sequencing question is exactly where a SAFe portfolio’s existing constructs matter, since some of the seven capabilities are already sitting inside the framework waiting to be recognized.


The Seven Team Archetypes: From Harmonious High-Achievers to Foundational Challenges

DORA 2025 sorts teams into seven archetypes generated from throughput, instability, and team well-being data, ranging from Harmonious High-Achievers at the top to teams facing Foundational Challenges at the bottom, with roughly 40 percent of the industry landing in the two strongest profiles. Two teams can post nearly identical delivery numbers and still belong to different archetypes, which is precisely the distinction the old tier system couldn’t make.

The Seven-Archetype Taxonomy

The taxonomy runs from Harmonious High-Achievers, who combine speed, stability, and high well-being, down through five middle profiles to Foundational Challenges, teams in survival mode with significant process gaps and high burnout.

ArchetypeOperating PatternApprox. Share of Industry
Harmonious High-AchieversHigh throughput, high stability, high well-being~20%
Pragmatic PerformersHigh throughput, high stability, average well-being~20%
Stable and MethodicalHigh stability, deliberate cadence, below-market speed,
High-Impact, Low-CadenceHigh-quality output, low deployment frequency,
Process-ConstrainedAdequate skill, long lead times from governance,
Legacy BottleneckHistorical systems block even willing teams~11%
Foundational ChallengesLow performance, high burnout, unstable systems~10%

Splunk’s review of the report describes the top two profiles plainly: Harmonious High-Achievers “excel across all dimensions, fostering a culture of high trust, performance, and psychological safety,” while Pragmatic Performers are “reliable workhorses” whose “team well-being and trust are only average” despite consistent delivery (Splunk. Scrum.org’s summary captures the bottom of the taxonomy just as directly: Foundational Challenges teams “don’t need AI; they need a reset,” while Legacy Bottleneck teams are “smart people trapped in old systems” where “AI can help, but slowly” (Scrum.org.

Reading the table left to right rather than picking one archetype in isolation is what turns this from a list of labels into a usable lookup for a specific train’s diagnosis.

Three Generating Dimensions

Three dimensions generate the seven-profile space: throughput, or how much moves through the pipeline; instability, or how often that movement breaks something; and team well-being, or whether the pace is sustainable for the people producing it.

These three collapse a wider set of underlying signals; DORA’s clustering method draws on eight measured factors, including burnout, friction, and time spent on valuable work, before reducing them to the three broad dimensions practitioners can actually reason about (efficientlyconnected.com. Throughput and instability were the report’s traditional territory; well-being is the genuinely new axis, and it’s the one that reshapes which teams count as strong performers.

An archetype only becomes actionable once an RTE can name which of these three axes is actually failing for a specific team: a train-wide AI rollout treats all three as the same problem, and that’s the mistake the next sections exist to prevent.

Well-Being as a Load-Bearing Dimension

Team well-being carries equal analytical weight to throughput and instability in the 2025 model, which is the single biggest scoring change from the old elite-tier system. Under the previous ranking, a team shipping fast while quietly burning out scored as elite, because the tiers measured only what moved through the pipeline and how often it broke.

Under the archetype model, that same combination produces a distinct, recognizably unstable profile: Pragmatic Performers score high on throughput and stability but land with only average well-being and trust, exactly the pattern Splunk’s analysis flags as a group that “consistently deliver with both speed and stability” while carrying hidden cost the old system never priced in (Splunk. The distinction between this profile and Harmonious High-Achievers is well-being alone; both groups deliver, only one of them can sustain it.

An RTE who only tracks velocity will keep mistaking a Pragmatic Performer for a Harmonious High-Achiever until the well-being gap surfaces the expensive way, through attrition or a burnout incident the metrics never predicted.

Profile Over Rank

Ranking teams against each other made sense under four tiers; profiling them against three independent dimensions does not, because two teams can occupy the same relative rank for entirely different underlying reasons. A Stable and Methodical team and a Legacy Bottleneck team might both post below-target deployment frequency, but the two numbers mean opposite things.

One team is choosing a deliberate, quality-first cadence that happens to run slower than the market expects; the other is being blocked by systems it doesn’t control. Coaching the first team to “move faster” ignores a deliberate tradeoff that’s working; coaching the second team the same way ignores the platform investment it actually needs.

Matching the wrong intervention to a shared rank wastes an entire PI’s improvement budget on a fix that was never going to address the team’s actual constraint.

Archetype as Operating Pattern

An archetype describes a team’s current operating pattern under present conditions, not a fixed maturity level the team is destined to occupy. Two teams profiled as Foundational Challenges this quarter are reporting where their system stands today, not a permanent identity assigned to the people on it.

The risk in misreading this distinction runs in a predictable direction: labeling a team by its archetype as though it were a character trait invites blame instead of intervention, and blamed teams tend to hide the instability data that made the diagnosis useful in the first place. The model exists to guide the next PI’s capability investment, not to sort people into permanent categories that follow them across reorganizations.

Mislabeling a team this way changes its incentives quietly; teams start managing the label instead of the underlying conditions the label was supposed to surface.


Measuring Delivery After AI: The Revised Metric Set and Rework Rate

DORA 2025 extends its classic four keys with a fifth metric, rework rate, because AI-generated code can inflate deployment and throughput numbers while the resulting defects only surface later as unplanned fixes. Read throughput alone under AI-assisted delivery, and a team drowning in its own rework looks identical, for a quarter or two, to a team that’s actually getting faster.

The Five-Metric Delivery Framework

DORA’s revised framework extends the four classic keys, deployment frequency, lead time for changes, change failure rate, and failed-deployment recovery time, with rework rate as the fifth formal metric, alongside reliability as a sixth, quasi-metric dimension.

MetricCategoryWhat It Captures
Deployment FrequencyThroughputHow often changes reach production
Lead Time for ChangesThroughputTime from commit to production
Change Failure RateStabilityShare of deployments causing failure
Failed Deployment Recovery TimeStabilityTime to restore service after failure
Rework RateStability (new)Share of work returning as unplanned fixes

As of 2025, the framework groups these into “five formal metrics and one quasi-metric,” organized under the same two broad categories the classic four keys always used, throughput and stability, with rework rate added to the stability side (Future Processing. The original four metrics, developed by Google’s DevOps Research and Assessment team, had already become the industry standard for measuring delivery performance before this year’s addition (GetDX.

A team can improve on all four classic metrics simultaneously while rework rate quietly climbs behind them, and only the fifth metric is positioned to catch that pattern before it reaches a retrospective.

The AI Case for Rework Rate

Rework rate measures the share of delivered work that returns as unplanned fixes rather than shipping cleanly the first time, and DORA elevated it to formal-metric status specifically because AI changes what “delivered” means. Plandek’s analysis of the addition describes it as tracking “how often teams have to push unplanned fixes to production,” picking up “instability that the original four metrics tend to miss” entirely (Plandek.

The mechanism is straightforward once named: AI-generated code can pass initial review and clear deployment gates while embedding defects that only surface as rework days or weeks later. A team’s deployment frequency and lead time can both look excellent through that entire window, because those two metrics only see the first pass through the pipeline, never the correction that follows it. RedMonk’s analysis frames rework rate as “a secondary proxy for measuring instability in systems” precisely because unplanned deployments are a more common early signal than outright failed ones (RedMonk.

An RTE reporting only the classic four keys during a quarter of heavy AI adoption is reporting a picture the report itself now considers incomplete.

Churn Detection: Change Failure and Rework

Change failure rate and rework rate together separate real acceleration from churn, because a rising change failure rate paired with rising rework signals the same code failing once, getting patched, and then failing again on the fix. Reading either metric alone hides this pattern: a stable change failure rate can coexist with a rework rate climbing steadily underneath it, since rework often surfaces as small, unplanned follow-up changes rather than outright deployment failures.

An RTE watching for throughput inflation should treat the two metrics as a single joint signal rather than two separate dashboard tiles: a train whose deployment frequency is rising alongside both change failure rate and rework rate is not accelerating, it’s churning through the same defects on a faster cycle.

Reading either metric alone lets a churning train look stable for a full PI before the underlying pattern becomes impossible to miss.

Value Stream as Verification Layer

Value Stream Management supplies the report’s end-to-end check on all five metrics, treating any local AI gain that doesn’t show up in measurable product or business outcomes as noise rather than progress. Google Cloud’s own framing puts this plainly: VSM “acts as a force multiplier, ensuring local productivity gains translate into measurable product performance instead of downstream chaos” (Google Cloud.

This is the measurement expression of the amplifier mechanism this guide’s second section already established: a local gain that never reaches the value stream is, by the report’s own logic, evidence the surrounding system absorbed rather than delivered it. Delivery Metrics measured only at team level can look uniformly positive while the Organizational Performance the value stream is supposed to produce stays flat.

A train can hit every team-level metric target while the value stream it feeds shows no measurable change, which is the report’s own definition of a false positive worth escalating.

Instrumenting Delivery Metrics First

The first instrument an RTE should stand up when AI adoption rises on a train is rework rate, because it’s the one addition to the classic four keys positioned to catch inflated throughput before a retrospective does. Change failure rate should go up alongside it from day one, since the pair together is what distinguishes acceleration from churn.

Value-stream-level tracking comes second, once team-level metrics have stabilized enough to be trusted; instrumenting VSM before the team-level numbers are reliable just adds noise on top of noise. This sequencing mirrors the capability-investment logic from earlier in this guide: measure the foundation first, then measure what it produces further downstream.

Skipping this instrumentation order means the first real evidence of a problem arrives as a production incident, not as a metric an RTE saw coming a sprint earlier.


Mapping DORA Capabilities onto SAFe: Where the Framework Already Delivers

SAFe already carries five of Storer’s seven AI capabilities through existing constructs, platform teams, the batch-size principle, customer centricity, and Built-in Quality, leaving clear AI stance and AI-accessible internal data as the two capabilities enterprises must add deliberately. That two-capability gap is smaller than most transformation roadmaps assume, and treating it as a five-capability rebuild wastes budget SAFe enterprises don’t need to spend.

Mapping Capabilities to SAFe Constructs

Five of the seven AI capabilities map onto SAFe constructs that already exist inside a functioning ART, which is why the crosswalk is a configuration exercise for those five rather than new construction.

DORA CapabilitySAFe ConstructStatus
Quality internal platformsPlatform teams, EnablersCarried
Small batchesBatch size principleCarried
User-centric focusCustomer centricity, design thinkingCarried
Strong version controlBuilt-in Quality practicesCarried
Healthy data ecosystemsArchitectural runway, enabler epicsPartially carried
Clear AI stance,Gap: requires deliberate addition
AI-accessible internal data,Gap: requires deliberate addition

The pattern behind the “carried” column is consistent: DORA’s own supporting research links solid source-control discipline, monitoring, observability, and small-batch delivery directly to teams “benefiting more from AI” than teams without those practices already in place (GetDX, which is exactly the profile a well-run SAFe ART already fits before AI enters the picture.

An RTE reading this table sees exactly two configuration decisions worth escalating to portfolio level, not seven separate capability-building initiatives.

Built-In Quality as Version Control

SAFe’s Built-in Quality practice carries the version-control capability’s intent, but the maturity tell is whether its gates are automated or dependent on a human remembering to run them. Automated CI pipelines, mandatory code review, and test-first practices under Built-in Quality already produce the rollback safety the capabilities model asks for when they’re enforced by tooling rather than by convention.

Where those same gates exist only as a checklist item in a Definition of Done, the capability exists on paper without existing in practice: an AI-generated change can still merge past a human reviewer who’s trusting the tool rather than verifying it, which defeats the entire point of the capability. The distinction between an automated gate and a manual one is invisible in a maturity self-assessment but decisive in an actual AI-assisted incident.

An ART auditing its AI readiness should check gate automation before anything else, since that single check predicts most of this capability’s real-world coverage.

Cognitive Load as Absorption Diagnostic

Team Topologies’ cognitive-load fit supplies the diagnostic for which teams on a train can actually absorb AI augmentation without their existing work suffering. A team already operating near its cognitive capacity doesn’t absorb new AI-assisted volume by shipping more: it absorbs it by producing more defects, since the added throughput has nowhere left to go but past whatever attention the team has left for review.

Matching platform and enabling-team support to the teams with the least remaining cognitive headroom, an approach drawn from the same integration literature connecting Team Topologies to SAFe delivery, turns this diagnostic into an actual prioritization tool rather than an abstract framework. A train running this check before an AI rollout can identify, ahead of time, exactly which teams need extra platform support rather than discovering it after defect rates spike.

Skip this diagnostic, and AI rollout lands hardest on precisely the teams least able to absorb it; usually the same teams already flagged as under strain before AI entered the conversation.

Platform Teams’ Coverage Limits

Platform teams cover the quality-internal-platforms capability only as far as their charter extends, and most ART-level platform teams were never chartered to also own AI-accessible internal data policy. A platform team that builds excellent shared tooling for AI-assisted development is solving a genuinely different problem than the one that governs which internal data those tools are allowed to read.

Mistaking platform-team coverage for full capability coverage is a natural error, since both problems involve the word “platform” and both sit adjacent to engineering infrastructure. But data-access policy and AI stance typically require enterprise-architecture or portfolio-level authority that no team-level platform charter carries, regardless of how well-resourced that platform team is.

This is exactly how the two real gaps in the crosswalk get missed: not through neglect, but through a reasonable assumption that good platform engineering already covers ground it was never chartered to cover.

Scaled Agile’s AI Integration Guidance

Scaled Agile Inc. has published its own Thriving in the Age of AI guidance acknowledging this integration point directly, positioning AI capability-building as a portfolio-level investment rather than a team-level tooling decision. The framework’s own steward treating this as a leadership-level concern rather than an engineering-team concern lines up with the two-gap crosswalk built out across this section.

That alignment matters practically: when Scaled Agile’s own guidance and DORA’s independent research point at the same two organizational gaps, clear AI stance and AI-accessible internal data, the crosswalk stops being one article’s interpretation of two separate bodies of work and becomes a shared vocabulary DORA readers and SAFe practitioners can both use in the same conversation.

That shared vocabulary is what makes the next section’s diagnostic procedure something an RTE can actually run, rather than a framework mapping exercise that stays on a slide.


The RTE Playbook: Diagnosing Archetypes and Targeting Interventions on Your Train

An RTE turns the archetype taxonomy into action by instrumenting throughput, instability, and well-being for every team on a train, placing each team in its archetype, and matching capability investment to that specific profile before adding any AI tooling. Skip the placement step, and every team gets the same generic AI rollout; exactly the scenario that produces seven different outcomes from one identical plan.

Diagnostic Profiling of Every Team

Profiling starts by instrumenting the three generating dimensions per team, throughput, instability, and well-being, using existing delivery telemetry plus a lightweight, recurring well-being pulse rather than a one-time audit. Throughput and instability data already exist in most ARTs’ delivery tooling; well-being data is the piece most trains have never systematically collected, and it’s the piece that separates a Pragmatic Performer from a Harmonious High-Achiever.

Once collected per team, the three numbers place each team in one of the seven archetypes described earlier in this guide, giving the RTE a train-wide map rather than a set of isolated team impressions gathered informally in hallway conversations.

A train that skips well-being instrumentation will keep mistaking a strained Pragmatic Performer for a genuinely healthy Harmonious High-Achiever until burnout makes the difference visible the expensive way.

Matching Interventions to Archetypes

Intervention matching means the capability investment differs by archetype rather than defaulting to the same AI enablement package for every team on the train. A Legacy Bottleneck team needs version-control and platform basics addressed before any new AI tooling; a Process-Constrained team needs governance and approval-path simplification before tooling can help at all; a Stable and Methodical team may only need small-batch coaching to convert its existing quality into faster cadence.

This is where the capability crosswalk and the archetype taxonomy meet in practice: an RTE isn’t choosing generically among seven capabilities, but reading a specific team’s archetype and asking which one or two of the crosswalk’s gaps that team is missing. Intervention matching turns a seven-item capability list into a two-or-three-item action plan per team, which is the only version of the plan that fits inside a single PI’s improvement backlog.

Applying a Harmonious-High-Achiever-style intervention to a Foundational-Challenges team wastes a PI and often makes the team’s existing instability worse rather than better.

Foundations Before AI Push

Foundational Challenges and Legacy Bottleneck teams need version control and platform basics in place before any AI tooling push, because amplification works on dysfunction too. This is the sequencing rule the earlier capability-cost analysis already established at the organizational level, applied here at the individual-team level where an RTE actually has to act on it.

In practice, this means an RTE should resist pressure to roll AI tooling out uniformly across a train just because leadership wants visible AI adoption numbers. A team without functioning version control or basic platform support isn’t ready to have its process amplified: it’s ready to have its process repaired, and only then to have AI applied on top of the repair.

Pushing AI tools ahead of these foundations is the single most common way a train converts a modest, contained existing problem into a train-wide incident that shows up in a PI-ending retrospective.

Re-Profiling Each PI Cycle

Re-profiling belongs inside the existing PI cadence rather than as a separate initiative, with Inspect and Adapt as the natural venue for reviewing each team’s current archetype. Because capability investments shift a team’s underlying conditions, a team’s archetype placement from two PIs ago may no longer describe what’s actually happening on the ground today.

Building re-profiling into the standing Inspect and Adapt agenda, rather than treating it as a one-time diagnostic exercise done once at the start of an AI initiative, keeps the map current as capability investments land and teams move between archetypes. This also gives the RTE a natural, recurring checkpoint to confirm an intervention actually worked rather than assuming it did.

A train that profiles once and never again is managing last quarter’s teams, not the ones actually running today’s PI.

Edwards and Fenton’s Practical Bridge

Edwards, a Developer Experience researcher at Google, and Fenton’s key-findings walkthroughs of the report supply the practical bridge from DORA’s research language to delivery-floor decisions RTEs can act on directly. Their sessions translate cluster-analysis findings and eight-factor scoring into a vocabulary practitioners already recognize, diagnosis, intervention, re-measurement, rather than leaving the report’s statistical framing as the only available language.

That translation matters because a research summary and an operational playbook read very differently: a summary tells an RTE what the seven archetypes are, while the practical bridge tells an RTE what to actually do on Monday morning with a train that has three teams showing up as Legacy Bottleneck. Without that bridge, the report’s findings stay accurate but inert.

Without a working translation layer, RTEs are left reading research summaries and improvising the operational equivalent themselves, one train at a time.

Recognizing Train-Level System Problems

A train-level system problem, not a coaching gap, is what most of a train’s teams clustering in low-stability archetypes actually signals. When the majority of teams on a train land in Legacy Bottleneck or Foundational Challenges profiles, the fix is capability investment at train or portfolio level, not another round of individual-team retrospectives asking the same teams to try harder.

This distinction determines where the escalation goes. A single struggling team is a coaching conversation between that team and its Scrum Master; five struggling teams on the same train sharing the same root cause is a platform, governance, or architecture-runway conversation that belongs with the RTE and portfolio leadership, because no amount of individual coaching fixes a shared system constraint six different teams are all running into independently.

Archetype Distribution

Archetype distribution is the shape of a train’s team-by-team profile results laid out side by side rather than read one team at a time; how many teams cluster in each of the seven archetypes, and where that concentration sits relative to the top two profiles that together make up roughly 40 percent of the industry. Reading distribution instead of individual scores turns a dozen separate team reports into one train-level signal an RTE can act on directly, rather than a stack of reports that each look like an isolated coaching problem.

Distribution matters because it separates two problems that look identical at the individual-team level: a single struggling team needs team-level coaching, while five struggling teams on the same train need shared platform or governance investment instead. A train where six of nine teams sit in Legacy Bottleneck or Process-Constrained profiles is reporting a system condition, not six unrelated coaching gaps, and the escalation path, from RTE observation to portfolio-level capability funding, depends entirely on reading the distribution rather than treating each team’s score as its own separate story.

When the distribution itself signals a train-level problem, the fix moves to the capability crosswalk this guide already built out, not to another round of team-by-team coaching that was never going to reach the shared root cause.


What DORA 2025 Cannot Tell You: ROI Caveats and the Agent Frontier

DORA 2025 is a strong diagnostic vocabulary but a limited proof: it’s survey-based and correlational, its archetypes are single-point snapshots with no longitudinal movement data yet, and LLM-powered agents are already changing what a “team” is for the taxonomy to profile. Trusting the model as a scientific guarantee rather than a shared vocabulary is the fastest way to over-commit a transformation roadmap to conclusions the data can’t yet support.

Foundations That Drive AI ROI

A 2026 InfoQ analysis of the report’s follow-on data links stronger engineering foundations directly to higher AI return on investment, validating the capability-first sequencing this guide has recommended throughout. Organizations with platform, quality, and version-control maturity already in place before scaling AI adoption see that investment translate into measured returns faster than organizations that bolt AI onto weak foundations and hope the tooling compensates.

This is the strongest available evidence that the two SAFe crosswalk gaps identified earlier, clear AI stance and AI-accessible internal data, are worth funding ahead of additional AI tool licensing, rather than treating tool spend and capability spend as interchangeable line items in the same budget conversation. Engineering Foundations, in this reading, function as a multiplier on every dollar spent afterward rather than a prerequisite that competes with AI spending for the same budget.

This finding is the single strongest argument available for sequencing capability investment ahead of AI tool licensing in any portfolio-level business case.

The Correlational Limits

DORA 2025 is a large-scale survey, and its findings are correlational rather than experimental, meaning the report can show that certain capabilities travel with better AI outcomes without proving that building any single capability causes the improvement on its own. Correlational Evidence of this kind is genuinely useful for prioritization, but it isn’t the same claim as a controlled study showing that adding version-control discipline to a specific team produces a specific measured gain.

Self-reported data on well-being and stability carries the usual limitations of survey research more broadly: recall bias, a non-random sample of participating organizations willing to complete a lengthy survey, and no independent verification of the underlying delivery numbers respondents report. None of this invalidates the findings, but it does mean the numbers describe patterns across thousands of organizations, not a guarantee about any single enterprise’s next quarter.

A portfolio business case built entirely on DORA correlations should still budget for its own before-and-after measurement, rather than assuming the survey’s aggregate findings transfer automatically to one specific train.

Archetypes as Static Snapshots

Each team’s archetype placement is a snapshot from one measurement window, and the 2025 report doesn’t yet include longitudinal data showing how teams actually move between archetypes over time. Nobody yet has published evidence on the typical path length from Foundational Challenges to Harmonious High-Achiever, whether that path is linear, or whether teams regress once capability investment attention moves on to the next train.

Longitudinal Evidence of this kind takes years of repeated measurement to build, and DORA’s archetype model is new enough that this first release simply hasn’t had the runway to gather it yet. That’s not a flaw in the model: it’s an honest limit on what a first-year taxonomy can claim, and treating the current snapshot as though it already describes a validated maturity curve overstates what the data supports.

RTEs should treat every re-profiling result as a fresh data point about a team’s current state, not as confirmation of a known trajectory the research has already mapped out.

Cognitive Agents Reshaping Teams

Cognitive agents built on large language models are beginning to fill functional roles inside Agile teams, and a 2024 study on agents operating within SAFe environments suggests the unit the taxonomy profiles is already starting to change under its own feet. The study investigates the integration of cognitive agents powered by Large Language Models within the Scaled Agile Framework to reinforce software project management, deploying virtual agents in simulated software environments to explore their potential to fulfill fundamental roles in IT project development and optimize outcomes through intelligent automation (Electronics, 2024.

Once a meaningful share of a team’s throughput comes from an agent rather than a person, human-focused well-being metrics capture a shrinking portion of what’s actually driving that team’s archetype placement. A team with two human developers and three cognitive agents filling supporting roles doesn’t fit cleanly into a model built entirely on human survey responses about burnout and friction, even if its throughput and instability numbers still register normally.

A taxonomy built entirely on human-team survey data may need a companion model of its own before agentic teams become common enough on SAFe trains to require one.

Diagnostic Vocabulary, Not Scripture

The right way to use DORA 2025 on a SAFe train is as a shared diagnostic vocabulary for naming what a team is experiencing, not as a scripture that dictates a fixed maturity path every train must follow in order. Adopting the archetype and capability language gives RTEs, Scrum Masters, and portfolio leaders a common way to describe what they’re all observing, which is valuable even before any of the correlational or longitudinal caveats above are fully resolved.

The honest next step after adopting that vocabulary is to re-measure your own train rather than assuming the survey’s industry-wide averages describe your specific teams: the 40 percent figure for the top two archetypes says nothing about whether your train’s teams sit inside or outside that share.

The model earns its keep the moment it changes which capability a portfolio actually funds next, not before.


Summary

DORA 2025 replaces a single performance rank with three things a SAFe train can actually act on: a seven-archetype diagnosis, a seven-capability build list, and a metric set honest enough to catch AI-inflated throughput before it becomes a production incident.

Read Archetypes, Not Rankings

Treating a team’s DORA placement as a diagnosis rather than a grade changes what an RTE actually does with a disappointing result. The taxonomy exists to answer a narrower question than “is this team good or bad”: it answers which of three dimensions, throughput, instability, or well-being, is the one actually constraining a given team right now, and the well-being axis is usually the one a velocity-focused RTE overlooks first.

The decision rule that falls out of this is straightforward to state and easy to skip under delivery pressure: match the intervention to the archetype’s specific weak dimension, not to a generic “improve delivery” backlog item that treats every underperforming team the same way. A Process-Constrained team and a Foundational Challenges team can both show up as “behind” on a dashboard, and coaching them identically wastes the PI on the team whose actual constraint was never coaching in the first place.

The sequencing consequence compounds across a whole train rather than one team at a time: profile every team before committing to an intervention, re-profile at every PI boundary rather than once, and treat the shape of the whole train’s distribution, not any single team’s score, as the signal that escalates a problem from team-level coaching to portfolio-level capability funding. A train that only ever looks at its best or worst team misses the pattern sitting in the other seven.

Build Capabilities Before You Scale AI

The capability crosswalk and the amplifier mechanism point to the same capital-allocation decision from two different directions, and a portfolio that funds them separately is solving one problem twice. Because SAFe already carries five of the seven AI capabilities through Built-in Quality, platform teams, the batch-size principle, and customer centricity, the actual new spend a portfolio needs to authorize is narrow: a clear, communicated AI stance and a policy making internal data genuinely AI-accessible, funded ahead of additional AI tool licensing rather than alongside it.

Instrumenting rework rate from day one of any AI push, rather than waiting for a quarterly review to reveal a problem, is the cheapest insurance available against the exact failure mode this guide has traced from the amplifier thesis through to the measurement layer: a train that looks faster for two quarters while its actual defect rate climbs underneath the numbers everyone is watching.

None of this is a one-time transformation project with a defined end date. Given the model’s correlational limits, its lack of longitudinal movement data, and the coming shift toward teams that blend human and cognitive-agent work, the honest posture for a SAFe portfolio is continuous re-measurement of its own trains; using DORA’s vocabulary to name what’s happening, and its own data, gathered PI over PI, to decide what to fund next.

Privacy Preference Center