Table of Contents

SAFe Inspect and Adapt: Driving ART Improvement

Can a single quarterly event prevent your ART from repeating the same dysfunctions PI after PI? SAFe 6.0 positions Inspect and Adapt (I&A) as the structured improvement mechanism that closes the Program Increment loop; without it, the same bottlenecks, quality gaps, and coordination failures resurface every 8–12 weeks. I&A is not a longer iteration retrospective; it is a three-part, data-driven event designed to surface systemic problems at the ART level and produce actionable improvements that feed directly into the next PI Planning session.


What Is SAFe Inspect and Adapt?

The Inspect and Adapt event is SAFe’s structured continuous improvement mechanism, held at the end of every Program Increment, where all ART stakeholders evaluate the current state of the solution through a PI System Demo, quantitative metrics review, and problem-solving workshop to identify and resolve systemic issues PI System Demo (SAFe). Without I&A, the ART has no cadenced mechanism for closing the improvement loop: each PI Planning session replans without data on what actually worked or broke in the previous one.

I&A as a PI-level structured improvement event

I&A operates at the Program Increment level, not the iteration level. Every PI, typically 8–12 weeks, ends with this event as the final activity before the next PI Planning session begins. SAFe 6.0 frames I&A as the “inspect” half of the Plan-Do-Check-Act cycle applied at ART scale: teams plan at PI Planning, execute across four to five iterations, then inspect and adapt before planning the next increment. The event runs approximately three to four hours for a full ART, with the problem-solving workshop consuming roughly half that time. ARTs new to SAFe or operating with more than 125 people may need up to four hours to complete all three phases (LeanWisdom). The cadence is what separates I&A from ad-hoc improvement efforts: it is not optional, and it is not moved when feature delivery runs late.

The three pillars: System Demo, Metrics, Problem-Solving

I&A breaks into three sequential phases, each supplying the factual foundation for the next. The PI System Demo runs first, typically 60–90 minutes, during which all ART teams demonstrate integrated working software against the PI Objectives committed at PI Planning. The quantitative and qualitative metrics review follows, taking 30–45 minutes, where the Release Train Engineer presents program-level flow data, predictability scores, and quality trends. The problem-solving workshop closes the event, running 90–120 minutes, where the ART identifies the highest-impact systemic issue and works through structured root cause analysis to produce concrete improvement actions (PM Expert). The dependency chain between these phases is strict: a problem-solving workshop convened without data from the demo and metrics review produces opinions, not evidenced root causes.

How I&A differs from iteration retrospectives

The most common misunderstanding about I&A is that it is “a quarterly retrospective at program level.” The differences are structural, not scalar. An iteration retrospective focuses on team-level process improvement within a single two-week iteration, driven by team sentiment and qualitative observation. I&A operates at ART level with three structural differences: it is quantitative-first (metrics run before discussion), it focuses on the integrated system rather than team-specific practices, and its output is an Improvement Backlog that competes for capacity against feature work at PI Planning. Iteration retrospectives address team dynamics and local workflow issues; I&A addresses systemic bottlenecks, cross-team coordination failures, and program-level quality trends. A team retrospective will surface that a team’s stand-up runs too long; an I&A will surface that the ART’s dependency management process is causing 40% of stories to miss their PI commitment.

Who attends and what each role contributes

The I&A event involves every role on the ART. Agile Teams present the PI System Demo, demonstrating the features and capabilities they built. Business Owners evaluate the demo against business value expectations and provide the stakeholder perspective during problem selection. Product Management ensures the demo covers the features committed at PI Planning and validates that business value was delivered. The Release Train Engineer facilitates the entire event, prepares the metrics dashboard, and owns the Improvement Backlog after the workshop. System and Solution Architects contribute technical context during root cause analysis, particularly when quality trends or architectural bottlenecks surface as the primary problem. Scrum Masters support breakout facilitation during the problem-solving workshop, and the SAFe Program Consultant (SPC) may co-facilitate or coach the RTE through the I&A process for newer ARTs SAFe Program Consultant (ScatterSpoke). The participants from the PI Planning session reconvene for I&A: the same cross-functional group that planned the work now inspects what they delivered.


The PI System Demo: Evidence Before Opinion

The PI System Demo is the first and most critical phase of I&A: it establishes the objective evidence base by demonstrating integrated working software from all ART teams against the PI Objectives committed at PI Planning, distinguishing fact-based improvement from opinion-driven discussion before the metrics review or problem-solving workshop begin PI Planning (Jira Align). Without a substantive System Demo, the remaining I&A phases operate without the factual anchor that prevents the event from devolving into generalised complaints.

PI System Demo vs iteration demo: scope and audience

DimensionIteration DemoPI System Demo
ScopeSingle team, single iterationAll ART teams, full PI
FocusTeam-level story completionIntegrated feature delivery against PI Objectives
AudienceTeam, Product Owner, Scrum MasterAll ART stakeholders, Business Owners, Product Management
Assessment criteriaDefinition of DoneBusiness value delivered
OutputIteration feedbackObjective evidence for I&A metrics review
Failure modeDemo-as-checklistDemo-as-theater (no integrated system to show)

The iteration demo confirms that individual teams completed their stories. The PI System Demo confirms that the ART delivered integrated value. An ART where every team passes its iteration demo but fails the PI System Demo has a coordination problem; teams are building components that do not integrate into a working system. This distinction is why SAFe 6.0 positions the PI System Demo as the primary inspect mechanism: it tests integration, not completion.

Integrated demonstration across all ART teams

The PI System Demo requires all ART teams to demonstrate together, in a single session, showing how their features combine into an integrated whole. This is not a round-robin of individual team demos: it is a unified walk-through of the working system, typically organised by feature rather than by team. The distinction matters: when teams demo sequentially rather than integratively, the audience cannot assess whether cross-team dependencies were resolved or whether the system behaves as a coherent product. ARTs that struggle with integration often discover the root cause during their first few I&A events: the PI System Demo reveals that what each team built independently does not compose into what the ART committed to deliver. The demo environment should mirror production as closely as possible, and teams should have the demo environment prepared before I&A day to avoid the “demo setup” time sink that derails the event’s first phase (Agile Seekers).

Business Owner and stakeholder evaluation criteria

Business Owners evaluate the PI System Demo against the PI Objectives that were committed at PI Planning. Each PI Objective carries a planned business value score (typically 1–10), and the demo answers the question: “Did we deliver this objective at a level that achieves its planned value?” The evaluation is not binary; Business Owners can assign partial credit when an objective was partially met or when the delivered solution creates value in a different way than originally planned. This scoring feeds directly into the ART Predictability Measure, which compares total planned business value against total achieved business value for the PI. A predictability score below 80% triggers investigation during the metrics review phase: either PI Objectives were poorly defined, scope changed mid-PI in ways that were not renegotiated, or the ART committed to more capacity than it could deliver.

Demo outputs that feed the quantitative review

The PI System Demo produces three outputs that flow directly into the metrics review phase. First, the scored PI Objectives provide the baseline for the ART Predictability Measure: the single most important quantitative signal for whether the ART delivered what it committed. Second, the demo surfaces integration issues that become candidate problems for the problem-solving workshop: features that could not be demonstrated together indicate coordination failures. Third, stakeholder feedback from the demo identifies features that met technical requirements but missed business outcomes: a class of problem that pure metrics cannot detect. An ART that delivers all its features on time but fails to satisfy Business Owner expectations has a requirements-validation problem, not a delivery-speed problem, and the demo is where this signal first appears.


Quantitative Metrics Review: Letting Data Drive Improvement

The quantitative metrics review runs as the second phase of I&A, anchoring improvement discussions in six SAFe Flow Metrics, flow distribution, velocity, load, time, efficiency, and predictability, so that the problem-solving workshop targets evidenced bottlenecks rather than the loudest complaint in the room SAFe Flow Metrics (Jira Align). SAFe 6.0 positions this review before the qualitative discussion deliberately: data establishes what is objectively true about the PI, and only after that foundation is laid does the ART move to interpretation and action.

PI predictability: measuring planned vs actual business value

The ART Predictability Measure is calculated by summing the actual business value achieved against each PI Objective and dividing by the total planned business value, expressed as a percentage. An ART that planned objectives totalling 100 business value points and achieved 83 points has an 83% predictability score. SAFe guidance identifies 80–100% as the target range for reliable ARTs; consistent scores above 80% indicate that the ART’s planning process is accurate and its execution is consistent. Scores below 80% trigger a specific line of inquiry during the metrics review: is the ART overcommitting at PI Planning, or is execution failing against realistic commitments? The distinction determines whether the corrective action targets planning discipline or execution capability. Trend matters more than any single data point: an ART scoring 75% for three consecutive PIs faces a different problem from an ART that dropped from 88% to 75% in a single PI, possibly due to a specific external dependency or capacity disruption.

Program performance metrics: flow distribution, velocity, time

Beyond predictability, the metrics review examines the SAFe Flow Metrics that reveal how the ART is performing as a delivery system. Flow Distribution tracks the proportion of work across Feature, Enabler, Maintenance, and Debt categories: an ART where Features consume more than 60% of capacity across multiple PIs is underinvesting in the architectural runway and will accumulate technical debt that eventually slows delivery. Flow Velocity measures the number of story points or features delivered per PI, trended across PIs to distinguish genuine throughput improvement from scope compression. Flow Time tracks the elapsed time from work item acceptance to delivery completion; rising flow time coupled with stable or rising velocity suggests that work items are spending more time in queues despite being completed faster once started, pointing to a WIP management problem. ARTs operating with Flow Efficiency below 15% face cascading delivery problems: the majority of elapsed time is wait state, not active work, and the I&A workshop must address WIP limits and handoff policies before any other improvement target Flow Efficiency (Flow Metrics framework).

Quality trends and their role in the metrics review

Quality trends override velocity signals. An ART that increases its velocity by 20% while defect escape rates rise from 4% to 17% is not improving: it is accelerating the production of defective work, and the velocity gains will reverse in subsequent PIs when rework consumes capacity.

Defect escape rates

Defect escape rate measures the percentage of production defects discovered after a feature was accepted as complete by the team. An escape rate below 5% is typical for ARTs with mature testing practices; rates above 10% indicate that team-level quality gates are insufficient. When defect escape rates spike in a single PI, the I&A workshop should investigate whether the root cause is compressed testing time due to overcommitment, a change in team composition that diluted testing expertise, or a systemic gap in automated test coverage. The trend across PIs is the signal to watch: a single bad PI can be an anomaly, but two or more PIs with rising escape rates indicate a structural quality problem that no amount of testing heroics will fix.

Test automation coverage trends

Test automation coverage tracks the percentage of regression tests that are automated. The target for mature ARTs is 70–80% unit test coverage combined with 40–60% integration test automation. An ART whose automation coverage is declining across PIs is accumulating manual regression debt: each PI adds features that must be manually retested, consuming more capacity and slowing release cycles. The metrics review should flag coverage drops of more than 10 percentage points between PIs for workshop investigation, particularly when combined with rising defect escape rates or lengthening regression cycles.

Technical debt accumulation metrics

Technical debt is measured through non-feature work proportion: the percentage of ART capacity consumed by refactoring, architecture upgrades, and defect fixes. ARTs that spend less than 20% of capacity on Enabler work across multiple PIs are consuming their architectural runway faster than they replenish it. The lag effect of technical debt means that the consequences do not appear in the same PI where the debt was incurred: an ART that underinvests in Enablers for three consecutive PIs will face a quality or performance crisis in PI 4 or 5, and the metrics review is where the leading indicators of that crisis first appear.

RTE pre-I&A dashboard preparation checklist

The quality of the metrics review depends almost entirely on dashboard readiness before I&A day. The RTE should distribute the metrics dashboard to all participants 48 hours before the event: this is the single highest-leverage preparation action available. The dashboard must include the ART Predictability Measure for the current PI trended across the last four PIs, Flow Distribution showing each category’s proportion, Flow Time and Efficiency with current-period and prior-period comparisons, defect escape rate trended across PIs, and test automation coverage percentages. The 48-hour distribution window gives Business Owners and stakeholders time to review the data before the event, so the metrics review phase focuses on interpretation and outlier identification rather than data discovery. ARTs using Jira Align or similar SAFe tooling can automate dashboard generation, but the RTE must still validate the data for accuracy; automated dashboards frequently double-count or miscategorise work items when team-level conventions for story tagging are inconsistent.


Facilitating the Problem-Solving Workshop

The Problem-Solving Workshop is the third phase of I&A: a structured two-hour root-cause analysis session following a six-step format that prevents the loudest voice from dominating and ensures the ART produces an actionable Improvement Backlog rather than a list of complaints Improvement Backlog (Agile Seekers). This phase consumes approximately half of the total I&A time budget, and its output is the sole deliverable that justifies the event’s overhead; without improvement stories that reach the Program Backlog, the demo and metrics review generated data but produced no change.

The facilitator role: RTE vs SPC; who leads?

The Release Train Engineer is the default facilitator for the Problem-Solving Workshop. The RTE knows the ART’s operational details, the metrics trends, and the personality dynamics in the room; all of which matter for keeping the session productive. The critical constraint is neutrality: the facilitator cannot be the person who owns the problem being solved or the person whose team is most directly implicated in the root cause. When the biggest problem identified during dot voting involves the RTE’s own facilitation patterns or the RTE’s team, a Scrum Master or SAFe Program Consultant must step into the facilitator role for that session. The SPC typically facilitates the first two or three I&A events for a newly formed ART, then hands off to the RTE once the six-step format is established as a team norm. Regardless of who facilitates, the core responsibilities are identical: enforce the time-box for each workshop segment, stay neutral on which root cause is correct, and actively push the room past its first plausible answer toward genuine root-cause depth (Agility at Scale). The single most common facilitation error is picking a root cause before the team does: a facilitator who telegraphs their preferred conclusion destroys the collaborative analysis that makes the workshop effective.

The six-step problem-solving format in detail

The Problem-Solving Workshop follows a structured six-step format built around rigorous timeboxing with specific facilitation prompts per step. Each step depends on the previous one, and skipping or compressing any step undermines the quality of the output.

Step 1: Agree on the problem statement

Timebox: 10 minutes. After the metrics review, the facilitator asks each participant to write down the single biggest problem they saw in the PI data. Each person writes their problem on a sticky note: one problem per note. All notes are placed on a shared board, and the group uses dot voting (three votes per person) to select the single problem the workshop will address. The facilitator then writes the selected problem as a specific, measurable statement: “The ART’s defect escape rate rose from 4% to 17% in PI 8, and the majority of escaped defects were integration-related” rather than “quality is getting worse.” A well-formed problem statement names the metric, the direction of change, and the scope.

Step 2: Root cause analysis tools

Timebox: 25 minutes. With the problem statement on the board, the facilitator introduces the root cause analysis tool best suited to the problem’s complexity. For multi-factor problems, which most ART-level problems are, the Fishbone (Ishikawa) Diagram is the recommended starting point. Teams split into breakout groups of 5–8 people, each group cross-functional by role and team affiliation, and begin mapping causes onto the four standard Fishbone categories: People, Process, Technology, and Environment (Agilephoria). The facilitator rotates between groups to ensure each group is identifying causes, not symptoms.

Step 3: Identify the biggest root cause

Timebox: 15 minutes. Each breakout group presents its Fishbone diagram to the full room. The facilitator consolidates the diagrams into a single board and asks for a second round of dot voting: each person votes for the root cause they believe has the highest impact on the problem. The root cause with the most votes becomes the workshop’s target. This step prevents the common failure pattern where the workshop tries to address all possible causes simultaneously and achieves none of them.

Step 4: Brainstorm solutions

Timebox: 20 minutes. With the target root cause identified, the room brainstorms countermeasures; concrete actions that would eliminate or mitigate that root cause. The facilitator enforces two rules during brainstorming: no criticism of any idea during the generation phase, and every proposed solution must be specific enough that someone could implement it after the workshop. “Improve testing” is not a solution. “Add automated regression tests for the payment integration module before PI 9 starts” is a solution.

Step 5: Create improvement action items

Timebox: 15 minutes. Each brainstormed solution is evaluated against two criteria: business impact and feasibility. The group votes on the top two solutions, and the facilitator converts each into an Improvement Story with a defined problem statement, proposed solution, success criteria, owner, and target PI. This step bridges the gap between workshop discussion and actionable Program Backlog items: an improvement that remains as a note on a shared board after the workshop has negligible probability of being implemented.

Step 6: Assign ownership and timelines

Timebox: 5 minutes. Each Improvement Story receives a named owner, a specific person, not a team or role, and a target PI for completion. The owner does not need to be the person who does all the work; the owner is accountable for ensuring the improvement is completed and demonstrated at the next I&A. The RTE records all Improvement Stories in the Program Backlog before the next PI Planning session begins.

Stakeholder management: keeping Business Owners engaged productively

Business Owners bring a valuable perspective to the Problem-Solving Workshop, they see the business outcomes that technical teams do not, but their presence also creates a power dynamic that can suppress honest discussion. Teams may hesitate to name problems that implicate strategic decisions or funding allocations when the person who made those decisions is in the room. The facilitator mitigates this by framing the workshop’s purpose explicitly: the goal is to find systemic causes, not assign blame. Practical techniques include asking Business Owners to speak last in each discussion round so their views do not anchor the conversation, positioning them in Fishbone breakout groups with facilitation training so they contribute as analysts rather than decision-makers, and explicitly inviting them to name problems that teams cannot see; such as misalignment between funded features and actual customer needs. An ART where Business Owners dominate every I&A workshop will systematically under-identify problems that originate at portfolio level.

From workshop output to Improvement Backlog

The immediate output of the Problem-Solving Workshop is a small set of Improvement Stories; typically two to five per I&A event. Each story follows a standard format: a problem statement referencing the specific metric or observation that triggered the investigation, a proposed solution describing the countermeasure, measurable success criteria that will determine whether the improvement worked, a named owner, and a target PI for completion. The RTE’s responsibility after the workshop is to ensure these stories are entered into the Program Backlog before PI Planning begins; if they are added during planning, they can be capacity-allocated alongside feature work. Improvement Stories entered after PI Planning has set its scope will almost certainly be deprioritised as “unplanned work” and deferred to the next PI, which means the I&A workshop’s output takes a full PI to begin producing value.


Root Cause Analysis Methods for the I&A Workshop

The I&A workshop employs three primary root cause analysis methods, Five Whys for linear causal chains, Fishbone (Ishikawa) Diagrams for multi-factor problems, and Pareto Analysis for frequency-based prioritisation, each suited to a distinct class of ART-level issue Pareto Analysis (Planview). The SAFe I&A Facilitator’s Guide recommends starting with Fishbone for most ART problems, then drilling into the largest branch with Five Whys to reach actionable depth.

Five Whys: when linear causal chains suffice

Five Whys is effective when the problem follows a traceable cause-and-effect chain without multiple independent contributing factors. The technique begins with the problem statement and asks “why” five times, each answer forming the next question in the chain, until the team reaches a process root cause that can be addressed with a specific countermeasure. The technique’s simplicity is also its limitation: it assumes a single causal path, which most ART-level problems do not follow.

Five Whys facilitation script for the I&A workshop

The facilitator writes the problem statement on the left side of a whiteboard or Miro board, then asks the full group: “Why did this happen?” The first answer goes next to the problem. The facilitator then asks: “And why does that exist?”: not “why did that happen,” which implies a one-time event, but “why does that exist,” which asks about systemic conditions. The facilitator repeats this for each successive answer, a maximum of five iterations, until the team reaches a response that names a process, policy, or system condition that can be changed. A common facilitation pitfall is accepting answers that name people rather than conditions: “the tester missed it” is not a root cause, but “the test environment does not include production-scale data” is a root cause that Five Whys can reach if the facilitator pushes past the first plausible answer.

Worked example: tracing a quality defect to its root cause

Problem: “The ART’s defect escape rate rose from 4% to 17% in PI 8.”

Why 1: Integration defects were not caught before release. Why? → Why 2: The integration test suite did not cover the new payment module. Why? → Why 3: The integration test environment was not updated to include the payment API. Why? → Why 4: Environment provisioning requires a two-week lead time, and the payment module API was finalised three days before PI end. Why? → Why 5: The ART has no policy requiring API contracts to be agreed by PI midpoint to allow environment preparation. The root cause is a missing architectural governance policy, not a testing failure; and the countermeasure is a contract-early policy, not more testers.

Fishbone (Ishikawa) Diagram: multi-factor root cause mapping

The Fishbone Diagram is the primary root cause tool for I&A workshops because most ART-level problems have multiple contributing factors. The diagram structures causes across categories to ensure the team considers all possible source domains before converging on a single causal path.

Standard I&A Fishbone categories: People, Process, Technology, Environment

The SAFe I&A Facilitator’s Guide recommends four standard categories. People covers skills, training, communication, and team composition; does the team have the right expertise, and do they have the information they need? Process covers policies, procedures, workflows, and governance; is the way work flows creating bottlenecks? Technology covers tools, infrastructure, automation, and data; do technical constraints prevent teams from delivering quality work? Environment covers organisational structure, culture, physical workspace, and external dependencies; are factors outside the ART’s control creating friction? Teams generating causes under each category produces a much richer causal map than unstructured brainstorming, which tends to cluster causes in whatever domain the loudest participant is most familiar with.

Worked example: mapping a PI predictability miss

Problem: “ART predictability dropped from 85% to 62% in PI 8.”

People: Three senior developers were reassigned to a different ART mid-PI; two replacement developers had a one-week onboarding ramp. Process: The dependency management board was not updated after the reassignments occurred, so downstream teams did not learn of the delay until the midpoint checkpoint. Technology: The CI pipeline failed for six days during the PI because the build server certificate expired and no one monitored certificate expiry. Environment: The PMO mandated a new compliance reporting tool mid-PI without adjusting ART capacity, consuming approximately 40 developer-hours that had been allocated to feature delivery. The Fishbone reveals that no single factor caused the predictability miss, four independent factors converged, and the countermeasures must address all four to prevent recurrence.

Pareto Analysis: frequency-based prioritisation

Pareto Analysis is used within the I&A workshop when the ART has multiple candidate problems and needs an objective method to select which one to address. The principle is that roughly 80% of the impact comes from roughly 20% of the causes; if the ART can identify that 20%, the workshop time is spent on the problems that matter most.

Building the Pareto chart from workshop data

The facilitator lists all problems identified during the quantitative metrics review and the PI System Demo feedback. Each participant assigns a relative weight to each problem based on its impact on ART performance; using a simple three-point scale (high = 3, medium = 2, low = 1). The facilitator sums the scores and plots them as a cumulative frequency bar chart. The problems that cumulatively account for 80% of the impact score become the workshop’s priority list. The technique is particularly useful when the ART is early in its I&A maturity and the room is overwhelmed by the number of issues surfaced in the metrics review; Pareto Analysis imposes an objective ranking that prevents subjective recency bias from dictating the workshop focus.

The 80/20 rule applied to ART improvement

ARTs that have run I&A for multiple PIs often discover that the same 20% of problem categories account for 80% of their predictability misses. Dependency management, incomplete acceptance criteria, and environment availability are recurring themes across most ARTs: the specific symptoms change each PI, but the underlying categories are stable. Pareto Analysis across PIs reveals these pattern categories, which is the signal to shift from reactive problem-solving to systemic improvement. An ART that identifies dependency management as its top problem category for three consecutive PIs has moved beyond “what caused this PI’s miss” to “what is the structural problem that produces dependency failures every PI”: a much higher-leverage improvement target.

Combining methods: Fishbone → Five Whys drill-down

The SAFe I&A Facilitator’s Guide recommends a specific method combination: start with a Fishbone to map all causal dimensions of the problem, then drill into the largest branch of the Fishbone using Five Whys to reach actionable depth. The Fishbone ensures the team does not prematurely converge on a single cause; the Five Whys ensures they do not stop at a surface-level cause that sounds plausible but cannot be addressed. An ART that used only Fishbone would produce a comprehensive map of causes but no prioritised action; every branch would look equally important. An ART that used only Five Whys on a multi-factor problem like a predictability miss would trace one narrow causal chain while ignoring the three other factors that contributed equally. The combination forces the team to be comprehensive before they are specific, which produces root causes that are both complete and actionable. This method also guards against the Correlation vs Causation trap: an ART that identifies “velocity dropped” as the problem may spend its workshop addressing velocity when the velocity drop was a symptom of dependency churn: the Fishbone surfaces the dependency issue as a causal branch, while a correlation-only approach would have treated the velocity drop as the root cause.


The I&A Workshop Agenda: A Timeboxed Facilitation Plan

A standard I&A workshop runs approximately three to four hours across three sequential phases, PI System Demo (60–90 minutes), Quantitative and Qualitative Metrics Review (30–45 minutes), and Problem-Solving Workshop (90–120 minutes), with each phase supplying the factual foundation for the next Problem-Solving Workshop (Agile Seekers). The total duration depends on ART size, maturity, and the complexity of the problems surfaced in the metrics review.

Full annotated agenda: timeboxes, facilitator notes, participant roles

TimeSegmentDurationFacilitatorParticipant RoleKey Output
0:00–0:10Welcome & Overview10 minRTEListen, review agendaAlignment on purpose and agenda
0:10–1:10PI System Demo60 minTeamsPresent, observe, evaluateScored PI Objectives; stakeholder feedback
1:10–1:15Transition5 minRTEShift context to data review,
1:15–1:45Metrics Review30 minRTEReview dashboard, identify outliersPrioritised problem candidates
1:45–2:00Break15 min,,,
2:00–2:10Problem Selection10 minRTEDot-vote on top problemSingle agreed problem statement
2:10–2:40Root Cause Analysis30 minFacilitatorFishbone breakout groupsMapped root causes
2:40–3:00Root Cause Voting & Solution Brainstorm20 minFacilitatorVote on root cause; brainstorm solutionsPrioritised root cause; solution ideas
3:00–3:25Action Items & Ownership25 minRTEDefine Improvement Stories, assign ownersDraft Improvement Backlog
3:25–3:30Close & Next Steps5 minRTEConfirm ownership, schedule follow-upCommitted Improvement Stories

The agenda is a template, not a straightjacket. Smaller ARTs (fewer than 50 people) can compress the System Demo to 45 minutes and still produce quality data for the metrics review. ARTs with distributed teams across time zones may need to schedule the three phases across two separate half-day sessions, with the Problem-Solving Workshop on a different day from the Demo and Metrics Review. The non-negotiable constraint is sequence: never run the workshop before the demo and metrics, because the workshop needs both the objective evidence and the quantitative signals to produce quality root cause analysis.

RTE pre-work: the 48-hour metrics dashboard distribution

The RTE’s pre-work determines whether the I&A runs smoothly or wastes the first 20 minutes on data discovery. Forty-eight hours before the event, the RTE distributes the metrics dashboard to all I&A participants via email or the ART’s collaboration platform. The dashboard must include: ART Predictability Measure for the current and prior three PIs, Flow Distribution (Feature/Enabler/Maintenance/Debt ratios), Flow Time and Efficiency with trend lines, defect escape rate by team and by feature area, and test automation coverage percentage. Participants receive a one-page guide showing how to read each metric and a prompt to come to I&A with their top three concerns ready. The 48-hour window is not arbitrary: it gives Business Owners and stakeholders time to correlate the quantitative data with their own observations before the workshop, so the Metrics Review phase focuses on interpretation rather than explanation. ARTs whose RTEs skip this pre-work consistently produce weaker problem statements during the workshop, because participants arrive without having reflected on what the data means.

Virtual I&A: adapting the workshop for distributed ARTs

Distributed ARTs face specific challenges that in-person I&A events do not: reduced nonverbal feedback during breakout sessions, difficulty maintaining engagement across time zones, and technology overhead that eats into workshop time. Three adaptations address these challenges directly.

Tooling: Miro/Mural boards for Fishbone collaboration

Pre-built Fishbone Diagram templates in Miro or Mural allow remote breakout teams to collaborate on root cause analysis with the same structure they would use on a physical whiteboard. Each breakout group gets a dedicated board with the four standard categories (People, Process, Technology, Environment) pre-labelled and the problem statement at the head. The facilitator monitors all breakout boards simultaneously, dropping into groups that appear stuck (no causes added in two minutes) or drifting (causes that describe symptoms rather than root causes). Virtual boards have the advantage of persistence: the completed Fishbone from PI 8 is available for reference during PI 9’s I&A, which helps the ART identify whether the same root cause categories recur.

The tech host role: separate from the facilitator

Every virtual I&A needs a dedicated tech host whose sole responsibility is managing the collaboration tools, monitoring chat for technical issues, and handling breakout room assignments. The tech host is never the facilitator or a participant in the workshop discussions. When the facilitator also manages tooling, admitting late participants, resetting breakout rooms, troubleshooting Miro permissions, the workshop loses 10–15 minutes to friction and the facilitator’s attention fractures at the exact moment it needs to be on the room dynamics. The tech host role is the single highest-impact investment a distributed ART can make in I&A quality.

Breakout room facilitation for root cause teams

Virtual breakout rooms require more structured facilitation than in-person Fishbone groups. Each breakout room receives a written brief containing the problem statement, the Fishbone template link, and a 20-minute timer with five-minute warnings. The facilitator designates a note-taker in each room before the breakout begins; in virtual settings, this prevents the “who types?” delay that loses the first 3–5 minutes of each breakout session. Rooms should be capped at eight participants; larger groups in virtual settings produce uneven participation because the video grid limits peripheral awareness. Rooms that finish their Fishbone early are instructed to identify the single largest branch and prepare a one-minute summary rather than defaulting to idle discussion.

I&A logistics: room setup, materials, tooling checklist

For in-person I&A events, the room setup determines workshop quality. The room needs: a large shared display for the PI System Demo (projector or large monitor visible to all participants), a separate wall or whiteboard for the Fishbone Diagram (or a flip chart per breakout group), sticky notes in three colours (one colour each for problems, root causes, and solutions), dot-voting stickers, and printed copies of the metrics dashboard for participants who prefer paper over screen. For virtual events, the checklist includes: a Miro or Mural board with pre-built Fishbone template and dot-voting widget, breakout rooms pre-configured in Zoom, Teams, or the chosen platform, a separate tech host with host privileges for breakout room management, and a shared-screen-friendly metrics dashboard (optimised for 16:9 display, minimum 14pt font for remote readability). The RTE validates every item on the tooling checklist 24 hours before the event: a failed projector or a Miro board that participants cannot edit is not a minor inconvenience; it is a structural barrier to the workshop achieving its purpose. The core facilitation techniques for virtual I&A, timebox enforcement, breakout room management, and silent voting, are the same as in-person, but the delivery requires more explicit structure and a dedicated tech host to handle tooling overhead.


From Insights to Action: Building the Improvement Backlog

The Improvement Backlog is the tangible output that justifies the entire I&A event: it transforms workshop insights into structured improvement stories with defined problem statements, proposed solutions, success criteria, owners, and target PIs that compete for capacity against feature work at PI Planning PI Planning (Agile Seekers). Without a functioning Improvement Backlog, I&A becomes a quarterly exercise in data analysis that produces no behavioural change.

The improvement story format: problem, solution, criteria, owner, target PI

Every Improvement Story follows a five-element format designed to survive the transition from workshop whiteboard to Program Backlog without losing its specificity. The problem statement names the metric or observation that triggered the story, “Defect escape rate rose from 4% to 17% in PI 8, with the majority of escaped defects being integration-related.” The proposed solution describes the countermeasure with enough detail for a team to implement it, “Implement API contract validation in the CI pipeline so that integration tests fail automatically when contract mismatches are detected.” The success criteria define how the ART will know the improvement worked, “Defect escape rate returns to 5% or below within two PIs of implementation.” The owner is a single named individual accountable for completion, not a team, not “the ART,” not “whoever has capacity.” The target PI specifies which PI the improvement is planned for completion. Stories that omit any of these five elements, particularly the success criteria or the named owner, have a dramatically lower completion rate because they lack the structure needed to survive the prioritisation that happens at PI Planning.

Prioritising improvement work against feature delivery

The single biggest failure mode in the I&A lifecycle is improvement stories that never get worked on because feature delivery feels more urgent. Improvement Stories entered into the Program Backlog after PI Planning has locked its scope face a near-zero probability of being allocated capacity; they become “stretch goals” that get dropped at the first sign of schedule pressure. The mechanism that prevents this is explicit capacity allocation at PI Planning: the ART agrees during planning to reserve a fixed percentage of capacity, typically 10–20%, for Improvement Stories from the previous I&A. ARTs that do not reserve capacity for improvement systematically starve their Improvement Backlog, and the same problems recur PI after PI while the Improvement Backlog grows as a graveyard of good intentions. WSJF or business-value voting can rank improvement stories against features when capacity contention arises, but the ranking mechanism is secondary to the capacity allocation decision; without reserved capacity, even the highest-ranked improvement story will lose to any feature delivery commitment.

RTE accountability: tracking improvement story completion rates

The improvement story completion rate, the percentage of Improvement Stories from each I&A that reach their Defined of Done within their target PI, is the single most important metric for whether I&A is working. An ART with a completion rate below 60% has an I&A that produces data and discussion but zero organisational change The RTE tracks this metric across PIs and reports it at the start of each metrics review phase. ARTs operating in the reactive stage of I&A maturity typically see completion rates below 40%: the workshop identifies valid problems, but the organisational system lacks the discipline to act on them. ARTs in the data-driven stage see rates of 60–80%; improvements are prioritised, allocated capacity, and completed. The RTE’s accountability mechanism is simple but uncomfortable: if improvement story completion rates remain below 60% for three consecutive PIs, the RTE must surface this during the I&A metrics review as itself a problem requiring a workshop, because the ART is spending four hours every quarter generating improvements it never implements.

Closing the loop: how this I&A’s backlog feeds the next PI’s planning

The Improvement Backlog from one I&A feeds directly into the next PI Planning session. Before PI Planning begins, the RTE reviews the Improvement Backlog and identifies which stories are proposed for capacity allocation in the upcoming PI. Each story is written on a card with its problem statement, proposed solution, and capacity estimate, inserted into the planning materials alongside feature cards. During PI Planning, the ART assesses improvement stories against the same capacity constraints as feature work and votes on which to include. This integration serves two purposes: it ensures improvement work is visible and committed rather than orphaned in a separate tracking system, and it forces the ART to make explicit trade-offs between feature delivery and improvement velocity. An ART that cannot find capacity for any improvement stories in a given PI is sending a signal that the PI was planned beyond available capacity; and that signal, captured during planning, becomes data for the next I&A.


Measuring I&A Effectiveness Over Multiple Program Increments

I&A effectiveness is measurable across three leading indicators, improvement story completion rate, PI predictability trend, and defect escape rate trajectory, tracked over a minimum of three to four PIs to establish statistically meaningful trend lines rather than reacting to single-PI noise. SAFe 6.0’s Continuous Learning Culture competency frames I&A as the operational mechanism for organisational learning, but that framing is aspirational unless the ART can demonstrate that its I&A practice is actually improving ART performance across time.

Leading indicators: improvement completion rate, PI predictability, defect trends

Three metrics, tracked together, distinguish effective I&A from performative I&A. Improvement story completion rate measures whether the ART acts on what it learns: a completion rate below 60% across two or more PIs means the ART is generating insights it does not implement. PI predictability trend measures whether the ART’s planning accuracy is improving: a predictability score trending from 65% toward 85% across PIs is the strongest signal that I&A-driven improvements are creating real reliability gains. Defect escape rate trajectory measures whether quality is improving: a declining trend signals that root cause analysis is correctly targeting quality problems. No single indicator is sufficient in isolation: an ART with high improvement completion rates but flat predictability is completing its improvement stories but targeting the wrong problems, and an ART with improving predictability but rising defect escape rates is sacrificing quality for reliability.

The I&A maturity model: reactive to data-driven to predictive

I&A maturity progresses through three stages, each defined by specific metric thresholds and behavioural indicators. The progression is not automatic: it requires intentional investment in facilitation capability, metrics infrastructure, and the organisational norms that make improvement work visible and valued.

Reactive stage (PI 1-2): ad-hoc workshops, low completion rates

In the reactive stage, I&A is a “quarterly retrospective”: the ART holds the event because SAFe says to, but the workshop lacks the data discipline and facilitation quality that drive action. Predictability is typically below 60%, improvement story completion rate is below 40%, and the same root causes appear in consecutive I&A workshops because no improvement stories were completed in between. The workshop is driven by participant sentiment rather than metric outliers: the problems selected are whatever the loudest participant thought was important, and the root cause analysis is superficial because the facilitator lacks experience with Fishbone or Five Whys. ARTs at this stage should focus on building the pre-requisites for effective I&A: RTE facilitation training, automated metrics dashboards, and the expectation that improvement stories will be allocated capacity at PI Planning.

Data-driven stage (PI 3-4): metrics-anchored, systematic root cause analysis

In the data-driven stage, the metrics review is the primary source of problem candidates, and the workshop selects the highest-impact issue through dot voting rather than loudest-voice. Predictability trends toward 70–85%, improvement story completion rates reach 60–80%, and the ART has completed enough improvement stories that the metrics review can distinguish between a one-off quality blip and a structural trend. The Fishbone and Five Whys techniques are used regularly workshop participants are comfortable with the format, and the facilitator can push past first-answer root causes without the room pushing back. ARTs at this stage are generating real improvement, the question shifts from “are we doing I&A?” to “are we solving the right problems?”, and the RTE begins tracking whether solved root causes from one PI recur in subsequent PIs.

Predictive stage (PI 5+): trend-based improvement, proactive defect prevention

In the predictive stage, the ART anticipates problems before they produce metric degradation and allocates improvement capacity proactively rather than reactively. Predictability is consistently above 85%, improvement story completion rates are above 80%, and the set of root causes addressed in I&A workshops shifts from recurring problems to emerging risks. The ART’s improvement trajectory is visible in trend lines across PIs: flow efficiency rises, defect escape rates decline, and the percentage of ART capacity consumed by unplanned rework drops below 10%. ARTs at this stage use I&A not to fix what broke in the last PI but to test hypotheses about what could break in the next one. The Workshop targets Enabler work, automation gaps, architectural runway, skill bottlenecks, before those gaps produce delivery failures. This is the continuous improvement cycle operating at its intended velocity, and the ART’s metric trends confirm that the cycle is working.

Continuous Learning Culture as the organisational enabler

Continuous Learning Culture is the SAFe competency that describes the organisational conditions under which I&A produces sustained improvement rather than episodic fixes. SAFe 6.0 positions it as one of the seven core competencies, and its relationship to I&A is straightforward: I&A is the event, Continuous Learning Culture is the system that makes the event’s output durable. An ART operating in an organisation that punishes failure, rewards heroics, or treats improvement stories as optional will produce an I&A that follows the format but never the intent. The maturity model progression from reactive to predictive is impossible without the organisational norms that Continuous Learning Culture formalises: psychological safety to name problems without blame, leadership commitment to allocate capacity for improvement, and the transparency practices that make ART-level metrics visible to every role. ARTs that measure their I&A maturity should also assess their Continuous Learning Culture maturity, because the two trajectories are linked: a data-driven I&A that operates in a blame-oriented culture will eventually revert to reactive patterns when the improvement stories implicate leadership decisions.

Detecting I&A stagnation: the performative-I&A checklist

The strongest signal that I&A has become performative is the same root cause appearing in three or more consecutive I&A workshops. If “dependency management” was identified as the primary root cause in PI 6, PI 7, and PI 8, and the ART’s predictability has not improved, then I&A is generating diagnoses but not treatments: the workshop identifies the problem correctly every quarter but never implements the countermeasure. Additional stagnation signals include: declining workshop attendance (busy participants stop attending when I&A produces no change), improvement stories that remain in the backlog for more than two PIs without being started, and workshop participants who disengage during the problem-solving phase because they have seen the same Fishbone exercise produce the same output for the last three quarters. When these signals appear, the RTE must escalate the stagnation as itself an I&A-worthy problem: the ART spends four hours every quarter on an event that produces no behavioural change, which is a waste of ART capacity that the metrics review should surface as a metric outlier.

Multi-PI trend tracking: why one data point is noise

Trend analysis across multiple PIs is the only reliable method for distinguishing systemic degradation from one-off disruptions. Single-PI data can mislead: an ART whose predictability drops from 88% to 62% in one PI may have experienced a reorganisation, a compliance deadline, or a key dependency failure; none of which signal a systemic problem. The same ART whose predictability trends from 88% → 82% → 74% → 62% across four PIs has a systemic degradation that demands root cause analysis. The SAFe guidance is explicit: trend lines across three to four PIs are the minimum viable signal for distinguishing between noise and pattern. ARTs that make major decisions based on single-PI data over-correct; they allocate improvement capacity to problems that would have resolved themselves and ignore slow-burn degradation that only becomes visible in the trend. The RTE should prepare the metrics dashboard with four-PI trend lines by default and flag single-PI outliers as candidates for investigation, not action.


Summary

Inspect and Adapt is the PI-level mechanism that prevents the ART from repeating the same dysfunctions quarter after quarter by coupling objective evidence from the System Demo with quantitative metrics to drive structured root cause analysis and an actionable Improvement Backlog. Its power comes from the dependency chain between its three phases: the System Demo supplies objective evidence, the metrics review supplies quantitative signals, and the problem-solving workshop transforms both into structured action. Without that chain, I&A produces discussion without direction.

Improvement is a Mechanism, Not a Meeting

The most important shift that an ART can make in its I&A practice is treating improvement as a structured delivery system rather than a quarterly discussion forum. The mechanism is straightforward: the Improvement Backlog from one PI feeds the capacity allocation in the next PI, the allocated improvement stories are completed during the PI, and the next I&A measures whether the improvement produced measurable change. An ART whose improvement completion rate is below 60% does not have an I&A problem: it has an organisational learning problem, because the ART spends four hours every quarter generating improvement ideas that the system cannot absorb. The RTE’s metric dashboard, tracked across PIs, exposes this gap. The intervention is not better workshops or more time spent in I&A; the intervention is capacity allocation discipline at PI Planning, where improvement stories compete for time against feature work and the ART explicitly decides how much of its capacity goes to getting better at delivering versus delivering the next feature.

The Root-Cause Habit Separates Mature ARTs from Performative Ones

ARTs that have progressed from reactive to data-driven to predictive I&A share one common behaviour: they do not stop at the first plausible root cause. The difference between an ART that identifies “we need better testing” as its biggest problem in three consecutive I&As and an ART that identifies “our API contracts are finalised too late for environment provisioning” is the depth of its root cause analysis practice. The Fishbone → Five Whys combination is not a facilitation technique; it is a discipline that forces the room to move past symptoms to structural causes. ARTs that invest in this discipline, training facilitators, building the metrics infrastructure that feeds accurate data into the workshop, and protecting improvement capacity at PI Planning, produce measurable improvement trajectories that compound across PIs. ARTs that run I&A as a 90-minute retro with flow metrics have an event that feels productive but produces no measurable change. The metrics review at the start of each I&A tells the truth about which category an ART belongs to, and that truth is the starting point for the next cycle of improvement.

Privacy Preference Center