SAFe Principles
19 MIN READ

Principle-Practice Diagnostic: Symptoms of Principle Violations

A SAFe train can pass every practice audit and still violate its principles. This principle-practice diagnostic traces symptoms to the violated principle.

A SAFe train can pass every practice audit on the books and still violate half its principles. In fact, most can. The scorecard checks whether ceremonies happened, not whether Take an Economic View or Apply Systems Thinking ever shaped a real decision. A principle-practice diagnostic closes that gap: it traces the symptoms of principle violations back to the specific principle each one corrupts.


Why Compliant SAFe Practice Doesn’t Prove Principle Adherence

A team can execute PI Planning, Iteration Reviews, and System Demos exactly as SAFe prescribes and still be violating the principles those ceremonies exist to serve, because compliance measures whether an event occurred, not whether the economic and systems-thinking reasoning behind it drove the decision made inside it. The gap doesn’t surface on a ceremony checklist. It becomes visible only when someone asks what changed inside the room, not whether the room met.

The John Deere Case: Compliant Practice, Stagnant Outcomes

John Deere’s Global IT group ran into exactly this gap in 2019, and the case is documented in a 2022 Scrum Inc. case study. The group had already been running a scaled Agile framework for years, PSI planning, two-month release cycles, the full ceremony calendar, and senior leadership still wasn’t seeing the results it wanted. Josh Edgin, the group’s Transformation Lead, named the specific failure mode instead of gesturing at poor execution: “With PSI planning and two-month release cycles, I think you can get comfortable transforming into a mini-waterfall.” The cadence was present on every calendar invite, though the economic reasoning behind flow and fast feedback never traveled with it.

Ganesh Jayaram, John Deere’s Vice President of Global IT, chose Scrum@Scale as the framework underneath the group’s new Agile Operating Model. His stated reason was structural: the group needed the reasoning behind decisions to run through every team, and Scrum@Scale gave the Agile Operating Model a home for that reasoning independent of whichever ceremony calendar sat on top of it. The case reports a 165% output increase attributed to the principle-led Agile Operating Model rather than the ceremony calendar underneath it: one company’s reported outcome, worth reading as a directional signal about where the actual lever sat rather than as a controlled benchmark. A 2022 academic review of SAFe adoption reached a parallel conclusion from a different angle: organizations can run process-compliant while the reasoning behind the process was never adopted by the people running it, which is precisely the mini-waterfall pattern Edgin diagnosed in his own group before any outside researcher established it.

A Century of Principle Lists Without a Verification Step

Henri Fayol’s 1916 Administration Industrielle et Générale set out fourteen management principles, and two more fourteen-item lists, Deming’s and Liker’s, arrived generations later carrying the same count and the same gap: no attached way to check whether an organization was living the list or filing it. The Lean Enterprise Institute’s retrospective on the coincidence notes that the matching count between Fayol’s fourteen principles, Edwards Deming’s fourteen points, and Jeffrey Liker’s fourteen Toyota Way principles is coincidental, not a shared method. Three separate eras converged on managing complexity through roughly a dozen statements, and none of them shipped a built-in test for whether a given decision actually ran through the reasoning the list described.

Liker’s own subject, Toyota, is the counterexample that shows why the gap matters. A June 2008 Harvard Business Review feature on the company’s success argued that Toyota’s results come from its willingness to hold contradictions, stable and paranoid, systematic and experimental, formal and frank, rather than resolve them into a single checklist item. A checklist is exactly what a fourteen-point or ten-point principle list becomes without a step that forces someone to test whether the reasoning, not just the label, is present in a given decision. That gap is what the smallest version of a diagnostic instrument targets first: a test that runs on one step, not an entire program.


The Step-Level Test: Applying Womack and Jones’s Five Criteria as a Principle Check

A single step inside a value stream fails the Womack and Jones lean test when it isn’t simultaneously Valuable, Capable, Available, Adequate, and Flexible, and each criterion it fails points to a specific SAFe principle rather than a vague process complaint. Running this test costs one step and a few minutes: no retrospective, no program-wide audit, no waiting for the next Inspect and Adapt to find out something is wrong.

The Five Criteria: Valuable, Capable, Available, Adequate, Flexible

James Womack and Daniel Jones proposed a five-step lean thought process in 1996 through the Lean Enterprise Institute, and buried inside step two, mapping the value stream, is a five-criteria test that decides whether a given step belongs in the process at all. A step earns its place only if it clears all five: Valuable (the customer would object if the step disappeared), Capable (it produces the correct result every time, not most of the time), Available (it can run whenever the value stream needs it to), Adequate (it has the capacity to keep the flow continuous rather than throttling it), and Flexible (it permits a range of outcomes without a full re-tooling of the step itself). Womack and Jones later folded their original five steps into a simpler Purpose-Process-People frame in 2007, but the step-level test survived that simplification unchanged, because it operates one level below the framing; on the individual step, not the overall value stream.

Each of the five criteria maps to a distinct kind of breakdown. A step that fails Valuable is often tolerated because it always ran, not because anyone can say what it produces that the customer would pay for. A step that fails Available or Adequate usually shows up first as a queue, since work backs up in front of a step that can’t run when needed or can’t keep pace once it starts. Capable and Flexible are the two criteria a team is least likely to test for itself, because both require comparing a step’s output against a standard nobody wrote down; which is exactly why the next test applies all five criteria to one concrete unit of SAFe work instead of leaving the check implicit.

Applying the Test to a Single PI Planning Step

Take one step inside PI Planning, an Agile Release Train drafting Program Increment objectives during a breakout session, and run all five criteria against it before judging the breakout as a whole. Valuable and Available usually clear without much argument: objective-drafting produces something the Business Owners act on, and the breakout slot exists on the agenda whether or not the team is ready for it. Capable and Adequate are where trouble tends to surface: a breakout that produces inconsistent Program Increment objectives from one PI to the next is failing Capable, and a breakout that routinely runs out of time before every team finishes is failing Adequate, since the step lacks the capacity to keep the planning flow continuous across the room.

Flexible is the criterion that carries a specific principle attached to it. A PI Planning step that is Capable, it reliably produces a committed objective, but not Flexible, because the team locks the objective early and treats any later adjustment as a planning failure, is not a training gap. It is a violation of Assume Variability and Preserve Options, the SAFe principle built specifically to keep design and commitment options open until the evidence justifies closing them. Naming that mapping is the entire value of the test: a Scrum Master who only sees “the breakout ran long” has a scheduling complaint, while a Scrum Master who runs the five-criteria test against the same breakout has a specific principle to raise at the next Inspect and Adapt.

Mechanics Failure or Principle Failure: What the Result Tells You

Every failed criterion answers one of two different questions, and conflating them is the most common way a five-criteria test gets misread. The mechanics question asks whether the step was performed correctly on the occasion someone observed it; did the facilitator run the breakout well, did the team show up prepared. The principle question asks whether the step is structurally capable of producing the intended value at all, regardless of who ran it or how well. A step that fails once under a distracted facilitator is a mechanics problem; a step that fails the same criterion under three different facilitators across three different Program Increments is a principle problem, because the defect persists through a change of personnel.

This same confusion shows up outside SAFe research specifically. A 2014 systematic review of Agile-UX integration studies, published through Agile Alliance, examined how consistently teams that formally adopted user-centered design practices actually let the design reasoning behind those practices shape their decisions, rather than treating user research as a step to complete and file. The pattern the reviewers found, practices adopted as artifacts while the reasoning behind them stayed optional, is the same pattern Edgin diagnosed at John Deere and the same pattern the five-criteria test is built to catch at the level of a single PI Planning step. A team that keeps failing the same criterion after swapping out the people running the step has stopped debugging execution and started debugging a principle.


The Symptom Table: Observable Behaviors That Flag a Failing Principle

Two behaviors documented in existing retrospective and ceremony notes, a time-to-points conversion and a stretched Sprint boundary, are enough to flag a specific principle failure before any program-level dashboard would show a problem. Scanning for them doesn’t require a new ceremony; it requires reading the ceremony notes a train already keeps with a different question in mind.

Observable Behavior Where It Shows Up Principle It Corrupts What Confirms It’s Structural
“Four hours equal one Story Point” (hours-to-points conversion) Backlog estimation, Sprint planning Economic sequencing Recurs across more than one Sprint’s estimation session
“We extend the Sprint until we get it done” (variable Sprint length) Sprint boundary, Iteration close Apply cadence and synchronize Sprint length varies across more than one cycle
Clean velocity and predictability dashboard alongside retro notes recording both symptoms Program-level reporting vs. team-level retro notes Both of the above, hidden from the reporting layer Dashboard period and retro record cover the same weeks

Time-to-Points Conversion: the Economic-Sequencing Symptom

Converting hours directly into Story Points, the LeadingAgile anti-pattern captured in the line “Four hours equal one Story Point”, corrupts economic sequencing, the discipline that turns relative-size estimates into prioritization decisions. Story Points exist specifically to strip time out of estimation, because a team’s velocity is meant to reflect relative effort across stories, not a disguised hour count that happens to use a different unit. When a team quietly runs the conversion in reverse, treating four hours as always equal to one point regardless of the story’s actual complexity or risk, every subsequent Cost of Delay calculation built on that team’s velocity inherits a distortion nobody flagged.

The distortion compounds because economic sequencing decisions, which Epic gets Program Increment capacity, which Feature gets pulled forward, depend on velocity numbers that are supposed to reflect genuine relative sizing. A team running the hours-to-points conversion is feeding the Portfolio Kanban’s prioritization math a number that looks like an estimate but functions like a disguised time log, which means every economic-view decision made downstream of that team’s Story Point count is reasoning from a number that never represented what the principle assumes it represents.

Variable Sprint Length: the Cadence Symptom

Stretching a Sprint past its committed boundary, “we extend the Sprint until we get it done”, abandons Apply Cadence and Synchronize in favor of exactly the batch-and-queue behavior cadence was built to remove. Cadence exists to convert unpredictable work into a predictable rhythm: fixed Sprint length turns “how long will this take” into “how much fits in the time we already have,” which is what makes forecasting possible across an Agile Release Train in the first place. A Sprint that extends whenever the committed scope isn’t finished quietly reverses that conversion, letting scope dictate the calendar instead of the calendar constraining scope.

Once a train tolerates one stretched Sprint, batch-and-queue creeps back into planning by a different door: teams start padding estimates to avoid the appearance of a stretch, which reintroduces the large-batch thinking cadence exists to eliminate, and cross-team synchronization, the second half of the same principle, breaks down because other teams can no longer trust that a Sprint boundary means what it says. The symptom is a single sentence in a retro. The principle it corrupts governs every downstream forecast the train produces.

Why Self-Reported Compliance Missed Both

Self-reported adoption metrics, the survey a team fills out claiming it follows SAFe’s cadence and estimation practices, routinely miss both anti-patterns, which is exactly why a 2022 study on SAFe implementation shortcomings argued for instrumented, AI-assisted behavioral tracking rather than relying on teams to flag their own deviations. Teams running the hours-to-points conversion or a variable Sprint length rarely report either one as a deviation, because from inside the team both behaviors feel like reasonable adaptations to real constraints rather than violations of a named principle.

The clearest illustration is a train that reports strong velocity predictability on its program dashboard for a given quarter while its own retrospective notes, covering that same quarter, record both the hours-to-points conversion and a stretched Sprint at least once. The dashboard and the symptom coexist in the same weeks, because the dashboard measures an output, did the numbers look stable, while the retro notes record the input behavior that produced them. Reading the retro notes against the symptom table catches what the dashboard alone cannot, which is exactly the layer the next distinction makes explicit.


Performance Versus Health: Why Metrics Miss a Principle Violation

A train’s velocity and predictability numbers measure its performance layer, and a principle violation typically lives one layer beneath that, in what Scott Keller and Bill Schaninger call organizational health. The two layers move independently of each other, and a metric built to track one will not reliably surface a structural defect in the other.

Keller and Schaninger’s Performance-and-Health Model

Scott Keller and Bill Schaninger, in Beyond Performance 2.0 (McKinsey, 2019), draw a distinction between performance, what an organization delivers in a given period, and health, how sustainably it can keep delivering that output over time. Their five-frames model treats health as a separate, measurable dimension rather than a soft afterthought to performance numbers, built from research the authors describe as the most comprehensive effort of its kind into what allows large-scale change to actually stick. The distinction matters for a principle diagnostic specifically because a principle violation rarely shows up as a performance failure on its own; it shows up as a health defect that performance metrics were never built to detect.

Coca-Cola’s Turnaround: Metrics Without Health

Neville Isdell took over as CEO of Coca-Cola during a period the company’s total shareholder returns had fallen to minus 26%, against PepsiCo’s positive 46% over the same comparison period. Isdell had a clear enough diagnosis of what the company needed, stronger core brands, new categories, a repaired relationship with bottlers, but no strategic redirection landed until the company’s declining morale, deficient capabilities, and flagging performance culture were addressed first. The pattern Keller and Schaninger draw from the case is that Isdell’s predecessors had already identified similar strategic priorities and still failed to make the changes stick, because they were treating a health problem as though better strategy alone could fix it.

Applying the Split to a SAFe Train

Velocity and predictability are a SAFe train’s performance layer, the numbers a Release Train Engineer reports at the System Demo. The health layer sits underneath: whether decisions actually route through economic reasoning, whether the train tolerates variability instead of forcing early commitment, and whether decisions sit with the people closest to the work rather than escalating by default. That split maps directly onto Keller and Schaninger’s five frames: a train’s decision-routing frame asks whether an Epic’s Cost of Delay actually gets argued in economic terms before it’s prioritized, its variability frame asks whether a team is allowed to hold two design options open past the point a status-conscious train would force a call, and its ownership frame asks whether a Product Owner or Scrum Master can make a call the framework assigns them without escalating it upward by habit; three structural checks a Release Train Engineer can run against a single team without waiting for a Program Increment’s worth of velocity data to accumulate.

A 2020 empirical study of SAFe adoption success factors found that the factors correlating with successful adoption clustered around alignment and capability rather than raw velocity, which grounds the performance-health distinction specifically inside SAFe rather than leaving it as a general management observation borrowed from Keller and Schaninger’s broader sample. The diagnostic implication follows directly: a principle violation is a health-layer defect, and health-layer defects don’t reach a performance dashboard until they’ve already compounded across several cycles; which is the reason a repeatable sequence, not a one-off audit, is what closes the gap.


Running a Principle Diagnostic: A Four-Step Attribution Sequence

Running a principle diagnostic means working through four ordered steps, scan, test, verify, and attribute, that turn a symptom noticed in a retrospective into a named principle and a named decision-maker, without adding a single meeting to the calendar. SAFe 6.0 already provides the ceremony slot the sequence runs inside; the work is a method for that slot, not a new one layered on top of it.

The Four Steps, in Order

Step one scans recent ceremony notes against the symptom table for candidate anti-patterns, before reaching for a general improvement action that treats the symptom as unexplained. Step two runs the five-criteria step-level test against the specific value-stream step where the candidate symptom occurred, which established the step is structurally deficient rather than poorly executed on a single occasion. Step three checks the flagged step against the performance-health split: if the step’s metrics still look acceptable, that demonstrated a genuine principle violation, since a violation that had already degraded performance would have become visible through the dashboard on its own. Step four names the one violated principle and pairs it with the one accountable decision point, the Product Owner’s story-ready call, the Release Train Engineer’s Sprint-boundary call, or the Business Owner’s Cost-of-Delay call, named inside the specific ceremony where that decision actually gets made, rather than issuing a general improvement action that nobody owns.

The order matters more than any single step in isolation. Running step four before step two produces a plausible-sounding principle name attached to a symptom nobody has established is structural, which is exactly the kind of unverified attribution the sequence exists to prevent.

Where the Sequence Runs: Inspect and Adapt

The four-step attribution sequence runs inside SAFe’s own Inspect and Adapt event, led by the Release Train Engineer, rather than as a separate audit competing for calendar time. SAFe 6.0 positions Inspect and Adapt as the train’s dedicated space for exactly this kind of structured reflection, a Program Increment System Demo followed by quantitative and qualitative measurement, then a structured problem-solving workshop, which means the attribution sequence slots into a workshop segment that already exists rather than requesting a new one. The Release Train Engineer who runs Inspect and Adapt is also the person best positioned to pull ceremony notes from every team on the train, which step one of the sequence depends on.

Running the sequence inside an existing event also keeps it cheap enough to repeat every Program Increment rather than reserving it for a crisis. A diagnostic that requires scheduling a special session gets deferred; a diagnostic that occupies twenty minutes of a workshop segment that was already happening gets run.

Naming the Principle and the Accountable Decision Point

The sequence ends by naming one violated principle and one accountable decision point, never a general improvement action assigned to “the team”, the same precision that steps one through three exist to earn, since a symptom that has survived the scan, the five-criteria test, and the performance-health check has already been narrowed past the point where a team-level attribution would be honest. An accountable decision point is a specific moment inside a specific ceremony where a specific role makes the call: the Product Owner deciding whether a story is ready, the Release Train Engineer deciding whether a Sprint boundary holds, the Business Owner deciding whether an Epic’s Cost of Delay justifies its position in the Portfolio Kanban.

Naming the decision point rather than the team matters because a diagnosis addressed to a team produces the same vague accountability the sequence was built to replace. A diagnosis addressed to a decision point, the moment inside the Portfolio Kanban where a Business Owner has to justify an Epic’s Cost of Delay against its queue position, gives the next Inspect and Adapt something concrete to check for improvement, rather than a general sentiment to feel good about.

Diagnostic Thinking Is a Skill, Not a Byproduct of Experience

Harvard Business Review’s brief on diagnostic thinking treats symptom-to-cause attribution as a discipline practitioners build deliberately, not something years of ceremony experience produces automatically. That framing matters here because the Agile community’s own foundational documents assume the opposite. Agile Alliance’s 12 Principles Behind the Agile Manifesto lay out what good practice looks like but supply no companion method for checking whether a given team’s daily decisions actually trace back to those twelve statements: the checking is left to the practitioner’s judgment, the same gap Fayol, Deming, and Liker’s lists left open decades earlier.

SAFe’s own framework team has made a version of the same admission. A Scaled Agile blog post on product innovation describes SAFe’s principles, practices, roles, and tools as “dispersed throughout the Framework,” with the Big Picture showing how they “all work together”: a description of coherence, not a test for it. Treating attribution as a skill to build, the way HBR frames diagnostic thinking, is what turns that dispersed coherence into something a Release Train Engineer can actually verify one Program Increment at a time, rather than something the framework merely implies.


Summary

A principle-practice diagnostic works because it runs cheap instruments in a fixed order, scan, test, verify, attribute, rather than trying to audit every ceremony on a train at once, and it treats a straightforward performance dashboard as a starting question rather than a closing answer.

Run the Cheapest Instrument First, Not the Broadest

The five-criteria step test and the symptom table both cost less than a Program Increment-wide audit, because each one runs against a single step or a single retrospective note rather than the whole train’s ceremony calendar. That ordering isn’t incidental. A team that reaches for a full audit every time a symptom appears will run the audit rarely, because full audits are expensive to schedule and disruptive to run, which means most symptoms go unchecked between audits. A team that can run the five-criteria test against one PI Planning step in a few minutes, or scan a retro against the symptom table’s two named anti-patterns, will actually run the check often enough to catch a principle violation while it’s still cheap to fix.

The John Deere case demonstrates the same ordering principle at the organizational level: the signal was already sitting in ceremony notes the group kept for other reasons, cheap enough to scan before anyone reached for a company-wide audit. The four-step attribution sequence formalizes that same instinct, scan first, test the smallest confirmable unit second, so that the diagnostic scales down to a single retrospective instead of requiring a crisis to justify running it.

A Clean Dashboard Is Never the Answer to a Principle Question

Keller and Schaninger’s performance-health split is the reason a train’s velocity and predictability numbers can never resolve whether a principle is intact, because those numbers were built to measure the performance layer, and a principle violation lives in the health layer underneath it. What resolves a principle question instead is the four-step sequence’s own output: one named principle paired with one named accountable decision point, the two things a metrics reading was never built to produce no matter which direction it points.

The distinction between a mechanics failure and a principle failure, raised earlier against the five-criteria test, resolves the same way here at the train level. A metrics dip is a performance failure, visible and self-correcting through normal ceremony pressure. A principle violation with clear metrics is a health failure, invisible to the dashboard by construction, which is exactly why the attribution sequence checks the performance-health split before naming a violation; showing the metrics still look fine is what indicates the defect is structural rather than a performance problem the train would have caught on its own.

Morné Wiggins · Agility at Scale · Talk to me

Privacy Preference Center