Delivery Predictability and Retrospectives: The Un-Examined Miss Problem

CASE STUDY · UNNAMED ENGAGEMENT.

What follows is drawn from a change programme delivered inside a global research and advisory business, in one of the largest units it operates. Nine product and technology teams carried a top line of hundreds of millions of dollars, with a target of doubling it. We put evidence under the programme by scoring the organisation on twenty-five dimensions, then layering a qualitative pass on top for what those scores leave out.

The analysis started with the map.

Start with the map

r ≥ 0.54 · Delivery predictability
System resilienceArchitecture roadmapAutomated testingBusiness value clarityFeature definitionDelivery predictabilityTeam alignmentPortfolio agilityPortfolio visionPrioritisationProduct intakeProduct management rolesProduct roadmapPI planning rolesQuality metricsQuality confidenceRelease processRisk managementFlow of workStakeholder managementSustainable paceTeam-level planningTechnical debtTestable requirementsCross-team planningTalent clock-speed
size: mentions in the written observationscolour: mean sentiment, low to highedge: significant score correlation, thicker is stronger

Correlation, not causation: every edge is a measured statistical association (p<0.05) from this engagement, not an asserted cause. Edges below |r|=0.35 are omitted for legibility; the rest are drawn thicker and more opaque the stronger they are. The map stays focused on Delivery predictability: hover any dimension to preview its own connections against it, and use the strength slider to keep only its strongest links.

The map opens centred on delivery predictability: whether the work lands when the organisation said it would. Look at its neighbourhood. Not calendars, not velocity charts: its strongest neighbours are smooth flow of work and quality metrics, with prioritization and obstacle removal just behind. Predictability here moves with the machinery of delivery and the ability to measure it. That second one turns out to be the tell.

What the numbers said

median 9 scale 1-10
012345678910

Landing at a median of 9 out of 10, the scores put seven of sixty-one respondents at 4 or below. On the scores, delivery predictability is nearly a non-issue. The words disagree, and they disagree about something the scores can’t see.

What the words said

What they scored
What they wrote: sentiment same axis
012345678910

The two dark blobs not lining up is the chart.

Produced by the team survey: seven coded units, five negative to two positive. Placed on the score axis the writing centres on 3, with the scores centring on 9. Across team survey and leadership interviews together, ten mentions, eight of them sore. In sequence, the sore units describe a commitment that costs nothing to make:

“The team says yes to too much during the ‘what’ meeting but there is no accountability in saying yes to too much.”

“We may be overestimating our team’s capacity, as some of the items included in our sprints are carried over to later sprints.”

“Rarely do we complete all tickets committed to in a sprint.”

“This year, we have seen several projects carry over from quarter to quarter.”

Then the sharpest unit in the set names the other missing loop: what happens after the miss:

“If IT doesn’t meet its target, there is no real accountability, we missed this. Why did we miss it? What’s next?”

Missed dates weren’t the core problem. Un-examined misses were. Saying yes had no cost at commit time, and the miss had no autopsy afterwards, so the same miss came back every quarter, wearing a new project’s name.

Where it lived

Team 9 med 10
Team 1 med 9
Team 2 med 9
Team 4 med 9
Team 3 med 9
Team 6 med 7
Team 7 med 7
Team 8 med 6
Team 5 med 4
012345678910

Nine teams, anonymised, best to worst. Thin rows render wider and flatter; less data looks uncertain, not falsely precise.

Of nine teams, seven hold medians of 7 or above, up to a perfect-looking 10. The other two sit at 6 and 4. Whatever the organisation-wide average says, unpredictability has an address, and from most seats in the building it’s invisible.

What predictability turned out to be entangled with

What the map puts next to it:

  • Smooth flow of work (r = .64), the strongest pull
  • Quality metrics (r = .63)
  • Prioritization (r = .60)
  • Risk and obstacle removal (r = .60)

Quality metrics in the top two is the reminder that a miss can only be examined against a measurement, and this client’s measurement layer was itself sore, as its quality-metrics case study shows from the other side. Forecasting was starving for the same reason:

“Right now, it’s hard to tell what our current velocity [is], so I can try and project 8 months from now but not 2 years from now.”

No measured velocity, no scored commitments, no review of misses: three absences that make every forecast a guess with a date on it. Which explains why predictability tracks flow and measurement instead of any forecasting technique. Correlations are what this reads; no cause is proved.

What happened next

The direction of the fix was already written down, by the people living the misses. The positive units in this dimension are both wishes, and they wish for the same thing:

“I’d rather the team is very clear on what they can take on and deliver against that, instead of us guessing which requests will have no progress against.”

“Would really like the team view this as a goal, as it would help communicate updates to business stakeholders on when they can expect certain deliverables.”

Not better estimates: an honest yes, and a visible score against it. The check that travels: take your last three missed commitments and ask two questions. What did saying yes cost? Who examined the miss? If the answers are ‘nothing’ and ‘nobody’, no velocity math will fix the forecast.

A note on the data

A real engagement; an unnamed client, and it stays that way. Team names and identifying details are removed, and quotes are lightly edited for anonymity. The figures and charts on this page compute fresh against the underlying assessment data each render, covering close to 2,000 data points, quantitative and qualitative alike, with nothing asserted. One organisation, correlations not causes.

Morné Wiggins · Agility at Scale · Talk to me

Privacy Preference Center