Quality Metrics and Delivery Confidence: The Strongest Signal Joined Two Things Nobody Had
This page comes out of an organisational change programme we ran at a worldwide research and advisory firm, in one of the largest business units it operates. Nine teams handled product and technology, and revenue there reached hundreds of millions of dollars each year against a target of doubling it. Evidence had to carry the programme, so we assessed the organisation: twenty-five scored dimensions, and a qualitative pass for what the numbers flatten.
And the analysis started, as it always does, with the map.
Start with the map
Correlation, not causation: every edge is a measured statistical association (p<0.05) from this engagement, not an asserted cause. Edges below |r|=0.35 are omitted for legibility; the rest are drawn thicker and more opaque the stronger they are. The map stays focused on Quality metrics: hover any dimension to preview its own connections against it, and use the strength slider to keep only its strongest links.
The map opens centred on quality metrics, and this case study exists because of a single line on it: the strongest relationship anywhere in the data, stronger than anything about intake, roadmaps or pace, joins quality metrics to quality confidence: whether people trust what the teams ship. Which is strange, because quality metrics is one of the least-discussed subjects in the study: five written mentions in the whole exercise. The strongest signal on the map runs through one of its quietest corners, and that asymmetry is the finding.
What the numbers said
Numerically this dimension looks like a strength: a median of 8 out of 10 across sixty-one respondents, one of the more comfortable strips in the assessment. People rate the measurement of quality highly.
What the words said
The lower strip is nearly empty. That is the chart.
Then you go looking for what people wrote about quality metrics, and you find one coded observation. Not one theme. One observation, in full:
“I’m not sure that we have metrics set up for our applications (it’s possible we do, I just am not aware).”
That sentence is the entire direct written record of the strongest-correlating dimension in the study. A subject people score at 8 out of 10, and the only person who wrote about it isn’t sure it exists. The scores rated the idea of measurement; the words couldn’t point to the thing itself.
Where it lived
Nine teams, anonymised, best to worst. Thin rows render wider and flatter; less data looks uncertain, not falsely precise.
Per-team medians stretch from 9 at best to 6 at worst. The spread is narrow across teams here, but the ordering matters: measurement here was a per-team accident rather than an organisational practice, and the teams’ quality confidence varied with it, project by project.
What quality metrics turned out to be entangled with
The numbers behind the map’s strongest line:
- Quality confidence (r = .77), the strongest edge in the entire network
- Automated testing (r = .70)
- Stakeholder management (r = .69)
Why would confidence track measurement this tightly? The written observations on quality confidence show what fills the space when measurement is missing. Quality gets adjudicated by attention:
“I have observed that we ignore bugs which we found internally and do not prioritize them till it becomes big or noticed by someone from business or leadership team.”
“There are some projects where business UAT ends up being a QA process and business finds very obvious issues that should have been caught by QA.”
“QA is shared across the dev teams and there isn’t dedicated QA for each team, which is why we see a varying degree of QA quality depending on the project.”
“The quality of QA that we get depends on the project.”
Without something stable to measure against, an internally-found bug carries no weight until leadership notices it; the customer’s acceptance test becomes the de-facto quality gate; and confidence varies project by project because it has nothing to rest on. Where a team had measurement, it had something to trust; where it had neither, both scores collapsed together. That is one coherent reading of why the two dimensions move almost in lockstep. It reads a correlation; it does not prove a cause. The third edge fits the same story: the business’s trust (r = .69) rests on the same substrate the team’s own confidence does.
What happened next
Measurement is buildable, which makes this an actionable finding. A portion of the post-assessment spend went to CI/CD and automation, the machinery that produces quality signals as a side effect of shipping. From there the organisation stayed close to the teams as circumstances shifted, and what the resulting numbers showed is not mine to report, that is theirs. The approach is what is worth copying: when quality confidence varies wildly by project, check whether anything is actually being measured before debating anyone’s standards.
A note on the data
A real engagement; an unnamed client, and it stays that way. Team names and identifying details are removed, and quotes are lightly edited for anonymity. The charts and numbers on this page are recalculated from the underlying assessment data on every run, spanning close to 2,000 quantitative and qualitative observations combined, none of them asserted. One organisation’s measurement, correlations not causes: a pattern to check for, not a rule.
Read Next
- Built-in quality vs. inspection
“Business UAT ends up being a QA process” is inspection pushed all the way to the customer: the exact anti-pattern this section names.
- Why automation is non-negotiable
Quality metrics tracked automated testing at r = .70 here: measurement and automation are the same substrate.