Skip to main content

← Crash log Run 016

What Late QA Really Costs (It's Not the 100x Curve)

The 1-10-100 bug curve is folklore. Late QA still costs you, in rework, escaped defects and regulator fines. Here's the evidence, charted.

Run 016 8 zones 0 samples 6 diagrams 0 compares

A while ago I started keeping a list in my notes of how QA gets quietly removed from teams. Goodbye defect reports. Goodbye reproduction steps. Goodbye test cases and end-to-end runs. Hello "everyone is responsible for quality", which in practice means nobody is. The usual defence of QA in those conversations is one chart, and it turns out that chart is the weakest argument we have.

This post is about the business side, so there's no code. It covers where the famous cost curve came from, what the better evidence says, what late QA costs in FinTech specifically, and the cheap things I put in place on a new project so the costs never show up.

The chart everyone shows

You've seen it in a slide deck. A bug costs 1 unit to fix in requirements, 10 in testing and 100 in production, credited to the "IBM Systems Sciences Institute".

Laurent Bossavit went looking for the source and traced it back to a 1987 textbook, which cites course notes from an IBM internal training programme. No underlying data has turned up (The Register, 2021). Hillel Wayne's verdict was that the chart is, as far as anyone can tell, made up (2021).

The real research is more careful than the slide:

  • Boehm and Basili (2001) wrote that a fix after delivery is often 100 times more expensive than one made in requirements or design, and that for small, non-critical systems the ratio is more like 5:1 (IEEE Computer, PDF).
  • Menzies et al. (2016) looked at 171 projects and about 47,000 defect logs, and found no delayed-issue effect. The median fix time for a coding defect went from 10 minutes in code review to 16 in acceptance testing (arXiv).
  • Capers Jones calls the 100x claim an urban legend (2012, PDF).
How much more a late fix costs, by source (multiple of an early fix, log scale)
The slide (no known data)100
Boehm, large systems ("often")100
Boehm, small systems5
Menzies, 171 projects (median)1.6

These aren't measuring the same thing (Boehm talks about cost, Menzies about engineer minutes before release), so don't read the bars as like-for-like. What they agree on is that the engineer's fix time is not where the money goes. If you're arguing for QA with the 100x slide, a sharp engineering manager can take your case apart in one meeting.

Where the money actually goes

Fixing the line of code costs about the same whenever you do it. What grows is everything around the fix.

Illustrative shape, not measured data blast radius incident fix rework incident redress regulator Requirements Design Build Test Release Production fix effort rework of work built on the bug incident response, customer redress, fines
The green layer is the part the 100x slide argues about, and the evidence says it stays roughly flat. The amber and red layers are where late defects get expensive, and the red ones only exist once customers are involved.

Four costs grow with time:

  1. Rework. Boehm and Basili put avoidable rework at 40–50% of project effort, with about 80% of it coming from 20% of the defects. A wrong assumption in a story gets built on: screens, endpoints, reports and tests all inherit it.
  2. Escapes. Every stage that doesn't catch a defect hands it to a later, wider audience. More on that below.
  3. Blast radius. In nine large IBM products, about 0.3% of defects caused about 90% of the downtime (Boehm and Basili again). Most bugs are cheap. The few that reach production in the wrong place are not.
  4. Regulators. In FinTech, a production failure can come with a fine, mandatory customer redress and a report you write for someone who doesn't care about your sprint.

Testing at the end catches less

Capers Jones has benchmarked defect removal efficiency (DRE, the share of defects removed before release) across more than 13,000 projects. A pipeline that relies on testing alone removes about 85% of defects. Add design and code inspections and a disciplined process, and it's about 98% (Jones, PDF).

That gap sounds small until you count what ships:

Defects delivered to customers per 1,000 function points (Capers Jones)
Testing only (85% removed)613
Inspections + testing (98% removed)76

Jones's data also shows where the escaped defects start. Most come from before anyone writes code:

Where delivered defects originate (defects per function point, Capers Jones)
Requirements0.23
Design0.19
Documents0.12
Bad fixes0.12
Coding0.09

Coding is the smallest bar. A QA process that only starts when the code is "done" is aiming at the smallest source of escaped defects and missing the two biggest.

What it looks like in FinTech

Every one of these is a well-documented case. None was one engineer's typo found late. They were missing checks around deployments, configuration, migrations and UI, which is where early QA earns its keep.

Incident What happened Cost What would have caught it
Knight Capital, 2012 New code wasn't copied to one of eight servers, and a reused flag woke up dead code. No written deployment procedure, no second-person review, and 97 warning emails before the market opened (SEC) Lost more than $460M in about 45 minutes, plus a $12M SEC penalty A deployment checklist and a post-deploy smoke test
TSB migration, 2018 A configuration difference between two data centres went undetected during testing (TSB) £330.2M in post-migration costs in 2018, then £48.65M in fines in 2022 (Bank of England) Testing in an environment that matches production
RBS, NatWest and Ulster Bank, 2012 A software upgrade left more than 6.5M customers unable to use their accounts, some for weeks. The FCA cited weak testing of software changes (FCA) £56M in fines (FCA and PRA) Change testing and a rollback plan
Citibank and Revlon, 2020 An operator ticked one box out of three on a payments screen, and three approvers missed it. Citi sent about $894M of principal with a $7.8M interest payment (The Register) About $500M was held by lenders until Citi won on appeal in 2022 Usability testing of a high-risk screen
Commonwealth Bank, from 2012 AUSTRAC alleged 53,506 missing threshold transaction reports. The bank blamed a coding error in a late-2012 update to its deposit machines (CIO) Part of an A$700M AML penalty, which also covered other failings A regression check on the compliance reports

Don't measure "cost per bug"

If your QA budget gets judged on cost per defect, it will look worse the better it works. Jones calls this the cost-per-defect paradox. Writing and running tests is a fixed cost, so the fewer defects there are left to find, the higher the cost per defect, even though every repair takes the same time.

Cost per defect by test stage when every repair costs the same $379 (Jones's illustrative model, US$)
Unit test419
Function test479
Regression test579
Performance test779
System test1045
Acceptance test2379

Nothing about the repair changed between those rows. Only the number of defects left to find did. Track these three instead:

  • Defect removal efficiency: bugs found before release divided by all bugs found, including the ones customers report.
  • Change failure rate: the DORA metric for how often a release needs a hotfix or rollback.
  • Rework rate: DORA added "deployment rework rate" as a fifth metric in 2024 (DORA).

What I put in place on a new project

None of this needs a test management tool or a big budget. Most of it comes from a QA guide I worked from years ago and kept using because it works.

Grooming Build Handover Release Production QA starts here QA reviews the acceptance criteria test cases written before code is done dev walks QA through the change smoke + post-deploy on a prod-like env tag where found, root-cause escapes every escaped bug becomes a new grooming check
One cheap QA checkpoint per stage. The red one is the expensive place to find things, and the loop back to grooming is how you stop finding the same class of bug there twice.
  • QA in grooming. QA reads the stories and acceptance criteria before the sprint starts and asks the awkward questions (what happens on a timeout, a duplicate submit, a currency with no decimals). Requirements are the biggest source of escaped defects, so this is the cheapest hour of QA you'll buy.
  • Test cases before the code is done. I keep them fast and dirty: a spreadsheet with ID, priority, component, title, steps, expected result, and four tick columns for smoke, full regression, post-deploy and automated. You don't need a tool on an early project. The tick columns tell you what to automate first.
  • A walkthrough before handover. The developer shows QA the change for ten minutes before it's marked ready. The point is to find bugs while the change is still fresh, before it turns into rework.
  • Tag every bug with where it was found: QA, release or production. After a couple of months you have your own cost-by-phase chart with real numbers, which beats any slide deck.
  • Root cause for anything found in UAT or production, and feed the fix back into grooming as a question you'll always ask.
Tip

Stop selling QA with the 100x curve. Sell it with rework, escapes and fines, and with your own "where was it found" numbers after a couple of months.

Why it matters more now

The 2024 DORA report found that more AI adoption came with lower delivery stability (−7.2% for every 25% increase in adoption) (Google Cloud). The 2025 report says throughput is up but instability keeps rising (Google). Teams are writing more code, faster, with less review per line. Every checkpoint in the timeline above gets more valuable, not less. (I wrote about the review side of that separately.)

Why it matters for your team: in software that moves money, the expensive defects are the ones regulators and customers find, and the cheapest place to stop them is a QA question asked before the code exists.

If you want the testing side of this in practice, every integration test has the same five parts and the golden rule of mocking are where I'd start.

Sources

  • Boehm and Basili, "Software Defect Reduction Top 10 List", IEEE Computer, January 2001 (PDF)
  • Menzies, Nichols, Shull and Layman, "Are delayed issues harder to resolve?", Empirical Software Engineering, 2017 (arXiv)
  • Capers Jones, "Three Harmful Software Metrics", 2012 (PDF) and "Software Defect Removal Efficiency", 2011 (PDF)
  • The Register on the origin of the 1-10-100 curve, July 2021 (link)
  • SEC order on Knight Capital, October 2013 (PDF)
  • Bank of England and FCA on TSB, December 2022 (BoE, FCA); TSB Annual Report 2018 (PDF)
  • FCA on RBS, NatWest and Ulster Bank, November 2014 (link)
  • DORA metrics history (link) and the 2024 and 2025 DORA reports