In the rebuild vs refactor software argument, the rebuild usually loses, and for a reason that has nothing to do with code. The first system failed because nobody wrote down what the business actually does, the scope grew as it went, and it was delivered in one big bang. Start a rewrite with the same three conditions and you get the same result, slower, because this time you are also keeping the old system alive while you do it.
A ground-up rebuild is the right call in one situation: when the platform underneath is dead. Dead means out of support, no way in and unable to take one more feature; ugly is not dead. Everything else gets replaced a piece at a time.
Why do software rewrites fail?
Because the hard part was never the code, it was the behaviour nobody documented. Martin Fowler's description of the strangler fig pattern gives the verdict on the switch-over rewrite: "we've seen this simple-sounding plan go down in flames most of the time." His reason is the one we see in discovery: "replacements seem easy to specify, but often it's hard to figure out the details of existing behaviour." Fifteen years of edge cases live in the old system and in the heads of the people who use it, and none of them are in the brief.
Fred Brooks named the second failure in 1975. The second-system effect is the tendency of a replacement to carry every improvement, optional feature and generalisation that was deferred from the first one. The rebuild becomes the wish list, and the wish list is why the first build ran late.
Then there is duration. McKinsey and Oxford's study of more than 5,400 IT projects (2012) found large projects ran 45 per cent over budget and delivered 56 per cent less value than predicted, with every additional year of delivery adding 15 per cent to the overrun. Flyvbjerg's 2022 analysis of 5,392 IT projects across 66 countries found the average project cost 1.8 times its estimate, with a fat tail of catastrophic overruns. A rebuild is, by definition, a long project delivered all at once. It sits in the tail.
And the firms that most want a rebuild are the least likely to finish one. McKinsey's 2023 research found companies in the worst fifth for technical debt were 40 per cent more likely to have modernisations left incomplete or cancelled than those in the best fifth.
Is it better to refactor or rebuild legacy software?
Replace it incrementally, by default, and rebuild only when the numbers force you. McKinsey's 2020 CIO research says "avoid a big-bang approach", because infrequent megaprojects carry "high execution risk", and a greenfield rebuild is "a last resort" that becomes defensible only when technical debt exceeds 50 per cent of the asset's value. DORA's finding that working in small batches predicts delivery performance applies to replacing a system as much as building one.
What you see | What it usually means | What to do |
|---|---|---|
Slow to add features, but it runs | Technical debt, not a dead platform | Refactor in place, funded as a share of every sprint |
Parts of it are fine, parts are a liability | A mixed estate | Replace the liabilities one workflow at a time behind a new front door |
Out of vendor support, no API, cannot take one more feature, nobody can change it safely | A dead platform | Rebuild, but scope it from the business, not from the old system's screens |
A cheap build with no tests, hard-coded credentials and a database that cannot take a second feature | A dead platform that happens to be new | Rebuild, and this time buy the discovery |
You can get an app built for £5,000, and we have inherited several of them. A rebuild there is right, and it costs more than the original build would have.
When is a rebuild the right call?
When the platform is gone, and only then, with the scope drawn from the business rather than from the screens you are replacing. The National Audit Office reported in January 2025 at least 228 legacy systems across departments, 28 per cent of them red-rated for operational and security risk, 53 per cent without fully funded remediation plans. The Public Accounts Committee's October 2025 report counted 319 systems, a quarter red-rated. Government spends around £14 billion a year on digital, and cannot big-bang its way out. A 40-person firm should not try.
The current temptation is to believe AI makes the rewrite cheap. Gartner predicted in June 2026 that more than 70 per cent of platform-exit projects started this year will fail to deliver their intended benefits because of overestimated generative AI tooling. It is a forecast, and it matches what we see: the model can translate the code, and the code was never the problem.
How do you replace a system without a big bang?
Put a new front door in, move one workflow behind it at a time, and keep the old system as the source of truth until the new one has proved itself on live work. That is the strangler fig in practice, and it has three properties a rewrite lacks: the business gets value in weeks, each move is small enough that a failure is recoverable, and the undocumented behaviour surfaces one workflow at a time, when it can still be handled.
One that came to us this summer: a subscription service had been built on a content-management system that had grown into a data platform, with a mobile app on top and a customer admin area that was the raw CMS back end. The instinct was to rip it out. The recommendation was to keep the mobile app, keep the CMS for what it is good at, which is content, and replace the parts that had outgrown it, starting with the one holding sensitive data.
Our technical debt post argued for funding it as a share of every sprint, and PatientGo is what that looks like over five years: from hundreds of patients to tens of thousands, 36 countries, more than seven external systems, and no year in which the whole thing was rewritten. The business systems we build now start with two weeks of discovery on what the old system actually does, because that is the document the first build never had.
