Module 04 · 25 min
Learning From Failure
Six cases, one pattern
Detailed case studies from 1992 to the present day. Read them for the mechanism, not the schadenfreude: in every case a series of individually survivable problems combined.
The classic: London Ambulance Service, 1992
1992
London Ambulance Service Computer-Aided Despatch
System withdrawn; 46 deaths linked to delayed response
On 26 October 1992 the LAS CAD system went live across London and collapsed within hours. Operators reverted to a manual system that then seized up entirely. The chief executive resigned within days.
What went wrong
- Poorly trained staff did not update the system with unit location and status.
- The increasingly out-of-date database despatched units non-optimally and sent multiple units to the same call.
- A software bug generated a flood of exception messages; unanswered exceptions generated repeat messages.
- Lists scrolled off the top of the screens and were lost.
- The public repeated unanswered calls, adding load to an already saturated system.
The lessonThe National Audit Office concluded that the small software error was only the straw that broke the camel's back. The attitudes of key LAS members toward the project, and the unreasonable constraints they placed on it, allowed the failure to occur.
That was 1992. The question that opened the original workshop still stands: have we learned the lesson?
Terminal 5, 2008
2008
Heathrow Terminal 5 opening
28,000 lost bags, 700 cancelled flights, 150,000 disrupted passengers
A £4.3bn terminal delivered on time and on budget — and then an opening day that became a national news story. The building worked. The sociotechnical system around it did not.
What went wrong
- Shortage of staff car parking spaces.
- Only one employee security checkpoint operating.
- Some staff unable to log on to the computer system.
- Hand-held communication software running slowly.
- No managers on the ground to re-allocate work.
- Shortage of bar-code reading storage bins.
- Baggage handling staff arriving late; 60 staff queuing to get into the terminal.
The lessonBy 6am three planes had left without bags; by midday 20 flights were cancelled; by 4pm the baggage conveyor stopped and check-in was suspended. Not one of those seven problems would have brought the terminal down on its own.
TSB core banking migration, 2018
2018
TSB migration to the Proteo4UK platform
Around £330m in costs and compensation; ~£49m regulatory fine; CEO resigned
TSB moved 5.4 million customers and roughly 1.3 billion records from Lloyds' systems to a new platform built by its parent Sabadell, in a single weekend cutover. Customers were locked out, saw other people's accounts, and support lines collapsed for weeks.
What went wrong
- A big-bang cutover for the entire customer base with no viable rollback.
- Testing that did not reproduce real production load or the full data estate.
- Two data centres configured differently, so only one behaved as tested.
- The board relied on assurance that the platform was ready rather than on independent evidence.
- Contact-centre and complaints capacity was sized for a good day, not a bad one.
The lessonThe independent review found the board did not sufficiently challenge the readiness assessment. This is the C-NOMIS pattern in modern dress: a project board accepting that everything was 'going well'.
Birmingham City Council Oracle, 2023
2018 – 2023
Birmingham City Council ERP implementation
Budget of roughly £19m; projected costs rose to around £100m+; contributed to a s.114 notice
Europe's largest local authority replaced its finance and HR systems with a cloud ERP. Heavy customisation to fit existing local ways of working, followed by a go-live in 2022 that left the council unable to produce reliable accounts or reconcile bank transactions for many months.
What went wrong
- Standard product customised extensively rather than processes being standardised to the product.
- No cumulative view of what the accumulated change requests were doing to cost and timescale.
- Insufficient in-house capability to hold the integrator to account.
- Go/no-go decisions taken under schedule pressure, with known defects carried into live running.
- Audit and reconciliation capability lost at exactly the moment it was most needed.
The lessonThe failure mode is identical to C-NOMIS in 2007: no resources allocated to simplifying and standardising business processes across a fragmented organisation, so the technology was asked to absorb the fragmentation.
Post Office Horizon, 1999 – present
1999 – present
Post Office Horizon
900+ wrongful prosecutions; compensation and inquiry costs measured in the high hundreds of millions
An accounting and retail system rolled out to thousands of sub-postmasters produced unexplained shortfalls. Rather than treating the discrepancies as a signal about the system, the organisation treated them as evidence of theft, and prosecuted.
What went wrong
- Defects and remote access capability known internally but not disclosed.
- A culture in which the system's output could not be challenged by the people using it.
- Sub-postmasters — the users — held no power in the governance of the programme.
- Contradictory evidence was managed as a reputational risk rather than a technical one.
The lessonThis is the extreme end of ignoring the social system. The technical faults were serious; the culture that made it impossible for users to be believed is what turned a defective system into a miscarriage of justice on an industrial scale.
The private sector is no better
2004
Hewlett-Packard ERP centralisation
$160m in order backlogs and lost revenue — more than five times the project's estimated cost
HP's project managers knew all of the things that could go wrong with the ERP centralisation programme. They simply did not plan for so many of them to happen at once.
What went wrong
- A series of individually small problems that arrived simultaneously.
- Risk assessed item by item rather than in combination.
The lessonGilles Bouchard, then CIO of HP's global operations: 'We had a series of small problems, none of which individually would have been too much to handle. But together they created the perfect storm.' That is the clue the rest of this course follows.
- MFI, 2004–05: a big-bang ERP changeover crashed, dispatch failed, over £30m was spent correcting it, and the retailer went into administration in 2008 with 1,500 job losses.
- C-NOMIS, 2004–07: an offender records system approved at a £234m lifetime cost; by July 2007 £155m was spent, it was two years late, the lifetime estimate had reached £690m, and it was suspended. The Public Accounts Committee called it 'a spectacular failure — in a class of its own'.
- Defence Stores Management Solution, 2002: halted after £130m, of which £118m was written off.
Key takeaways
- In every case, individually survivable problems arrived together and the system tipped.
- Boards repeatedly accepted assurance instead of evidence of readiness.
- Where organisations refused to standardise their own processes, the technology absorbed the mess — and broke.
Take it back to work
- Which of these six patterns is most alive in your current programme?
- If your go-live went wrong at 6am on a Monday, what is the actual rollback plan — and who has practised it?