Memfida

Article7 min read

The shutdown plan is seven years old

A legacy ecommerce core was scheduled for shutdown in 2019 and has more systems depending on it now than it had then. Three replacement projects put it there. Outlived three CIOs, one office move, and a vendor that is not reachable. The cost of keeping it is growing, split across ten cost centres, and attributable to nobody. The cost of removing it requires one owner and one budget line, but yet, the legacy runs today.

The e-commerce core I am working to decommission was scheduled for shutdown in 2019; it is still running in production today, and it has outlasted three replacement platforms, two Chief Information Officers, one office move, and one supplier that is no longer reachable. It also has more systems depending on it now than it had on the day the shutdown was signed off. Every attempt to remove it put them there.

The plans were good

I found the original deck in an old folder of sharepoint. Phased carve-out of the order management system, then payments, then the product catalog, with dates against each phase and a named owner for every workstream. The sequencing is correct. I would sequence it the same way today, and I have the advantage of knowing what happened next.

The two plans that follow it are also good, produced at intervals of roughly eighteen months, each one written by people who had read none of the others, each one confident, and each one accurate about the target state and ignorant of the reasons that impeded the previous attempts. The system is still running. Nothing here was caused by a bad plan.

Every replacement left an integration behind

The most recent attempt was a new team, formed to carve the main functionality out of the legacy platform and rebuild it as a new microservice to run on the modern container platform. The carve-out did not complete. The traffic never fully moved, the old code paths stayed live as a fallback, and the project lost its funding at the point where it had built enough of the new system to need the old one. The legacy platform came out of that project with three more consumers than it went in with, because the partial replacement now calls it for the parts that were never finished.

I have written one of those calls. On a previous engagement I built the adapter that let a new service read customer records out of a system everyone agreed was going away, and I argued for it on the grounds that it unblocked the delivery date, which it did. That system is still there. The adapter is one of the reasons.

That’s how it works. A replacement project that succeeds removes a dependency. A replacement project that stops halfway adds one, because everything it built has to reach back into the system it was replacing to be useful at all. The legacy platform does not survive these attempts by resisting them, it survives by absorbing them, and it comes out of each one with more callers, and a stronger claim on the next year’s budget meeting. The system is harder to switch off when more dependencies pile up.

Three replacement attempts have produced a platform that is more central to the business than it was in 2019, and every one of those attempts was launched specifically to reduce its protagonism. The organization is failing to remove it, and it is feeding it.

Decommissioning is staffed as cleanup

The team assigned to the decommissioning is small on purpose, and the reasoning is defendable on its own terms: the work is understood, basically housekeeping, that’s what the project/engineering managers say; it is less relevant than you think, another lie; and it is not where the growth is, this one is sort of true. It gets four halftime engineers, 20 percent of a delivery manager, and no product owner, because there’s no product to own 😄. The replacement platform running alongside it gets a full team, a roadmap, a launch date, and a capable product owner 😉.

Both teams draw on the same handful of engineers who understand the legacy system. Every quarter, those engineers are allocated to the thing with the launch date. This is not a failure of prioritization. It is prioritization, executed correctly, against a storyboard where one of the two projects does not appear…guess which one.

A launch produces an announcement, a demo, a slide in a quarterly review, theater, and a promotion case. A shutdown produces a decrease in a number that nobody was watching. I have not met the engineer who was promoted for a deletion, and I have met several who were promoted for an integration that made a deletion harder, including myself.

Meanwhile, the old platform works. It processes orders; it has processed orders every day for fifteen years, and on the metric the business cares about most, it outperforms every system proposed to replace it, because those systems have never processed a live order at full volume, and it has processed several billion.

The cost is real, and nobody has ever seen the total

I ran the numbers because I was losing an argument. The decommissioning had been neglected for the fourth time; the reasoning given was that the platform was cheaper to keep than kill, and I wanted to know whether that was true before I disagreed with it in a room where everyone outranked me.

Infrastructure only: €300,000 per year, thirty percent of the total cloud bill, €2,100,000 since the first shutdown plan was approved.

That excludes the engineers who maintain it, the operations staff who carry its pager, the retained vendor that costs as the infrastructure because the people who wrote the system have left and knowledge is scarce, and the engineer-days spent by teams that do not own it but still have to integrate with it. All these costs are hidden and hard to calculate, but they are not cheap.

The reaction inside a business of this size is measured and calm, which surprised me until I understood why it is the reasonable reaction. Nobody in the organization has ever been shown that number. It does not exist anywhere as a single figure. It is split across ten cost centers, and inside each one of those, it sits below the threshold that triggers a review.

Waste at this scale is not tolerated because no one is comfortable with waste. It is tolerated because of the geometry of how it appears. The cost of keeping the system is continuous, distributed across many owners, and attributable to none of them. The cost of removing it is a single project with a single budget line, a single owner, and a named person who has to defend it in a steering committee against way more attractive and visionary project proposals that have a launch date and green positive numbers attached to them.

A manager that inherited this beast along with ambitious goals aligned with the new company vision has no reason to touch the one project whose full scope is not visible and easily camouflaged under a pile of failed and unfinished attempts that only bring bad memories.

What one quarter buys

An organization that cannot fund a shutdown will fund an adoption within the same quarter, and the two facts have one cause. An adoption is a launch. It arrives with a vendor, a business case, a target architecture, and a date, and it is legible to a steering committee in a way that removal is not. That is how a company ends up running four order management systems at once.

The work that would finish a shutdown is not large. It needs the legacy specialists exclusively rather than partially, which is the only thing it needs that it has never once had.

Inventory every consumer and produce a call graph from the traffic rather than from the documentation, since the documentation describes just a portion of what is running. Cut the remaining integrations over one at a time, with the old path live behind them, in the order that removes the most callers first rather than the order that delivers the most visible functionality first. Only then rebuild what is left, which by that point is a fraction of what the original decks scoped, because most of what the platform does is serve consumers that should not exist.

That sequence needs fulltime engineers who know the system, undivided, for six months. Not a reorganization, not a transformation program, not a new strategic platform. One quarter of the innovation track was spent on subtraction.

Everyone I have spoken to agrees with it. It does not happen because a quarter spent on subtraction produces nothing to announce, and the organization has now paid €2,100,000 in infrastructure alone to avoid a quarter with nothing to announce.

And if the underlying issue is that the engineers don’t find this work interesting, give them motivation. Make adoption of new technologies conditional on retirement. A new platform gets funded when the decommissioning of its predecessor is in the same plan, with the same owner and the same date. Without that condition, every adoption is an addition.