A laptop screen showing time-series metrics tracked over several months
Insights
Operations·5 min read·August 2026

Nobody budgets for year two

By Zero2Launch

The pilot-to-production gap gets all the attention. The one that quietly kills more systems comes after launch, when the assumptions expire and the person who understood the thing has moved on.

The system went live in March and it worked. Six weeks in, the numbers were easy to defend: faster triage, fewer escalations, a support floor that had stopped dreading Mondays.

Nobody looked closely again until November, when someone noticed escalations had crept back to roughly where they started. Reading back through the logs, the drift had begun sometime in June. No alert fired, because no one had built one.

The gap between a pilot and a production system is well documented, and everybody in this industry has an opinion about it. The gap between production and year two barely gets discussed. It has killed more of the systems we have been called in to rescue than any modeling mistake.

Things quietly stop being true

A deployed system is a set of assumptions about the world, frozen on the day you shipped it. The world keeps moving.

Products get renamed. A policy changes and a chunk of your retrieved documents are now wrong in a way that still reads as confident and well-sourced. An upstream team adds a field, or drops one, and extraction starts failing on a category nobody thought to test. Users work out what the system is bad at and route around it, which shows up in your dashboard as lower volume rather than lower quality.

None of that is a model failure. All of it looks exactly like the model getting worse.

The person who understood it leaves

Almost every production system we meet has one person who genuinely understands it. Not the vendor, not the platform team. One person, usually whoever pushed hardest to get it funded.

Then they get promoted, or reorganized, or they leave. The system carries on running, because that is what software does. What goes is the judgment: a feel for what a normal week looks like, which failures are worth chasing, whether a dip in usage is the summer or a symptom.

Documentation does not transfer that, and neither, in our experience, does a handover meeting. What transfers is a dashboard somebody actually opens and a recurring half hour on a real calendar.

This is also why governance keeps slipping. Organizations treat it as a launch requirement, tick it off, and discover it was never a milestone in the first place. Most expect the work to run well past the first year.

Budget for it as a line item

Ongoing ownership tends to get filed under contingency, which is another way of saying it gets cut in the first tight quarter. Name it, cost it, and assign it to somebody before launch rather than after the first bad month.

The work itself is dull. Re-run your evaluation set whenever a provider ships a new model version and changes behavior without telling you. Refresh the retrieval corpus. Read a sample of what got escalated. Retire the prompts written against a process that no longer exists. Keep one metric that would have caught the June drift, and put it somewhere a human sees it weekly.

None of this will ever get a slide at the quarterly review. It is the whole difference between a system still earning in year three and one switched off with a shrug and the words “the AI didn't really work out for us.”

So the question worth asking before launch is not whether the system works. You already know it works. That is why you are launching it. The question is who will notice when it stops.

Sources

  1. 1.Deloitte, “State of Generative AI in the Enterprise” (Wave 4), 2025

Start with an audit

Find your highest-ROI AI in weeks, not quarters.

A focused audit of where AI moves the needle in your business, with a prioritized roadmap and an ROI model. No lengthy consulting cycle.

Average response time: 24 hours