The Most Dangerous Moment in a Network Change Is When Everyone Thinks It Worked

Raja TadimetiCo-founder & CEO

There is a moment in almost every maintenance window that feels safe and isn't.

The commit went through. The prompt came back. The device is reachable. The BGP session you were worried about is back to Established. Someone on the bridge says, "Looks good."

That moment gets treated like the finish line. It is usually the start of the part that matters.

Execution is not the hard part

The hard problem in change management is not executing the change. Most competent teams can already do that. They write runbooks, sequence the steps, schedule the window, get the approvals, push the config with discipline. The hard problem is knowing whether the network is actually healthy once the work is done.

That sounds obvious. It is also where a surprising amount of operational risk still hides.

A network does not fail only when a device goes dark. It fails when a path comes back asymmetric. When a routing policy does something slightly different from what the author intended. When a tunnel re-establishes but with half the throughput it had before. When a dependency outside the blast radius quietly absorbs the damage and nobody notices until Monday morning. When everything looks stable right up until production traffic finds the one thing the window did not exercise.

So the most dangerous words in change management are often not "the change failed." They are "the change completed."

Completion is an execution milestone. It is not proof.

What the 2026 outage numbers say

The numbers back this up, and they are not getting better. In Uptime Institute's 2026 outage analysis, 57% of respondents said their most recent major outage cost more than $100,000, and for the second year running one in five said it cost more than $1 million. The leading driver of human-error outages is still the same thing it was last year: people not following the procedure, or the procedure itself being wrong. Uptime also notes that outages tied to fiber and connectivity are rising and tend to last longer than other kinds.

I read that as a verification problem more than a tooling problem. Almost nobody in those postmortems forgot to run the change. They ran it, called it done, and found out later that "done" and "healthy" were two different things.

Did the network absorb the change safely?

Most change processes are still much better at answering "Did we do the thing?" than "Did the network absorb the thing safely?"

The second question is the one operators actually care about. A router upgrade is not successful because the router is back. A firewall policy change is not successful because the rule committed. A WAN change is not successful because BGP is Established again. A change is successful when the surrounding network, the expected traffic, and the services that depend on it all behave the way they were supposed to afterward.

That takes a different mindset. Post-change validation cannot be a checklist stapled onto the end of the window. It has to be the center of the process. The team needs to know what should change, what should not change, what evidence counts, and what would make the outcome ambiguous rather than clean.

Ambiguous is the case that gets people. A clean failure is easy: you roll back. A clean success is easy: you close the window. The dangerous outcome is the one where the tunnel is up but latency doubled, or the route is present but learned from the wrong neighbor, or the interface counters are clean because traffic silently moved somewhere else. Nothing is red. Nothing is right either.

Where AI actually fits

This is also where I think the next useful wave of AI in network operations shows up.

Not as a system that improvises production changes on its own. Nobody I talk to wants an agent pushing config at 2 a.m. without a human in the loop.

The more important shift is toward systems that help a team prove the network is healthy with more rigor than a static runbook can. Systems that know the difference between an expected flap and the first sign of a broader regression, that account for a device's role, its path dependencies, and any known weakness when deciding what gets tested, and that put pre-change and post-change state side by side so the operator can continue, pause, or roll back based on evidence rather than relief.

That is the work we are doing at REAP, and it is a much more useful future than "AI that makes changes."

The teams that get better at change over the next few years will not just automate more steps. They will get more rigorous about proof. The standard moves from "the procedure ran" to "the network demonstrated that it is healthy." A green ping, a returned prompt, and a quiet bridge are not proof.

That is a higher bar. It is the right one.

See REAP in action

Watch REAP reason through a live incident on your network.