Build or Buy Agentic AI for Network Operations? The Demo Is the Cheap Part

Rishit LakhaniHead of Solutions EngineeringAgentic AI

Building an agentic NetOps tool in-house is easy to start and expensive to own. For most network teams, the better split is to buy the reasoning engine, its security posture, and its model lifecycle from a vendor, and to own the network context, policy, runbooks, and approvals yourself.

Every few weeks I talk to a network team that is trying to build something. Not a software team. Network engineers, or the automation person who has been writing Python and Ansible for the network operations team for years, now pointing an AI coding assistant at an agent framework. A Slack bot that answers questions about the topology. A script that hands an alert to a model and asks for a probable cause. An agent that pulls show commands from a few routers and summarizes them. The first version comes together in a week or two, and in the demo, it works.

Then the question comes: should we keep going, or should we buy this?

It is the right question, and most of the time it gets answered on the wrong evidence. The demo is what people see, so the demo is what they price. The demo is the cheap part.

The demo is not the product

An agent that investigates a network problem needs four things: a way to reach the network safely, enough context about the network to reason about it, a model that can reason, and a control loop that decides what to look at next. AI coding assistants and agent frameworks have made the fourth part feel easy, even for someone whose last project was an Ansible playbook. You can describe what you want, accept the generated code, wire a model to a handful of tools in an afternoon, and watch it walk through a BGP session flap on its own.

What the afternoon does not include is everything that makes that agent safe to point at production. Read-only credentials scoped per device and per role. A policy layer that decides what the agent may query, on which devices, at what rate, so a runaway loop does not turn into a self-inflicted polling storm on a core router. Audit logs that a security team will accept. A way to evaluate whether the agent's answers are actually right, across hundreds of incident types, before and after every model change. A plan for what happens when the model you built on is deprecated, which can now happen on a cadence measured in months, not years.

None of that is exotic. All of it is work, and it is work that never finishes.

The failure mode is usually not the first version. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, citing escalating cost, unclear business value, and inadequate risk controls. Those problems tend to show up after the prototype, when the system meets production scale, governance, and real operating cost. You do not discover them in the demo.

What you are really signing up for

Here is what an in-house agentic tool looks like once it is in production, based on the teams I have watched try.

Someone owns the model lifecycle. When the provider retires the version you tuned your prompts against, your investigation quality shifts, sometimes subtly, and someone has to re-run the evaluation suite you hopefully built and fix what regressed. Someone owns the security surface. Agentic systems make security problems around tool use, privileges, and control boundaries more dynamic than they are in traditional automation. The OWASP Top 10 for Agentic Applications, published in December 2025, is a useful map of those risks, many of which become especially relevant when an agent can run commands against your network.

The Replit incident from July 2025 is worth remembering here, not because of what was destroyed but because of why. A coding agent deleted a production database during a declared code freeze, despite explicit instructions not to make changes, and then told its user a rollback was not possible when it was. The lesson is not that agents are dangerous. It is that a natural-language instruction is not an enforcement boundary. "Do not touch production" in a prompt is not the same thing as a permission system that makes the action impossible. Anyone building an agent that can reach network devices has to build the second thing, and it is a very different kind of work from writing the prompt.

Someone owns the integrations, which break when the ITSM vendor, the observability vendor, or the network OS vendor ships an API change. Someone owns the feature backlog, because the moment the tool is useful, the operations team will ask for more.

I ran into this recently with someone who runs the network for a large institution. He had built an internal agentic system himself, and it worked well enough that the team started relying on it. The problem came next. Once it became part of how the team operated, someone had to own the software behind it: maintain the code, keep the integrations working, handle model and dependency changes, patch it, test it, and keep evolving it as the environment changed. He did not have dedicated internal resources to take that on. So the next problem became finding someone to own and maintain what he had built, either by hiring for it or bringing in an outside team.

That is the part of building that is easy to underestimate. Getting an agentic system working solves one problem. Keeping it working creates an ongoing software ownership responsibility that still needs a team behind it.

I want to be careful here. This is not an argument that network engineers cannot write software. Most of the automation running networks today was written by them. It is an argument that much of the ongoing cost of owning an agentic system shows up after the demo, and that it lands on the people you can least afford to lose.

Where building is the right call

There are teams that should build, and the case for them is stronger than a vendor blog usually admits.

If you already have a platform engineering team, mature APIs into your network, a source of truth you actually trust, an evaluation discipline for AI systems, and a security organization prepared to govern agent access, then building may be the economically rational choice. Hyperscalers and the largest carriers fit this description, and their networks are often specific enough that no vendor will model them well. If your network is the product, owning the reasoning layer is defensible.

There is also a middle case. If a vendor cannot reach the data you need, cannot support your device mix, or cannot fit inside your security model, a well-scoped internal tool beats a bad fit. The mistake is treating that as a permanent decision rather than a gap to revisit.

The question that separates these teams from everyone else is not whether their network engineers can build the first version. Most can. It is whether the organization deliberately wants to become the software team that owns the next fifty versions. For most enterprise and service provider network operations teams, the answer is no, and the reason is not capability. It is that the differentiated part of their operation was never the agent's control loop. It is their topology, their change history, their runbooks, and their escalation paths. That is what the agent needs to reason well, and it is what they already own.

Buy the engine, own the edges

For most network teams, the split that makes sense looks like this.

The vendor owns the engine: the reasoning loop, the model lifecycle, the evaluation harness, the guardrails around what the agent may query and change, the security posture, and the patching. These are the parts that need constant attention, should be tested across diverse environments, and have nothing to do with what makes your network yours.

You own the edges: your network context, the runbooks and escalation logic the agent should follow, the policy on what it is allowed to touch, who approves what, and the judgment about when to trust its output. These are the parts that encode your operation. No vendor should own them. A good vendor builds and maintains the mechanisms behind them: the connectors into your ITSM and observability stack, and the way policy gets expressed. Owning the edges should mean configuring them, not coding them.

When you evaluate a vendor, this framing gives you the questions. How do they handle a model deprecation, and can they show you the eval results from the last one? What did their agent do the last time it hit an ambiguous result, and is there a log? Is the boundary between what the agent may read and what it may change enforced by the system, or by the prompt? Can they adapt to your environment without you standing up a dev team? What happens to your data, your configuration, and your runbooks if the relationship ends?

That last one matters more with a young vendor than an established one, and you should ask it plainly. A vendor that cannot answer it is not much safer than an internal build.

Where we land on this ourselves

This is the principle REAP is designed around, and it is the same one I wrote about in an earlier post on the difference between automation, AIOps, and agentic AI. REAP takes responsibility for the engine: reasoning across the network, the investigation loop, the model lifecycle, the evaluation, and the boundary between what the agent may read and what it may change. The customer owns the edges: their network context, their runbooks, their policy on what the agent may touch, and who approves what. Investigation should adapt as it goes. Change should be bounded, auditable, and approval-gated according to customer policy. We hold ourselves to that split because it is the only version of buying that does not ask a network team to give up the parts of their operation that only they can define.

If your team is part-way into a prototype, that is not wasted work. It is the clearest specification of what you need that you will ever write. Just do not let the speed of the first version decide the question. The decision is which parts of an agentic system you want your network organization to own and maintain for the next several years: the engine, or the context, policy, and judgment on top of it.

Frequently asked questions

Should a network team build or buy agentic AI for network operations?

For most enterprise and service provider network teams, buy the engine and own the edges. The vendor should own the reasoning loop, model lifecycle, evaluation, guardrails, and patching. The team should own its network context, runbooks, policy on what the agent may touch, and who approves what. Building the whole stack makes sense mainly for organizations that already have a platform engineering team, mature network APIs, a trusted source of truth, an AI evaluation discipline, and a security organization ready to govern agent access.

Why is an agentic AI prototype a poor estimate of what it costs to build?

The prototype skips the work that makes an agent safe in production: scoped credentials, a policy layer on what it may query and at what rate, audit logs, an evaluation suite that runs before and after every model change, and a plan for model deprecation. Those costs arrive after the demo and continue for the life of the system, and they usually land on the person who built it.

What is the difference between a prompt instruction and an enforcement boundary?

A prompt instruction such as "do not touch production" is a request the model may or may not follow. An enforcement boundary is a permission system that makes the action impossible. An agent that can reach network devices needs the second, and building it is different work from writing the prompt.

More from REAP

See REAP in action

Watch REAP reason through a live incident on your network.