Skip to main content
Decision framework

Build or buy AI agents?

A straight framework for deciding whether to build agents in-house or buy them operated — including the costs that are routinely missed, and the questions worth asking any provider, us included.

The short answer

Build AI agents in-house when the workflow is core to how you compete, you already employ engineers who have run models in production, and you can fund that team past the first release. Buy an operated service when the workflow is important but not differentiating, when you need it working in weeks rather than quarters, or when nobody on staff owns production AI today. Most organizations underestimate the second cost — building an agent is a matter of weeks, while operating one is indefinite.

The decision

When each answer is right

Building is the correct choice more often than vendors admit. These are the conditions under which each path actually works.

Build in-house

When building is right

The workflow is a genuine source of competitive advantage, not overhead.
You already have engineers who have shipped and operated ML or LLM systems in production — not just prototyped them.
You can fund that team through year two, past the point where the initial release stops being interesting.
Your data or regulatory position makes any external processing genuinely impossible.

Buy operated

When buying is right

The workflow matters commercially but is not what differentiates you.
You need it running in weeks, and a hiring cycle alone would take longer.
No one currently owns AI in production, and the roadmap has no room to add it.
You want the operating burden — evaluation, drift, incidents — to be someone else's obligation.
What gets underestimated

The four costs that arrive later

These are not arguments against building. They are the line items that turn a confident build decision into a stalled one, and they are worth pricing honestly before you choose.

01

The work starts at the demo, not before it

A working prototype is now a matter of days. What follows is the actual system: evaluation harnesses, approval workflows, retry and failure handling, observability, permissions, and the operational habit of checking whether output is still good. Teams budget for the prototype and inherit the rest.

02

Agents decay without anyone touching them

Conventional software is stable until changed. Agents are not. Model versions shift behavior, upstream data drifts, prompts that worked stop working, and vendor APIs change. Quality erodes silently, and it erodes fastest when nobody is measuring.

03

The talent is scarce and does not stay

Engineers who can operate production agents are expensive, in demand, and rarely satisfied maintaining a single internal workflow. A team of one is a single point of failure; a team of three is a material line item for a workflow that is not your product.

04

Governance arrives late and costs the most

Approval gates, audit trails, data isolation, and permissioning are easy to design in and painful to retrofit. They are also the first things asked about in diligence — and retrofitting them after an agent is live usually means rebuilding how it handles data.

Diligence

Six questions for any provider

These work on us as well. If a provider cannot answer them with enforced mechanisms rather than intentions, that is the answer.

Who is accountable if the agent stops working correctly six months from now?

How is agent quality measured after launch, on what cadence, and against what reference set?

What actions can the agent take without a human approving them first?

Where does our data live, and can the system run inside our own cloud account?

If we ended this relationship, what would we keep, and in what state?

How is one agent prevented from reading another agent's data?

Common questions

Build vs buy, answered

Should we build or buy AI agents?

Build when the workflow is genuinely differentiating, you already employ engineers who have operated models in production, and you can fund that team beyond the first release. Buy an operated service when the workflow matters but is not a competitive differentiator, when you need production in weeks rather than quarters, or when no one on staff currently owns AI in production. The decisive factor is usually not whether you can build an agent, but whether you can commit to operating it indefinitely.

Why do AI agent pilots fail to reach production?

Most pilots fail after the demo rather than during it. A prototype proves the model can do the task; production requires evaluation, approval workflows, error handling, observability, permissions, and continuous monitoring for drift. Teams commonly scope and fund the prototype, then discover the operating system around it was never budgeted, and the pilot stalls with nobody accountable for finishing it.

What does it actually cost to run AI agents in production?

Production cost has three parts, and inference is usually the smallest. There is model and infrastructure spend, which scales with usage; engineering time to monitor, evaluate, and repair agents as models and data shift; and the governance layer — approval workflows, audit trails, and isolation. Organizations that budget only the first are the ones surprised at month six.

How long does it take to get an AI agent into production?

A well-scoped first agent can be live in production in one to two weeks under a forward-deployed model, where engineers build against the real workflow rather than a specification. Timelines extend when the agent needs several system integrations, when data access requires security review, or when approval workflows must be designed alongside the agent. An internal build typically takes considerably longer because hiring precedes building.

What should we ask an AI agent vendor during diligence?

Ask who is accountable when the agent degrades, how quality is measured after launch and on what cadence, what actions the agent can take without human approval, where data resides and whether it can run in your own cloud, how one agent is prevented from accessing another's data, and what you retain if the relationship ends. Answers that describe enforced platform mechanisms are stronger than answers describing policy or intent.

Can we start with a managed service and bring AI agents in-house later?

Yes, and it is a reasonable sequence. Running a managed deployment first tells you which workflows actually justify agents and what operating them really requires — knowledge that makes a later in-house build far more likely to succeed. What matters is that the agents run somewhere you control or can take control of, so the option remains open rather than theoretical.

Talk it through before you commit

If building is the right call for your workflow, we will say so. If it is not, we will show you what operating it actually involves.