Build or buy AI agents?
A straight framework for deciding whether to build agents in-house or buy them operated — including the costs that are routinely missed, and the questions worth asking any provider, us included.
The short answer
Build AI agents in-house when the workflow is core to how you compete, you already employ engineers who have run models in production, and you can fund that team past the first release. Buy an operated service when the workflow is important but not differentiating, when you need it working in weeks rather than quarters, or when nobody on staff owns production AI today. Most organizations underestimate the second cost — building an agent is a matter of weeks, while operating one is indefinite.
When each answer is right
Building is the correct choice more often than vendors admit. These are the conditions under which each path actually works.
Build in-house
When building is right
Buy operated
When buying is right
The four costs that arrive later
These are not arguments against building. They are the line items that turn a confident build decision into a stalled one, and they are worth pricing honestly before you choose.
The work starts at the demo, not before it
A working prototype is now a matter of days. What follows is the actual system: evaluation harnesses, approval workflows, retry and failure handling, observability, permissions, and the operational habit of checking whether output is still good. Teams budget for the prototype and inherit the rest.
Agents decay without anyone touching them
Conventional software is stable until changed. Agents are not. Model versions shift behavior, upstream data drifts, prompts that worked stop working, and vendor APIs change. Quality erodes silently, and it erodes fastest when nobody is measuring.
The talent is scarce and does not stay
Engineers who can operate production agents are expensive, in demand, and rarely satisfied maintaining a single internal workflow. A team of one is a single point of failure; a team of three is a material line item for a workflow that is not your product.
Governance arrives late and costs the most
Approval gates, audit trails, data isolation, and permissioning are easy to design in and painful to retrofit. They are also the first things asked about in diligence — and retrofitting them after an agent is live usually means rebuilding how it handles data.
Six questions for any provider
These work on us as well. If a provider cannot answer them with enforced mechanisms rather than intentions, that is the answer.
Who is accountable if the agent stops working correctly six months from now?
How is agent quality measured after launch, on what cadence, and against what reference set?
What actions can the agent take without a human approving them first?
Where does our data live, and can the system run inside our own cloud account?
If we ended this relationship, what would we keep, and in what state?
How is one agent prevented from reading another agent's data?
Build vs buy, answered
Should we build or buy AI agents?
Build when the workflow is genuinely differentiating, you already employ engineers who have operated models in production, and you can fund that team beyond the first release. Buy an operated service when the workflow matters but is not a competitive differentiator, when you need production in weeks rather than quarters, or when no one on staff currently owns AI in production. The decisive factor is usually not whether you can build an agent, but whether you can commit to operating it indefinitely.
Why do AI agent pilots fail to reach production?
Most pilots fail after the demo rather than during it. A prototype proves the model can do the task; production requires evaluation, approval workflows, error handling, observability, permissions, and continuous monitoring for drift. Teams commonly scope and fund the prototype, then discover the operating system around it was never budgeted, and the pilot stalls with nobody accountable for finishing it.
What does it actually cost to run AI agents in production?
Production cost has three parts, and inference is usually the smallest. There is model and infrastructure spend, which scales with usage; engineering time to monitor, evaluate, and repair agents as models and data shift; and the governance layer — approval workflows, audit trails, and isolation. Organizations that budget only the first are the ones surprised at month six.
How long does it take to get an AI agent into production?
A well-scoped first agent can be live in production in one to two weeks under a forward-deployed model, where engineers build against the real workflow rather than a specification. Timelines extend when the agent needs several system integrations, when data access requires security review, or when approval workflows must be designed alongside the agent. An internal build typically takes considerably longer because hiring precedes building.
What should we ask an AI agent vendor during diligence?
Ask who is accountable when the agent degrades, how quality is measured after launch and on what cadence, what actions the agent can take without human approval, where data resides and whether it can run in your own cloud, how one agent is prevented from accessing another's data, and what you retain if the relationship ends. Answers that describe enforced platform mechanisms are stronger than answers describing policy or intent.
Can we start with a managed service and bring AI agents in-house later?
Yes, and it is a reasonable sequence. Running a managed deployment first tells you which workflows actually justify agents and what operating them really requires — knowledge that makes a later in-house build far more likely to succeed. What matters is that the agents run somewhere you control or can take control of, so the option remains open rather than theoretical.
Talk it through before you commit
If building is the right call for your workflow, we will say so. If it is not, we will show you what operating it actually involves.