Article

Why AI agents fail: What we learned building them

Last updated 
Oct 8, 2026
7
 min read
Episode 
7
 min
Published 
Oct 8, 2026
7
 min read
Published 
Oct 8, 2026
7
min

We've already built a fair share of AI agents at Aubergine, and we've noticed a pattern that's easy to miss when you're still working with prototypes. Getting an agent to work isn't particularly difficult. Getting it to keep working in changing scenarios, users, models, and workflows around it is much harder.

More than 80% of AI projects fail, roughly twice the failure rate of non-AI IT projects. AI agents raise the stakes further, since they don't just generate an answer, they act on someone's behalf. Our experience building them has followed a familiar shape:

✓ We build something.
✓ We test it.
✓ It works.
✓ Someone else uses it. It still works.
✓ We start believing the system is ready.

Then someone uses it differently. And a failure that looks like a model problem turns out to be something else entirely.

The model isn't the whole story

Our early instinct, like most teams building AI agents, was straightforward: the model isn't good enough. The more of these systems we've built, the more we've found that failures usually trace back to how the system around the model is designed, not the model itself.

That's the thesis of this piece. AI agent failure is rarely a model problem. It's a systems problem, and systems problems need a different kind of AI architecture to solve them.

The problem with "it worked when we tested it"

Two lines show up constantly in early-stage AI agent projects:

  • "If it works for me, it will probably work for everyone else."
  • "I tested it and it's working, so it's tested."

This reflects a familiar development mindset: build, test, confirm, deploy. That approach holds up when systems behave predictably. AI agents don't behave that way, for a few reasons:

  • The underlying models are non-deterministic.
  • Model providers change how they expose and control behavior.
  • Prompts and workflows shift the results a model produces.
  • Hallucinations remain possible even in a well-tested pipeline.

So the question changes. It stops being "did we test it?" and becomes "have we designed a way to continuously know whether it still works?" That second question is what an evaluation layer is built to answer, and we come back to it below.

What building AI agents taught us about AI architecture

For a non-technical reader, AI architecture doesn't need to mean choosing between RAG, agents, fine-tuning, or a specific framework. Our experience building agents has pushed us to think about architecture more broadly. A working AI agent system needs answers to questions like:

  • Who is responsible for this task?
  • When should responsibility move somewhere else?
  • What information needs to move with it?
  • How do we know a step actually worked?
  • How do we know the entire system is still performing as expected?
  • What happens when the system doesn't know?

For us, AI agent architecture isn't only about how the AI reasons. It's about how the entire system behaves when things don't go as expected. That idea shapes everything in the next section.

The four things we now design into an AI agent from the start

As we've built and tested AI agents, four areas have proven particularly important: ownership, evaluation, handoffs, and verification.

A flowchart illustrating the steps to apply for a job, including researching, applying, and interviewing.
Design elementQuestion it answers
OwnershipWho is responsible for this task?
AI evalIs the system still working reliably?
HandoffsWhat happens when the task moves?
VerificationDid the agent actually do what we asked?

1. Ownership: who is responsible for what?

An AI agent shouldn't simply be told to figure it out. We define what a given agent is responsible for, and just as importantly, what it isn't.

A web agent, for instance, should handle web verification or web interaction because it has the tools and capabilities built for that task. It shouldn't take on responsibility for something outside its defined role.

Without clear ownership, the failure path looks like this:

Task → wrong agent → wrong action → wrong result

Ownership isn't something we define once. We also need to check whether the system routes tasks to the right agent, and that check is a job for evaluation.

2. AI eval: how do we know the system is still working?

Evaluation is the continuous measurement layer, not a test we run once at the end. An AI agent can work well against the examples we test initially and still fail once we change:

  • the model
  • the prompt
  • the workflow
  • the tools
  • the domain
  • the type of user request

That's why we rerun a defined set of evaluation cases whenever something changes, and check the results against a set performance threshold. This is also where we draw a clear line: verification checks a step, evaluation checks the system.

CheckWhat it asks
VerificationDid the agent fill this form correctly?
EvaluationAcross our representative test cases, does the agent reliably complete this class of task?

We don't try to test every imaginable scenario. Instead, we define what scenarios matter, what success means, what threshold is acceptable, and when the tests need to run again. In our experience, benchmarks somewhere in the 90 to 95% range work well, depending on how consequential the use case is.

 Flow diagram illustrating the steps to effectively use the autoresponder feature in a software application.

The goal of evaluation isn't to prove an agent never fails. It's to know how often it fails, where it fails, and whether that level of failure is acceptable for the workflow it sits in. We've written more on how to build this layer properly in What a 50% accuracy score taught us about building AI agents.

3. Handoffs: what happens when the task moves?

As AI agents take on more complex work, one agent rarely does everything. A task can move agent to agent, agent to system, or agent to human. The hard part isn't getting the next agent to perform its task. It's making sure context and responsibility move correctly with it.

Aubergine example: For one of our customers, we helped build a workflow where an agent interacted with the web and filled out a form, with a separate verification agent checking the result before it moved forward. If something needed fixing, the work could be sent back for correction rather than shipping as-is. 

Every handoff needs to define:

  • What is being handed over?
  • What context travels with it?
  • What is the receiving agent responsible for?
  • What counts as completion?
  • What happens if the next step fails?

Each handoff is a point where the system can lose context, responsibility, or accuracy. Designing for that risk upfront is part of the AI architecture, not a fix applied after something breaks.

A karaoke agent interface displaying song options, user selections, and a microphone icon for singing.

4. Verification: did the agent actually do what we asked?

Even when ownership is correct and the handoff works, the action itself can still fail. That's what verification catches.

The Kiara workflow is useful here too. An agent can complete an action successfully from its own point of view while still missing what the user actually needed. A verifier adds a layer of checking before the workflow moves on.

We don't advocate verification everywhere. It adds complexity, so it earns its place on actions where the consequences of getting it wrong are high. The more consequential the action, the less comfortable we are asking the same system that performed it to confirm on its own that it worked.

What AI failure looks like when we skip these steps

SkipResult
OwnershipThe wrong agent can take responsibility for a task.
EvaluationRegressions can enter the system unnoticed.
Handoff designContext or responsibility gets lost mid-task.
VerificationAn incorrect action gets accepted as complete.

All four gaps lead to the same outcome: the system can technically run while failing to deliver what the user actually expected.

This gap is expensive at scale. Despite roughly $30 to 40 billion in enterprise generative AI spending, MIT's 2025 State of AI in Business report found that 95% of pilots showed no measurable P&L impact within six months, tracing most of the shortfall to workflow and integration gaps rather than model quality. Gartner makes a similar point about agentic AI specifically, predicting that over 40% of agentic AI projects will be canceled by the end of 2027 due to rising costs, unclear business value, and inadequate risk controls, none of which are model problems.

The business consequence is simple. Users tolerate an AI system while it works. Once they run into a hallucination, or even a smaller reliability failure, they tend to stop using it. Small AI failures are enough to kill adoption.

So when do we stop testing?

We're never going to test every possible interaction an AI agent might face. Our approach isn't to chase infinite coverage. It's to set a defined threshold based on the use case:

Use case → critical scenarios → evaluation set → acceptable threshold → repeated passes

Once the system consistently meets that threshold, we have a defensible basis for moving forward. We focus evaluation where it adds the most signal, then carry those tests forward as part of ongoing regression and release checks.

The questions we would ask any AI agent vendor before a pilot

Before starting an AI agent pilot, these are the questions worth asking a vendor, not just "how do you build an AI agent."

CategoryQuestions to ask
OwnershipWhich agent owns each critical task? How do you prevent responsibilities from overlapping?
EvaluationWhat scenarios will you evaluate before the pilot starts? What performance threshold defines success? What happens to the evaluation when you change the model, prompt, or workflow?
HandoffsWhere are the handoffs in the workflow? What context and responsibility move at each handoff?
VerificationWhich critical actions get independently verified? What happens when verification fails?
Failure handlingWhat does the agent do when it doesn't know the answer or can't complete the task?

What we've changed about how we approach AI agent pilots

We deliberately try to find where an AI agent will fail before that failure becomes someone else's production problem. That means being willing to discover, during a pilot, that:

  • the use case isn't ready
  • the workflow isn't reliable
  • the handoffs aren't working
  • the evaluation threshold isn't being met
  • the AI architecture needs to change

This is a more credible position than promising a successful deployment every time. A pilot that surfaces a real weakness early has done its job, even if the honest answer is "not yet."

Conclusion: the model is only one part of the system

We've learned that "which model should we use" is only one of the questions that matters when building AI agents. The more important ones are:

  • Who owns the task?
  • Where does responsibility move?
  • How is the result verified?
  • How do we know the system continues to work?
  • What happens when it doesn't?

A capable model inside a poorly designed system can still produce an unreliable AI agent. The intelligence of the model matters less than the discipline built around it.

Ready to pressure-test your AI agent architecture?

Reliable AI agents aren't built by picking the right model and hoping for the best. They're built by designing for failure, evaluating continuously, and finding the weak points before your users do. Our AI Strategy Sprint helps teams pressure-test their AI opportunities, architecture, and path to production, so they can move forward with confidence and unlock the full value of AI. Let’s talk.

FAQs

What's the biggest reason AI agent pilots fail?

Most failures trace back to how the system around the model is designed, not the model itself. Missing ownership, no continuous evaluation, weak handoffs, and no verification layer are the recurring causes we see.

How is an AI agent different from a chatbot or a copilot when it comes to reliability?

A chatbot generates text a person reviews before acting on it. An AI agent often takes the next step itself, whether that's filling a form, calling an API, or handing a task to another system. That extra autonomy means a failure has a real-world consequence, not just a wrong sentence.

Do we need a dedicated AI eval framework before running a pilot, or can we add it later?

Build it before the pilot, even in a lightweight form. Without a baseline evaluation set, you have no way to tell whether a change to the model, prompt, or workflow made the system better or worse.

How many AI agents should handle a single workflow?

There's no fixed number. The right question is whether each agent has a clearly scoped responsibility and whether handoffs between agents are explicitly designed, not whether you're using one agent or five.

What's a reasonable success threshold for an AI agent before it goes into production?

It depends on how consequential the task is. For lower-stakes, internal workflows, a lower threshold with human review may be fine. For anything customer-facing or financially consequential, we typically look for evaluation results in the 90 to 95% range before promoting an agent to production.

Should every action an AI agent takes be verified?

No. Verification adds complexity and cost, so it's worth reserving for actions with real consequences if they go wrong, such as sending something externally, updating a system of record, or completing a transaction.

Authors

Aaditya Brahmbhatt

Associate Technical Lead
Associate Tech Lead who enjoys tackling the infinite possibilities that Artificial Intelligence and software engineering can bring to life. An expert in working with LLMs and creating new AI tools.

Podcast Transcript

Episode
 - 
7
minutes

Host

No items found.

Guests

No items found.

Have a project in mind?

Read