A pilot that never leaves experimentation costs you the time testing it. An agent that ships and gets pulled costs you the pilot, the confidence of whoever championed it, and whatever business it impacted on the way out.
Sinch's 2026 research found that 74% of enterprises have rolled back or shut down a live AI customer communications agent. In professional services the figure is 85%, the highest of any sector they measured, and higher than technology at 66%.
One finding in that set is worth pausing on. The rollback rate rises to 81% among organizations with mature governance. Governance isn't failing to catch these. It is catching them, which is what a rollback is.
The causes split into two categories, and they need different fixes
Among organizations that reported a governance-failure rollback specifically, Sinch ranks the causes: exposure of personal or customer data at 31%, hallucination or brand risk at 22%, and lack of auditability at 16%, meaning nobody could diagnose what went wrong.
Sinch separates those into two groups.
Exposure, context loss and audit gaps sit in the infrastructure layer. They are failures the platform around the agent should catch before anything reaches the agent at all. Careful agent instructions, and a context architecture you can edit independently of the LLM, are what let you refine a workflow until it holds.
Hallucination and off-brand responses are a model and prompting problem, and Sinch is explicit that no amount of infrastructure investment prevents them. Better context and clearer instructions reduce how often it happens and make the gap visible when it does. They don't eliminate it, and anything that tells you otherwise is selling something. It is surprising that an industry earning this much hype and money is allowed to be so accident prone without consequence, but narrowing the opportunity for an agent to plainly make things up is part of the maintenance.
That distinction changes what you do. The first group is a configuration decision you make once. The second is a standing condition you manage, which means somebody keeps reading the output.
What testing left open
Testing establishes that the agent can do the task. Three other things decide whether it survives.
What it can reach. In a test it saw a folder you chose. In production it carries whatever permissions its account has, which in a small firm is often everything.
What it does when it doesn't know. Ask a model something outside its context and you get a confident answer where you wanted a question. In testing you knew the right answer. A client doesn't.
Who reads the output. In testing, you did.
The cost nobody prices in
Sinch found that 84% of AI communications engineering teams spend at least half their time building guardrails and safety controls. 35% spend most of their time there instead of on the next thing.
At enterprise scale that is a staffing line. At your size it is you, on a Thursday, and it is the reason a second agent takes longer to ship than the first one did if you skipped this on the first one.
What is different at twenty people
These numbers describe enterprises, and your situation differs in ways worth naming rather than assuming.
You can see all the output. At two thousand people an agent's output disappears into a volume nobody can read, so monitoring has to be built before you can trust anything. At twenty, somebody can read everything one agent produced this week in a few minutes. That is a real advantage and it expires as volume grows, so it is worth using while it lasts.
You also carry more concentrated risk per incident. An enterprise rollback is embarrassing. A small firm exposing one client's material to another may lose the client, and there is no communications department between you and that conversation.
Related questions
Why do AI agents fail in production?
Sinch's 2026 research found 74% of enterprises have rolled back a live AI customer communications agent. Among those reporting a governance-failure rollback, the leading causes were exposure of personal or customer data at 31%, hallucination or brand risk at 22%, and lack of auditability at 16%.
Which industry rolls back agents most?
Professional services, at 85%, the highest sector in Sinch's 2026 research. Technology was lowest at 66%.
Why is the rollback rate higher in organizations with mature governance?
It reaches 81% there. Governance is what detects the failure, so a higher rollback rate reflects better detection as much as worse agents.
Can better prompting prevent hallucination?
It reduces frequency and makes gaps visible. Sinch classes hallucination and off-brand responses as model and prompting problems that infrastructure investment doesn't prevent, which is why client-facing output keeps a human reader.
What should a small firm decide before shipping an agent?
The specific records it may read, the specific place it may write, and what it does when information is missing. Then who reads its output, and for how long.
How we read this
The two-category split is the part worth carrying, because it tells you which problems have an end and which don't.
Scope and auditability are decisions. You make them once, they stay made, and the work is bounded. Hallucination is a condition. You reduce it, you catch it, and you keep a person between it and a client, and that arrangement is permanent for anything where being wrong is expensive.
Most of the disappointment we see comes from treating the second like the first: shipping something, watching it work, removing the human, and discovering months later what it got wrong quietly. The review load should shrink as you learn what an agent is reliable at. It should not reach zero on anything client-facing, which is the same position the portfolio takes everywhere: the system carries the work and your judgment stays in the loop.
The 85% in professional services is the number we would sit with longest. That is your sector, at enterprise scale, mostly failing at this. It is also a group with more budget and more staff than you, which suggests the constraint isn't resources.
What you can do this week
Take the agent or automation you are closest to shipping and write down two things.
What exactly can it read, and what exactly can it write to. If the honest answer is "whatever my account can," that is the scope decision and it is most of the task. Don’t give the agent access to everything and don’t give the agent long instructions.
Then: what does it do when it doesn't know? What is the consequence of being a little wrong versus very wrong?
Expect the scope decision to break something. Narrow the permissions and you will usually find the agent can no longer do part of the job, because the records it needs live in three systems. That isn't a reason to widen the scope again. It is the clearest map you will get of what needs consolidating first, and it is worth more than the agent was.
Working together
Flow State Found works with a limited number of businesses to make their best work their baseline. Most firms we speak with had an agent that worked for a month and then quietly stopped being right. We build the maintenance in from the start, because an agent is a process you own rather than a purchase you made.
We take on limited engagements, so it starts with a conversation.
Start a conversation