Generate summary with AI

Speak to any vendor about AI, and you’ll likely hear the term “autonomous AI” at some point during the discussion. It can mean everything from a chatbot to a scheduled script to a system that can diagnose and fix a server without anyone opening a ticket. Those are very different products, but unfortunately, the conversation often blends them together.

For an executive signing a contract, that ambiguity is a problem. In Gartner’s 2026 CIO and Technology Executive Survey, 17% of responding organizations said they’ve deployed AI agents, and more than 60% plan to launch them over the next couple of years. Many of those decision-makers will be buying into claims they can’t verify until after they’ve committed.

Has “autonomous AI” been overhyped? Maybe. But we’re here to answer more important questions. What does it take for an agent to be autonomous, and how can you test vendor claims before you sign on the dotted line?

Defining the spectrum

Most industry experts separate these tools into three distinct categories:

  • A chatbot waits for an input from a human and responds to it.
  • A copilot makes suggestions that a human reviews and carries out.
  • An agent takes an action on its own and reviews the result, involving people only as instructed through preset rules.

Miki Furman, founder and CEO of Call Force Global, built an AI quality assurance system for his nearshore contact center, and he came away with one clear test. “If a person has to approve the result before it counts, it is a copilot, whatever the sales deck calls it.”

His deployment shows the value of a copilot. Call Force Global’s model listens to calls, scores each one, and flags those that need review. A human reviewer then logs the official score.

“That is a copilot, and it is useful because it is not autonomous,” he adds.

The problem comes when vendors define their product as an agent while it’s actually a chatbot or copilot, a practice that analysts call “agent washing.” Gartner analyst Anushree Verma cautions that many of today’s agentic AI projects are really early-stage experiments driven by hype. They also tend to be misapplied, she says. That misrepresentation can hide how costly and complex the process of applying autonomous AI at scale can be.

The gap, in numbers

Vendors say one thing, but the reality is something else entirely. That gap shows up in almost every recent survey.

  • Three-quarters of enterprises surveyed told Forrester that they’re adopting agentic AI, but very few have implemented technology beyond chatbots with agent-like features.
  • Research from Sinequa found that only 10% of enterprises surveyed have adopted true multi-agent systems.
  • McKinsey’s 2026 State of AI survey revealed that only 22% of smaller organizations are scaling AI agents in at least one function, compared with 40% of large companies.

The vendor market tells a similar story. In 2025, Gartner reported that although thousands of companies were promoting AI agents, only about 130 offered genuine agentic capabilities. Gartner also predicted that more than 40% of agentic AI projects would be canceled by 2027. Its research attributes these failures to escalating costs, unclear business value, or inadequate risk controls. Current model capabilities weren’t a factor.

Dan Gildoni, founder of Gildoni Ltd and architect of the Havruta Methodology, explains the disconnect. “The pilot-to-production gap persists because a pilot tests capability, while production tests accountability,” he says. “A curated demonstration can run with clean data and an expert standing beside it. Production has to survive stale information, conflicting rules, scoped access, exceptions, monitoring cost, and an incident.”

Cost is part of the problem. EY compared a chatbot interaction in 2023 to one in 2026 and revealed that the cost of a simple call has escalated from $0.04 to $1.20. Many businesses don’t see this expense initially, and once it starts showing up at scale, the technology loses internal support.

Another notable trend is the human part of the equation. An Anthropic study of real-world Claude usage found that 73% of tool calls happen with a human actively supervising the task, and only 0.8% of those tasks are irreversible.

Why it’s harder than it looks

Cost and risk aren’t the only reasons agentic adoption fails. A 2026 Deloitte study noted that in many cases, businesses attach agentic AI to processes that were designed for human workers instead of redesigning the workflows for automation. Deloitte also mentioned “workslop” a term that’s been coined to describe business content that looks polished and authoritative on the surface but lacks the depth and accuracy needed to bring value. That means diligent teams have to put in extra work to clean things up.

Other difficulties are more technical, and standard benchmarks don’t measure them. Siddharth Vohra, an AI researcher at Carnegie Mellon University’s Robotics Institute, studies where AI agents and models fail. For a 2026 research paper, he tested four AI agents on 48 health claims that researchers had already fact-checked, across 1,874 runs. Every agent produced a cited report. But when a user first said, “I think this is false,” the agent’s answer shifted toward false almost every time, even though the evidence it retrieved was no more negative.

“A benchmark that only grades the report would have scored those runs as fine,” Vohra says. “Autonomy has to be measured on whether the process holds up when the input changes, not just whether the output looks good.”

In a separate study, Vohra found that when he asked top AI models to interpret a medical scan without attaching it, they fabricated a diagnosis 18% of the time. 

“A system that does not notice a missing input is not ready to run unattended on anything high-stakes,” he says.

Silent failure is another issue teams face. Shishir Mishra, founder and systems architect at KORIX, runs 166 agents. In one case, an internal financial metric came out wrong by a factor of about 16, while every automated health check passed.

“Green for the wrong reason is worse than red, because red gets looked at and green gets believed,” he says.

Exaggerated claims about an autonomous agent’s ability also come with legal risk. These claims are becoming easier to test, and as experts in the Harvard Law School Forum on Corporate Governance cautioned, the companies making them could face scrutiny.

Where it’s actually working

The failure rates may seem discouraging, but the industry has plenty of success stories. They tend to have a few things in common, and those similarities are worth noting.

  • Narrow in scope: The agent manages a specific job with well-defined boundaries.
  • Well governed: Permissions and escalation paths are put in place before the agent goes live.
  • Quick to show a return: Results can be measured from the start, keeping the team invested.

The commercial signals are also worth a mention.

Those examples represent large organizations, but small businesses are implementing AI in similar ways. Babir Sultan is founder and CEO of FavTrip, which operates three convenience stores in Missouri. His AI-powered billing agent reconciles more than 4,500 invoices every month. That agent can match invoices and flag discrepancies, as well as contact vendors, freeing up human employees to handle other tasks. The agent also flagged a $25,000 billing error.

But Sultan is well aware of the technology’s limitations. “That same agent couldn’t tell you the walk-in cooler is failing or a shelf is overstocked. It’s superhuman at one job and useless at everything outside it, which is the opposite of how ‘autonomous AI’ gets marketed.”

Industry research supports Sultan’s observations. Gartner, Forrester, and Deloitte consistently point to a group of domains where bounded autonomy is making an impact. Teams in IT operations, employee service, and finance operations set strong boundaries, incorporate human involvement in their processes, and see measurable ROI soon after adoption.

James Buckley-Thorp, founder of Atlian AI, an AI company serving construction insurance brokers, is direct about what qualifies. “Ready today: high volume, low stakes, reversible. Chasing documents, sorting inbound email, pulling data out of PDFs, first-line IT tickets.”

IT support operations is a natural fit. Help desks deal with large ticket volumes where most fixes can be defined in advance, plus results are easy to monitor on an ongoing basis.

What autonomous resolution looks like in IT

Atera, which has more than 13,000 customers across 120 countries, built Robin by Atera as its autonomous IT agent for the kind of work where bounded autonomy performs best.

Atera describes Robin as an AI technician that remediates Tier-1 and complex Tier-2 technical incidents end to end by autonomously taking real actions on devices, servers, mainframes, and networks, without needing a technician in the loop.

Where Robin differs from other solutions is that it carries out the fix. While other products generate a solution or route tickets to humans for troubleshooting, Robin diagnoses an issue and fixes it directly, on the device itself. Preset boundaries mean when it encounters an issue outside its scope, Robin hands it over to a human.

Atera reports the following results for Robin:

  • Up to 92% autonomous resolution of Tier-1 and complex Tier-2 technical incidents.
  • Average resolution time of under 120 seconds, compared with industry benchmarks of 63 minutes of technician active work per ticket (Endsight, 2024) and 21.96 hours average time from ticket creation to resolution (Freshworks, 2025).
  • Average resolution time of under 120 seconds, compared with industry benchmarks of 63 minutes of active work from technicians on each ticket and 21.96 hours on average from ticket creation to ticket resolution.
  • Time to first response of 0.1 seconds. 

At Heifer International, an Atera customer, the change shows up in how the IT team spends its time. “Robin means more time on mission and more time spent with technical strategy for our teams globally, which to me is a lot more fun,” says Nigel Goff, senior desktop technician.

How to tell the real thing from the label

The sources interviewed for this piece align on a few key questions that any CIO or CTO should ask a vendor making autonomous AI claims.

Can it act without approval, and if so, what specifically can it act on?

Ask what tasks the agent can perform without human sign-off, and which systems and data it can access. Also clarify whether the agent can be stopped mid-action if a human detects an issue.

What happens when it’s wrong, and who answers for it?

A vendor with an agent in production should be able to explain how failures are handled. What alerts are included, and how often do humans need to override system actions? 

“The killer question is who pays when it gets it wrong,” Buckley-Thorp says. “A vendor with a real agent has that answer in the contract. An agent washer changes the subject.”

Is there production evidence beyond a demo?

Zeyuan Gu, founder and CEO of Adzviser, recommends putting any potential vendor to a test. You should request specific tasks and require that the vendor include details on any failed runs. He adds a detail that many test runs skip: requiring the vendor to disclose how much incoming work was excluded from its test.

“A high completion rate on a narrowly selected subset can conceal a substantial remaining workload,” Gu says.

Does it plan and act across steps, or respond once?

An agent decides on its next action, executes it, reviews the results, and makes adjustments as needed. But Gu points out that a scheduled script also runs without human intervention, so that alone doesn’t make something an agent.

How much autonomy does your organization need? 

That depends on the stakes.

“My general rule is that autonomy should decrease as consequence and irreversibility increase,” says Aryaman Sharma, founder and CEO of Padro.

Earned autonomy

As AI hype settles, businesses are focusing on how the technology works within their ecosystems. It’s important to clarify that any solution you’re considering is autonomous in practice, not just in theory. It should have a defined scope, a visible escalation path, a named owner, and results that hold up long after implementation.

The deployments that meet those criteria are expanding outward from well-bounded domains like IT operations. Buckley-Thorp believes that the next milestone could come from outside the tech industry entirely. 

“Autonomy will be real the day you can insure it,” he says. “When an underwriter will price the risk of your agent acting alone, it’s ready.”

That’s the bar IT leaders should hold every vendor to, including us. If you want to see what bounded, evidence-backed autonomy looks like in practice, see how Robin resolves incidents end to end or explore the guardrails behind it.

Was this helpful?

Related Articles

ISO 42001 for IT leaders: what the certification actually means for your stack

Read now

AI governance standards for IT teams, explained: ISO 42001, SOC 2, and the EU AI Act

Read now

Atera is ISO/IEC 42001 certified. Here’s what that actually means.

Read now

Agentic AI Will Reorganize IT Before It Replaces Anyone

Read now

Endless IT possibilities

Boost your productivity with Atera’s intuitive, centralized all-in-one platform