Skip to content
Customer Experience agents
G2 Agent EvaluationsInsightsPublished August 2026 · Updated September 2026

We put 10 AI support agents to work at the same company. Here is what we learned.

“AI can now do tier 1 support well, but performance varies up to 2x as we noticed some of the AI agents sold today struggle to follow a basic one-page policy.”
Alexis ZhengChief Product and Technology Officer, G2

G2 runs each vendor's live agent inside a simulated company with a support policy, customer database, working tools, and customers with real problems. Every agent runs the same support desk under the same rules, and we record every message and tool call. These findings are about the market, not any one vendor.

G2 interviews and conversations with more than 20 buyers of AI customer service agents, conducted from July to September 2026, add a view of what it takes to put these agents into production. The interview findings below reflect buyers’ experiences and expectations, rather than scores from the simulated evaluation.

Agents Evaluated
10
Scenarios per Agent
46
Conversations Recorded
~700
Live Tools in the Environment
38
Dimensions Scored per Task
4

AI agents are ready for much of Tier 1 support

The average support representative stays in a seat for less than a year, so companies retrain endlessly or hand the queue to business process outsourcers, then live with unwanted ramp time and uneven quality. Seasonality makes the problem worse: retailers can grow support staff by as much as five times for the holidays, hiring in September and letting people go in January.

This is the work AI agents absorb best: repetitive, policy-bound, and spiky. Teams getting value from agents can redeploy people to complex and proactive work instead of rehiring and retraining every year.

Buyers also want that capacity gain to show up in operating results. The Head of Customer Experience at Flipkartdescribed starting at 60–65% efficiency and building toward roughly 85%. That is one buyer's experience, not a universal ceiling or a direct comparison with G2's task accuracy scores. Interviewees assess resolution rates alongside satisfaction on resolved conversations and the effect on support staffing.

The bottom line every institution looks at: a real reduction of FTEs without reducing satisfaction. Until that materializes, it's just numbers on paper.

CX Transformation Executive, large Canadian financial services company

The same job produced a two-to-one performance spread

Every agent received identical tasks, tools, and policies inside G2's simulated company. The best Overall Score was more than double the lowest, while task accuracy ranged from 64% to 88%. Measuring actual performance helps buyers look beyond vendor claims.

We expect agents to improve quickly and take on more complex work involving payments, sensitive private data, and multiple integrations across a complete CX workflow.

Score range across evaluated agents
Accuracy64%–88%
Policy compliance56%–92%
Relevance82%–100%

Percentages show the lowest and highest observed dimension score in this evaluation.

Success rates fall as process complexity rises

Nearly every agent stayed on topic: relevance ranged from 82% for the lowest-performing agent to 100% for Retell. Larger gaps emerged in policy compliance, which ranged from 56% to 92% for Tidio.

Recurring failures included answering before checking the customer record, escalating tickets the agent could have resolved, and taking the wrong action while reporting success.

“The gap we see is not the models, but in the harness. Products that let the agent work straight from the policy and tools scored at least 20% higher than the ones that route everything through their own workflow builder.”
Frankie LiVP Data Science, G2

Buyers expand autonomy in stages

The cost of a wrong action shapes how buyers roll out agents. Interviewees described human checks for refunds, payments, and account changes, including at companies using Sierra and Decagon. A senior customer care leader at an ITSM software vendor described progressing from drafts to routine answers, then low-risk tasks, before real transactions.

Everything informational is autonomous. Everything transactional needs a human validating before action, AI is mostly used for recommendation.

CX Transformation Executive, large Canadian financial services company

For buyers, policy compliance is therefore a deployment decision as well as a benchmark score: define which actions the agent can take alone, which require approval, and when it must hand off to a person.

Integrations decide how much work an agent can do

Salesforce and Jira came up as must-have integrations in separate buyer conversations. One buyer reported dropping a vendor over a three-hour Salesforce sync and walking away from another when the integration was quoted at $100,000. Buyers need to test the connection to their own systems, including sync speed and implementation cost, before committing.

Without integration to our membership platform, the AI agent can't do anything except surface FAQs. It has to connect to the system that drives our business.

Head of Customer Experience, PureGym

The category is moving quickly

Since G2 began evaluating CX agents, the market has continued to consolidate, including Fin being acquired by Salesforce for $3.6 billion and Forethought by Zendesk. Sierra, which is under evaluation, and Decagon, which declined evaluation, are differentiating through a forward-deployed engineering strategy rather than a self-serve model. Both are also developing frontier skills intended to handle longer-running tasks.

Pricing and implementation determine the return

Interviewees differed on paying per resolution versus per conversation. Some preferred paying only for solved issues; others valued predictable conversation costs. Usage-based credits and long commitments drew objections. Buyers also described how changes in vendor ownership affected their plans and contract expectations.

Per resolution puts me in a position where the better I configure the agent, the more I pay.

Customer Care Director, racket-sports marketplace

Time to value matters alongside price. Buyers weighed a six-month setup resolving 50% of tickets against a one-month setup resolving 30%, and wanted clarity on whether implementation was self-serve or required the vendor's engineers. Ashley Wagner, VP of Customer Experience at Blackthorn, described how both vendor support and cost per resolution affected the economics of deflection.

Selection is only the start of testing. One buyer described running two agents side by side on 1% of traffic for three to six months before choosing. Buyers plan to keep monitoring after launch because agent quality can change over time. A pilot should establish both the service outcomes and the ongoing effort needed to sustain them.

Buyers want independent evidence

G2's August 2026 AI agent buyer study of more than 1,000 respondents found that buyers now count AI search (17%) and evaluation sites (15%) among their key research sources. Seventy-three percent trust open research or independent test labs most, compared with 11% for vendor-published tests.

Buyers' most requested benchmarks include task completion (33%), safety and robustness (31%), and tool-calling reliability (29%).

The interviews reinforce that demand for evidence. Buyers want a third party that has actually run the agents, plus customer references they can speak with independently of the vendor. As PureGym's Head of Customer Experience put it: “I'd trust the numbers, but I'd ask to speak to one of their customers, ideally without the vendor present.”

This is honestly a game changer... a reputed third party that has actually gone and implemented it, seen results, and knows what to do.

Sanjay Kini, Head of Customer Success, Birdeye

Use independent evaluations to build a shortlist, then test the finalists against your own policies, integrations, approval rules, and cost model. The strongest benchmark result still needs to translate into a deployment that works for your customers and team.

Next CX evaluation: November 2026

G2 plans to refresh the Customer Experience evaluation every quarter as agents and the benchmark evolve.