Skip to content
Customer Experience agents
G2 Agent EvaluationsInsightsAugust 2026

We put 10 AI support agents to work at the same company. Here is what we learned.

“AI can now do tier 1 support well, but performance varies up to 2x as we noticed some of the AI agents sold today struggle to follow a basic one-page policy.”
Alexis ZhengChief Product and Technology Officer, G2

G2 runs each vendor's live agent inside a simulated company with a support policy, customer database, working tools, and customers with real problems. Every agent runs the same support desk under the same rules, and we record every message and tool call. These findings are about the market, not any one vendor.

Agents Evaluated
10
Scenarios per Agent
46
Conversations Recorded
~700
Live Tools in the Environment
38
Dimensions Scored per Task
4

AI agents are ready for much of Tier 1 support

The average support representative stays in a seat for less than a year, so companies retrain endlessly or hand the queue to business process outsourcers, then live with unwanted ramp time and uneven quality. Seasonality makes the problem worse: retailers can grow support staff by as much as five times for the holidays, hiring in September and letting people go in January.

This is the work AI agents absorb best: repetitive, policy-bound, and spiky. Teams getting value from agents can redeploy people to complex and proactive work instead of rehiring and retraining every year.

The same job produced a two-to-one performance spread

Every agent received identical tasks, tools, and policies inside G2's simulated company. The best Overall Score was more than double the lowest, while task accuracy ranged from 64% to 88%. Measuring actual performance helps buyers look beyond vendor claims.

We expect agents to improve quickly and take on more complex work involving payments, sensitive private data, and multiple integrations across a complete CX workflow.

Score range across evaluated agents
Accuracy64%–88%
Policy compliance56%–92%
Relevance82%–100%

Percentages show the lowest and highest observed dimension score in this evaluation.

Success rates fall as process complexity rises

Nearly every agent stayed on topic: relevance ranged from 82% for the lowest-performing agent to 100% for Retell. Larger gaps emerged in policy compliance, which ranged from 56% to 92% for Tidio.

Recurring failures included answering before checking the customer record, escalating tickets the agent could have resolved, and taking the wrong action while reporting success.

“The gap we see is not the models, but in the harness. Products that let the agent work straight from the policy and tools scored at least 20% higher than the ones that route everything through their own workflow builder.”
Frankie LiVP Data Science, G2

The category is moving quickly

Since G2 began evaluating CX agents, the market has continued to consolidate, including Fin being acquired by Salesforce for $3.6 billion and Forethought by Zendesk. Sierra, which is under evaluation, and Decagon, which declined evaluation, are differentiating through a forward-deployed engineering strategy rather than a self-serve model. Both are also developing frontier skills intended to handle longer-running tasks.

Buyers want independent evidence

G2's August 2026 AI agent buyer study of more than 1,000 respondents found that buyers now count AI search (17%) and evaluation sites (15%) among their key research sources. Seventy-three percent trust open research or independent test labs most, compared with 11% for vendor-published tests.

Buyers' most requested benchmarks include task completion (33%), safety and robustness (31%), and tool-calling reliability (29%).

Next CX evaluation: November 2026

G2 plans to refresh the Customer Experience evaluation every quarter as agents and the benchmark evolve.