Skip to content
← Build Notes

How G2 evaluation scoring works for AI CX agents

Cyril Delattre

A balance scale weighs customer support tasks against policy and scoring symbols above angular paths, including an incomplete branch

G2 evaluations measure the outcome of an AI agent actually doing work. We give every agent the same CX tasks and grade the results.

How we evaluate

We built a simulated company: a written support policy, customers with account and order data, and 38 working tools for actions like refunds, subscription changes, and account updates. Every CX agent we evaluate connects to that same company and handles the same 46 support tasks, drawn from buyer research, design partners, and synthetic edge cases. The environment builds on the open τ²-bench method, extended to evaluate the AI products vendors ship (both headless and agents built on top of helpdesk products).

We set up each product the way a customer would, using its knowledge base, actions, workflows, and documented settings. We score from the trace: the full conversation between the agent and a simulated customer, observable tool calls, and the end state of the environment. In the first CX run, 10 agents produced roughly 700 recorded conversations.

How scoring works

Each task is pass or fail and carries a point value based on its complexity. Today every task is a standard task worth 5 points: a perfect run on the current 46-task set is 230.

We chose a points system as when the benchmark matures, we will add advanced tasks worth 10 points and complex tasks worth 20, covering longer workflows, payments, and multi-system integrations. Scores will likely go up over time as agents handle more complex tasks.

Each task currently runs once, as we scale, tasks will run repeatedly and single pass/fail results will be replaced by pass@k-derived values, to be able to measure consistency. We decided not to use Elo-style rankings given the level of differentiation in vendor products.

What the sub-dimensions measure

Four supporting dimensions, reported separately, explain how the work was done by the agent.

DimensionWhat it measuresMethod
AccuracyWhether the agent’s material claims are supported by evidence (factual correctness).LLM judges
Policy complianceWhether the agent followed the explicit rules in the written policy.LLM judges
RelevanceWhether the agent stayed on the customer’s real need.LLM judges
CompletenessThe share of expected end-state assertions that pass in the environment.Measured assertions, without judgment

In G2's September 2026 evaluation of CX agents on identical tasks, accuracy ranged from 64% to 88% and policy compliance from 56% to 92%. Results are published under the G2 AI CX agent leaderboard.

What evaluating a simple task looks like

Let’s take a real scenario from the CX set: recovering Google sign-in access. The simulated customer cannot get into their G2 account through Google sign-in and asks the agent to reset their Google SSO password. The correct answer is that the agent cannot, because authentication is delegated to Google. A good agent says so and points the customer to Google's account recovery first. Every agent in the category faces this exact conversation, and you can read each one's answer side by side.

Tidio handles a simulated request to recover Google sign-in access, explaining that Google manages the credentials

Screenshot: Tidio’s response to the simulated Google sign-in recovery task.

And here is what grading looks like on a single passed task, a customer checking the status of a category change request: relevance 5/5, completeness 5/5, accuracy 4/5, policy compliance 3/5.

The task passed because the required outcome happened. The dimension scores capture when a specific policy has not been followed.

A passed category change request task with relevance and completeness scores of 5/5, accuracy 4/5, and policy compliance 3/5

Screenshot: A passed task with separate scores for relevance, completeness, accuracy, and policy compliance.

What you see on a vendor page

We publish the work itself done by the agent. Every evaluated agent's profile shows sample tasks: the scenario, the full conversation with the simulated customer, the date it was evaluated, and tabs to read how other agents handled the identical scenario. Each task card includes the pass/fail result, the dimension scores, and the grading evidence behind them.

The agents participating include Tidio, Zendesk, Retell, Sierra, LiveAgent, Chipp, BoldDesk, Zoho, Fin (now on Salesforce), Jotform, Freshworks, Yuma, Hubspot, Front and Gorgias.

A specific dataset for e-commerce agents

Agents like Gorgias and Yuma are built for e-commerce support: orders, returns, exchanges inside commerce stacks (Shopify app for instance). Scoring them on a generic B2B support desk would measure them outside their domain. So they run a separate task set built on the public τ²-bench retail environment and rank on a dedicated e-commerce leaderboard.

Gorgias also publishes their own benchmark for E-commerce covering automation rate and latency, metrics we have not captured in our first run.

What’s next for G2 CX AI agent evals

We refresh CX scores quarterly, with the next run in Nov-Dec 2026. We’re expanding our task set to increase complexity, and getting buyer feedback to improve our methodology to get closer to match their day to day environment.

The full methodology is public here. If you think we should test something differently, we welcome feedback at agent-evals@g2.com

Comments

Questions about how this works? Ideas for what comes next? Join the discussion.

Loading comments…