How G2 evaluation scoring works for AI CX agents

G2 evaluations measure the outcome of an AI agent actually doing work. We give every agent the same CX tasks and grade the results.
How we evaluate
We built a simulated company: a written support policy, customers with account and order data, and 38 working tools for actions like refunds, subscription changes, and account updates. Every CX agent we evaluate connects to that same company and handles the same 46 support tasks, drawn from buyer research, design partners, and synthetic edge cases. The environment builds on the open τ²-bench method, extended to evaluate the AI products vendors ship (both headless and agents built on top of helpdesk products).
We set up each product the way a customer would, using its knowledge base, actions, workflows, and documented settings. We score from the trace: the full conversation between the agent and a simulated customer, observable tool calls, and the end state of the environment. In the first CX run, 10 agents produced roughly 700 recorded conversations.
How scoring works
Each task is pass or fail and carries a point value based on its complexity. Today every task is a standard task worth 5 points: a perfect run on the current 46-task set is 230.
We chose a points system as when the benchmark matures, we will add advanced tasks worth 10 points and complex tasks worth 20, covering longer workflows, payments, and multi-system integrations. Scores will likely go up over time as agents handle more complex tasks.
Each task currently runs once, as we scale, tasks will run repeatedly and single pass/fail results will be replaced by pass@k-derived values, to be able to measure consistency. We decided not to use Elo-style rankings given the level of differentiation in vendor products.
What the sub-dimensions measure
Four supporting dimensions, reported separately, explain how the work was done by the agent.
| Dimension | What it measures | Method |
|---|---|---|
| Accuracy | Whether the agent’s material claims are supported by evidence (factual correctness). | LLM judges |
| Policy compliance | Whether the agent followed the explicit rules in the written policy. | LLM judges |
| Relevance | Whether the agent stayed on the customer’s real need. | LLM judges |
| Completeness | The share of expected end-state assertions that pass in the environment. | Measured assertions, without judgment |
In G2's September 2026 evaluation of CX agents on identical tasks, accuracy ranged from 64% to 88% and policy compliance from 56% to 92%. Results are published under the G2 AI CX agent leaderboard.
What evaluating a simple task looks like
Let’s take a real scenario from the CX set: recovering Google sign-in access. The simulated customer cannot get into their G2 account through Google sign-in and asks the agent to reset their Google SSO password. The correct answer is that the agent cannot, because authentication is delegated to Google. A good agent says so and points the customer to Google's account recovery first. Every agent in the category faces this exact conversation, and you can read each one's answer side by side.

Screenshot: Tidio’s response to the simulated Google sign-in recovery task.
And here is what grading looks like on a single passed task, a customer checking the status of a category change request: relevance 5/5, completeness 5/5, accuracy 4/5, policy compliance 3/5.
The task passed because the required outcome happened. The dimension scores capture when a specific policy has not been followed.

Screenshot: A passed task with separate scores for relevance, completeness, accuracy, and policy compliance.
What you see on a vendor page
We publish the work itself done by the agent. Every evaluated agent's profile shows sample tasks: the scenario, the full conversation with the simulated customer, the date it was evaluated, and tabs to read how other agents handled the identical scenario. Each task card includes the pass/fail result, the dimension scores, and the grading evidence behind them.
The agents participating include Tidio, Zendesk, Retell, Sierra, LiveAgent, Chipp, BoldDesk, Zoho, Fin (now on Salesforce), Jotform, Freshworks, Yuma, Hubspot, Front and Gorgias.
A specific dataset for e-commerce agents
Agents like Gorgias and Yuma are built for e-commerce support: orders, returns, exchanges inside commerce stacks (Shopify app for instance). Scoring them on a generic B2B support desk would measure them outside their domain. So they run a separate task set built on the public τ²-bench retail environment and rank on a dedicated e-commerce leaderboard.
Gorgias also publishes their own benchmark for E-commerce covering automation rate and latency, metrics we have not captured in our first run.
What’s next for G2 CX AI agent evals
We refresh CX scores quarterly, with the next run in Nov-Dec 2026. We’re expanding our task set to increase complexity, and getting buyer feedback to improve our methodology to get closer to match their day to day environment.
The full methodology is public here. If you think we should test something differently, we welcome feedback at agent-evals@g2.com
Comments
Questions about how this works? Ideas for what comes next? Join the discussion.
Loading comments…