G2 is Building the New Trust Layer for AI Agents
Alexis Zheng

For over a decade, G2 has been the trust layer for software markets: independent, verified feedback from real users, so buyers can choose with confidence.
But the buying journey has changed. We used to sell to people who spent months researching software. Today, we increasingly sell to their agents, which do the research, build the shortlist, and shape the decision. Agents, like people, have a discovery problem. They need trust signals too.
Meanwhile, every buyer hears the same pitch on AI agents, and every demo looks like a Claude demo. Buyers respond with bake-offs on their own data, and brands publish evaluations against their peers. The results rarely agree, because the industry has no standard for comparing agents beyond model benchmarks.
A new project: agents evaluating agents
For the past few months, we have worked with enterprise agent builders and buyers to test whether agentic evaluation can give buyers a grounded signal.
We built a simulated company with customer support, sales, legal, and finance departments — each with its own data, policies, artifacts, and tools — and drew tasks and ground truth from real enterprise operational data. We then hand an agent a set of tasks, capture every interaction as a trace, and grade it against a rubric. The trace matters because an agent can hit its goal and still land on the wrong end state; for example, a request to pause a subscription should not be resolved by canceling it.
Grading separates verifiable facts from judgment and allows any path that stays within the task's factual and policy constraints. Agents perform the grading because a single run produces more steps than humans can consistently audit. A panel of buyer and vendor design partners helped us iterate on the methodology.

Simulated vendor agent tasks
Our early findings
Our first leaderboard for Customer Experience AI Agents is live in beta. Across agents running identical tasks, accuracy ranged from 64% to 88%, and policy compliance from 56% to 92%. This is what buyers are seeing in their own bake-offs as well and why hiring an agent is difficult.
We expected the results to sort agents by performance. They did. They also sorted them into two camps we had not set out to find, and the line ran along product design rather than model choice. On one side is the last decade of enterprise automation, built on custom workflows. On the other is a new class of agents, built as LLMs got powerful enough to trust with judgment, that work through delegation interfaces: Cowork-style, chat, or headless.
Workflow builders suit tasks that require guaranteed decisions. Every decision point is encoded and enforced in logic. That matters to enterprises that have spent years refining their workflows, and to buyers trained around them. But it is a deterministic layer over a probabilistic model, and that tension never fully resolves.
Headless interfaces relax those guarantees. Users work with the agent like a colleague, setting policies and goals rather than steps. Onboarding is faster and out-of-the-box performance is strong, especially on greenfield projects. Natural language is becoming the only interface flexible enough to match what the latest models can do, and headless modes are now on the roadmaps of some of the strongest incumbents.
A third class, vertical-specific agents such as CX agents built directly into e-commerce interfaces, meets different requirements and performs best inside its domain. We are adding domain-specific features so the leaderboard can show that.

Score range across evaluated agents
The tradeoff is control versus iteration speed. Our beta leaderboard favors headless interfaces because their workflows come directly from SOPs and policy documents. For agents built on carefully designed workflows, the same setup measures a new customer navigating an unconfigured system, not a team upgrading with workflows in place. Configuring those agents well also take more rounds with each vendor.
That makes the current leaderboard a launchpad for headless challengers and an early read on where new entrants are pushing the category. It is not yet a complete picture.
What comes next: evaluating configured agents
The shift to headless makes sense once you see where the capability moved. With today's reasoning models, you build the product from the model rather than the workflow. This is because the model now defines the workflow better than a human expert can. You do not micromanage an excellent team.
That raises a question we cannot answer yet. If the model defines the workflow, how much of the product design we inherited from the last decade still earns its place? Every builder in this category, ourselves included, may be due for a fresh look. We will let the evals speak to it.
Frontier models will keep improving, handling more uncertainty with less onboarding across a wider range of work. But steering and alignment remain hard. Enterprise-critical work, legal and financial compliance especially, still demands guaranteed rule adherence, and configured workflows remain the answer there.
There is a trust problem underneath. As the dispute over OpenAI's claimed proofs of 10 open math problems shows, models can now produce answers that humans had not reached, but they arrive as a black box. It is hard to trust an answer when you cannot see how it was reached. That evolution, from hand-built workflows to headless configuration, is why we are building these evals: to work with buyers and vendors on a methodology that lets enterprises adopt agents at scale.
So the next phase is the incumbents, and it requires a different method. Fairly evaluating a configured agent means working with each vendor's forward-deployed engineers to stand up a standard deployment before scoring. That takes several rounds, and it measures the product at its best under a configuration both sides recognize as standard, not how well it survives being dropped cold into someone else's harness. The open question is how to ensure fair cross-vendor comparison in that setup.
Build with us: request an evaluation, and tell us what you would do differently
We are launching G2 Agent Evaluations as a public beta because there is no playbook for evaluating enterprise agents yet. It means making calls on what counts as success, which edge cases matter, and how to keep the comparison useful as products change. We have a point of view, but the first version should not be the final one.
In the spirit of building publicly and transparently, I have a few resources to share, and we would love your input:
- G2 resources: Check out the full methodology, the early insights from evaluating 10+ CX agents, and a simulated evaluation you can walk through. You can also follow Cyril Delattre on X for regular updates as we continue to build.
Requests for input:
- If you work on agent evaluation, enterprise AI, safety, or procurement, tell us where the methodology is strong, where it is incomplete, and what you would test differently at agent-evals@g2.com.
- If you want your agent evaluated, request an evaluation.
My team and I will be listening closely. Thank you in advance to everyone contributing to the future of agent evaluations on G2.