Public methodology
How G2 evaluates AI agents
G2 Agent Evaluations is a hands-on evaluation program for AI agents, beginning with customer support agents. The methodology combines category-specific rubrics, consistent real and synthetic scenarios, model judging, and human review.
Overview
Measured performance instead of marketing claims
G2 evaluates customer support AI agents hands-on and publishes the results so buyers can compare measured performance. The first category is customer support, with more categories expected to follow.
Within a category, G2 intends to apply the same methodology consistently to every vendor.
Evaluation process
Consistent scenarios and category rubrics
G2's market research team works with buyers and design partners to define what good performance looks like. Those expectations become category rubrics broken into skills.
Each agent in a category completes the same real and synthetic scenarios through direct product access, which may be provided through an API, sandbox, or white-glove setup.
Score interpretation
One score cannot describe every capability
The Overall Score is the percentage of the 46 assigned tasks that passed: tasks passed divided by 46, rounded to the nearest whole number. Missing or failed executions remain in that fixed denominator.
Accuracy, Policy Compliance, Relevance, and Completeness are separate rubric-based dimensions. They explain response quality but are not blended into the Overall Score.
Only verified G2 evaluation results are included in the current comparable leaderboard.
Category design
Skills reflect buyer needs
Category rubrics are organized into buyer-relevant skills. The methodology source names refund handling, deflection, and escalation as examples for customer support agents.
Product profile skills describe documented product-specific work. Evaluation dimensions describe performance in the verified task run; the two should not be interpreted as the same signal.
Data provenance
Each signal keeps its own source
Evaluation scores come from G2's agent evaluation system. Products appear together on a leaderboard only when their task set, rubric, judge, dataset, and version fingerprint match. Ties in Overall Score are ordered by Accuracy, then Policy Compliance.
Product skills, integrations, operating choices, and vendor-reported claims come from reviewed product sources. G2 ratings and review metrics are separate buyer-feedback signals and are never used to calculate the task-based Evaluation score.
Quality control
Model judging anchored by people
An LLM judge scores responses against the category rubric. Human-in-the-loop review from buyers and practitioners anchors that judging process.
Before publication, G2 walks vendors through the initial run to capture operational context that raw scores may miss.
Point-in-time results
Evaluations evolve with the products
An evaluation is a snapshot, not a permanent label. Results may change as an agent improves and as G2's category standards evolve.
Counts, product versions, exact dates, score formulas, and scenario evidence appear only when verified evaluation records supply them.