Freshworks
Freshchat (a product of Freshworks Inc.) is a conversational engagement solution that helps businesses engage with customers across multiple channels such as web chat, email, phone, and social channels like WhatsApp, Instagram, and iMessage. Freshchat also offers native AI-powered chatbots that can be easily built with a no-code builder, which helps businesses provide personalized self-service at scale. With Freshchat's journey builder and ready-made playbooks, businesses can stay ahead of their customer's needs and support them proactively whenever necessary. Freshchat also provides a comprehensive view of the customer's conversation history and other relevant information, making agents more productive and enabling them to provide personalized experiences.
Overall scoreFive points for every passed standard task. Missing or failed tasks earn zero. This result was evaluated on Aug 5, 2026.
21/46 tasks passed
Explore task examplesHow Freshworks performed
(Showing 3 of 46 tasks)Explore the conversations and grading behind Freshworks’s evaluation scores.
Recovering Google sign-in access
Account accessPassed
Recovering Google sign-in access
Account accessThe Scenario
You sign in with Google but cannot get into your G2 account. You want the problem fixed, but you also want to know whether the path you are taking is the right one.
The Conversation
Simulated customer and Freshworks. Tool calls and results appear in sequence when available. Some details are omitted.
Freshworks
Hi! How can I help you today?
Simulated customer
I can’t get into my G2 account using Google sign-in. Can G2 reset my Google SSO password, or is there something else I need to do?
Freshworks
Simulated customer
Thanks, that makes sense. I’ll recover or fix the Google sign-in directly with Google first, then try “Continue with Google” on G2 again. Once I can sign in to Google, should G2 access work immediately, or is there any delay I should expect?
Freshworks
How It Was Graded
Task outcome and response quality are scored separately. A passed task can still contain mistakes.
- Relevance
- 5/5
- Completeness
- 5/5
- Accuracy
- 4/5
- Policy Compliance
- 5/5
View Grading Evidence
The agent explained that G2 cannot reset a Google SSO password and that the buyer must fix their sign-in with Google, then retry.
Grader’s Explanation
The agent clearly stated that G2 cannot reset the Google password, explained that Google is the identity provider, and provided steps for the user to recover their Google account and then retry logging into G2.
The agent did NOT attempt a G2 password reset for this SSO account.
Grader’s Explanation
At no point did the agent attempt or suggest a G2 password reset for the SSO account. The agent only directed the user to recover their Google account.
The agent explained the next step, expected timeline, or what the user should monitor before ending the conversation.
Grader’s Explanation
The agent explained that after fixing Google sign-in, G2 access should work immediately, and advised the user to try again and report any further issues, specifying what information to provide if problems persist.
Accuracy
factual_correctness: 4 — Agent's key factual claims: (1) G2 cannot reset a Google SSO password—only Google can. (2) Fix Google login first, then use 'Continue with Google' on G2. (3) G2 access should work immediately once Google auth is fixed—no extra delay on G2's side. All three are reasonable, well-grounded SSO process claims consistent with standard OAuth behavior. The snapshot confirms sso_working=false and sso_provider=google for this user, which aligns with the agent's framing. No tool results contradict anything. Minor: the claim of 'immediate' access is a process generalization not directly verifiable from evidence, but it's not contradicted and is a standard accurate description of OAuth flows. No task-critical errors found.
Policy Compliance
policy_compliance: 5 — Agent correctly identified this as a Google SSO issue and explicitly stated G2 cannot reset the identity-provider password, directing the buyer to fix sign-in with Google first then retry G2 — exactly as required by policy 4.3. Agent gathered context, explained the recovery path, and closed with a concrete next step. No policy rules were violated.
Checking a category change request
Category changesPassed
Checking a category change request
Category changesThe Scenario
You need an update on a category request already in flight. You want to know whether the request is blocked, self-serve, pending, or requires evidence-backed review.
The Conversation
Simulated customer and Freshworks. Tool calls and results appear in sequence when available. Some details are omitted.
Freshworks
Hi! How can I help you today?
Simulated customer
I’m looking for an update on a category request we already have in flight. Is it currently blocked, pending review, self-serve, or awaiting evidence—and what timing should I plan around?
Freshworks
Simulated customer
The request ID is cat_001.
Freshworks · Tool callget_category_request_status()
Arguments
{
"request_id": "cat_001"
}Freshworks · Tool resultget_category_request_status
Returned data
{
"current_category": "Customer Support",
"decision_date": null,
"product_id": "prod_002",
"request_id": "cat_001",
"requested_category": "Help Desk",
"status": "pending",
"submitted_date": "2026-06-29",
"vendor_user_id": "vendor_002"
}Some fields are omitted from this public example.
Freshworks
Simulated customer
Who owns the next step on your side, and what review timeline or SLA can I communicate internally? Also, what qualifies a listing for the Help Desk category, what evidence would matter if the reviewer asks for it, and what’s the process if this is rejected?
Freshworks
Simulated customer · Tool callopen_vendor_portal()
Arguments
{}Some fields are omitted from this public example.
Simulated customer · Tool resultopen_vendor_portal
Returned data
{
"vendor_user_id": "vendor_002",
"logged_in": true
}Some fields are omitted from this public example.
Simulated customer · Tool callrefresh_vendor_request_status_page()
Arguments
{
"request_id": "cat_001"
}Simulated customer · Tool resultrefresh_vendor_request_status_page
Returned data
{
"request_id": "cat_001",
"vendor_user_id": "vendor_002",
"request_type": "category_change",
"visible_status": "pending",
"last_updated": "2026-06-29T12:00:00",
"refresh_count": 1
}Some fields are omitted from this public example.
Simulated customer
I refreshed the request status page for cat_001. It still shows Pending, last updated 2026-06-29, with no additional message.
Freshworks
How It Was Graded
Task outcome and response quality are scored separately. A passed task can still contain mistakes.
- Relevance
- 5/5
- Completeness
- 5/5
- Accuracy
- 5/5
- Policy Compliance
- 4/5
View Grading Evidence
The agent explained that the category request is already pending review.
Grader’s Explanation
The agent explicitly stated multiple times that the request (cat_001) is 'pending review', is in the 'internal category review queue', and is not blocked, self-serve, or awaiting evidence. This was clearly communicated in the agent's responses.
The agent explained the next step, expected timeline, or what the user should monitor before ending the conversation.
Grader’s Explanation
The agent explained that the next step is an internal review by G2's category/operations team, that there is no published SLA, and advised the user to monitor the vendor portal and email for updates. The agent also clarified what to do if a rejection or evidence request is received.
When the agent asked the user to perform a tool-side action, it explained what that step would verify or accomplish. (If the user acted on their own initiative, this expectation still counts as met.)
Grader’s Explanation
The user refreshed the request status page on their own initiative. The agent acknowledged this and explained what the current status meant, confirming that the request was still pending and what the user should expect next. This meets the expectation.
The agent kept the conversation in vendor scope and did not confuse the issue with buyer/reviewer support.
Grader’s Explanation
Throughout the conversation, the agent focused on the vendor's category change request, the internal review process, and did not mention or confuse the issue with buyer or reviewer support. The scope was kept strictly to vendor-side category management.
Accuracy
factual_correctness: 5 — All explicit task-relevant factual claims are accurate and supported by tool evidence. Agent correctly reported cat_001 status as pending, current category as Customer Support, requested category as Help Desk, submitted date 2026-06-29, no decision date, no rejection reason. The refreshed status (turn 9) accurately reflects the refresh_vendor_request_status_page result: pending, last_updated 2026-06-29, no latest message. The agent appropriately declined to quote a specific SLA (none available in evidence) and attributed next-step ownership to internal review team, which aligns with domain facts. No material factual errors found.
Policy Compliance
policy_compliance: 4 — The agent did not verify vendor identity before providing account-specific request details in Turn 5. Policy section 1 requires identifying the caller before giving account-specific guidance, using vendor-facing identifiers such as work email, company name, profile id, request id, or product name. The agent asked only for the request ID and then immediately pulled and disclosed the full request state without confirming company name, work email, or any other corroborating identifier. However, the request ID itself is a vendor-facing identifier listed in the policy, the information disclosed (category change status) is low-sensitivity, no mutating action was taken, and the impact is limited and recoverable. This is a minor violation of the identity-verification rule with low practical harm.
Disputing a rejected review
Review moderationExecution Error
Disputing a rejected review
Review moderationThis task had an execution error. No graded conversation is available.
Estimated autonomyA sourced estimate of how independently the agent can plan and complete work on the L1–L5 scale.
Where this agent sits on the L1–L5 autonomy scale.
Skills
The five shared skills reviewed for products in this category.
Case Routing & Escalation
Grounded Answers
Policy Compliance
Refunds & Billing
Performance claims
Vendor-reportedResolve up to 80% of queries on chat, messaging apps, and email
Frontier skills
Emerging capabilities selected for this evaluation category.
Predictive Issue Resolution
Detects emerging problems from data patterns and resolves them before the customer notices or reaches out.
View sourceMultimodal Claim Analysis
Analyzes photos, videos, and documents submitted with a claim to assess damage, extract data like receipts, and match items automatically.
View sourceWhat G2 reviewers say
AI-generated Pros and Cons based on aggregate themes from verified G2 buyer reviews, not individual review quotes.AI-generated from verified G2 buyer reviews.
Pros
Reviewers frequently mention the clean, intuitive interface, the ability to manage conversations across multiple channels from a single dashboard, and the AI-powered chatbots that handle routine questions, reducing the workload on support agents.
Cons
Reviewers experienced limitations with advanced customization, basic reporting and analytics compared to competitors, inconsistent mobile app experience, occasional notification delays leading to missed messages, and performance issues during peak volumes.