Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Companies should evaluate their products against named competitors on live products, publish the scoring method, and show where they lose. Gorgias’s ecommerce AI-agent benchmark is a useful case study—not independent certification—because Gorgias publishes the comparison while also competing in it. The strongest lesson is not that one vendor is definitively best; it is that a comparison becomes more useful when its scope, evidence, trade-offs, and commercial interests are visible.

Why publish direct competitive evaluations?

Product comparisons are often built from demos, feature lists, or claims chosen by the vendor. A live-product evaluation can answer a harder question: what happens when competing products face comparable tasks under stated conditions? That matters for AI agents, whose behavior in a configured storefront may differ from a polished demo.

Jason Lemkin, SaaStr founder, argues that companies should “Run them on live products, against named competitors, and include the categories where you lose.” The last part is essential. A ranking without disclosed weaknesses can read like advertising; a comparison that identifies a real shortcoming gives readers a way to judge the result and decide whether that trade-off matters to them.

What the Gorgias benchmark measures

Gorgias’s benchmark, marked refreshed October 2026, evaluates ecommerce AI agents in two jobs: helping a shopper find and buy a product, and answering support questions about shipping, returns, and store policies without a human. Gorgias reports 9,226 conversations captured, 9,220 blind-judged, 18 vendors, and 224 live stores. These are changing benchmark counts, not market-wide statistics. The Gorgias benchmark says every vendor receives the same questions, adapted to the store catalog.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the comparison is run

  • Each run uses a cold browser session. The auditor does not ask for a human, so any handoff must be initiated by the agent.
  • Vendor names are hidden for blind scoring, and claims count only when they are quoted from the conversation transcript.
  • Quality is derived from binary, evidence-forced checks rather than a judge simply selecting an overall impression.
  • A vendor needs at least 15 judged conversations in a job to qualify for a head-to-head rank.
  • Gorgias says the benchmark is rerun weekly. Results therefore need to be read with their reporting date and sample counts, rather than as permanent standings.

These choices make the method easier to inspect, but they do not make the publisher independent. A blind rubric can reduce some forms of evaluator bias; it cannot by itself eliminate choices about task design, sample selection, or which outcomes matter.

Three separate measures, not one universal definition of “best”

The benchmark scores automation, answer quality, and speed. Automation is the share of conversations resolved without human involvement; answer quality is a blind score out of 100; speed is time to a complete answer. Gorgias combines those measures with different weights for each job:

Job Automation Answer quality Speed
Shopping assistant 40% 35% 25%
Support agent 50% 40% 10%

The composite ranking reflects those publisher-stated weights. A reader who values speed more heavily, or prioritizes resolution without human intervention, could reasonably reach a different conclusion from the same underlying measures. This is why a published score should come with its rubric and weights rather than standing in for them.

Rank #2
Adams Service Call Book, 5.25 x 11 Inch, Spiral Binding, 2-Part, Carbonless, 4 Messages per Page, 200 Sets, White and Canary (SC1155), White/Canary
  • Record all incoming calls needing service
  • 2-part carbonless
  • Spiral bound on left
  • Part one is perforated to give to service person, part two remains in book for records
  • White, canary paper sequence

What the October 2026 results say—and where Gorgias loses

In the current display, Gorgias ranks #1 overall, #2 in support, and #3 in shopping. Gorgias reports support answer quality of 74/100 and says its shopping answer quality is the highest in the field. It also identifies speed as its weakness: shopping answers take about 18 seconds, compared with about 8 seconds for Envive and about 10 seconds for Sierra; support answers take about 14 seconds. These are claims from the vendor-published benchmark, not independently certified findings.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is the useful shape of an honest comparison: a leading overall position alongside a clearly named gap. A buyer evaluating shopping use cases may care more about that speed difference than an overall composite rank. Another may accept slower responses if the answers are better. The benchmark makes that trade-off legible instead of hiding it behind a single headline score.

Gorgias also cautions that no vendor leads automation, quality, and speed at once, and that the same vendor can perform differently across stores. Configuration matters. It further reports that almost a third of detected “AI chat” widgets did not produce a real conversation. Treat these as findings reported by Gorgias, not independently audited market statistics.

Keep the older SaaStr snapshot separate

Jason Lemkin’s September 26, 2026 SaaStr article describes an earlier benchmark snapshot: 8,356 live conversations, 18 vendors, and more than 212 storefronts. It should not be blended with the Gorgias page’s October 2026 counts as though both describe the same reporting period.

The SaaStr article reports an Envive pre-sale composite score of 72 against Gorgias at 65, Gorgias answer quality of 76, and average shopping response times of 18.4 seconds for Gorgias versus 7.9 seconds for Envive. It also says 28% of Gorgias shopping answers took longer than 20 seconds, and that Gorgias scored 74.3 when the support weights were applied. Those figures belong to Lemkin’s earlier account; the current Gorgias page presents a newer snapshot and different score descriptions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Lemkin’s article calls Gorgias a roughly $100 million ARR company and says about 80% of its revenue came from AI support for ecommerce brands. Those are company figures reported by the article, not independently verified here. The relationship matters too: SaaStrFund led Gorgias’s seed round, and Lemkin says SaaStr encouraged Gorgias to publish the evaluation. Readers should consider that context when weighing the case study.

How to judge a vendor-run benchmark

Check what was tested

Look for the task type, store configuration, sample size, reporting window, and any threshold required to appear in the rankings. Shopping assistance and customer support are not interchangeable workloads, and performance in one does not establish performance in the other.

Separate the measurements

Read automation, answer quality, and time to a complete answer individually before using a composite rank. Check what “resolved” means, what evidence supports a quality score, and whether speed means first response or a complete answer. A fast but incomplete reply is not equivalent to a resolved conversation.

Inspect the publisher’s stake and method

Ask who runs the benchmark, who competes in it, and who chose the rubric. Gorgias says it applies the same blind rubric to itself and competitors, and says a former Gorgias-only exclusion rule was removed in July 2026. That disclosure helps readers understand the process, but the company remains both benchmark publisher and participant.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Lemkin’s article says the harness is open-sourced; the public GitHub repository lets readers inspect project files. Open code improves inspectability, but its existence does not show that every reader has independently rerun or validated the benchmark.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical standard for publishing your own evaluation

  1. Name the products and define the jobs. State which live products were tested and which real tasks they were asked to complete.
  2. Make conditions comparable. Use the same task set and explain any adaptation, such as tailoring questions to each store’s catalog. Record configuration differences that could affect outcomes.
  3. Publish the evidence and scoring rules. Explain how automation, answer quality, and completion time are measured; show the weighting used for any composite score.
  4. Report sample and timing. Give conversation counts, qualification thresholds, reporting dates, and how often the test is rerun. Do not present a changing snapshot as a permanent market fact.
  5. Show weaknesses as well as strengths. Identify where the publisher trails competitors and let readers decide how important that trade-off is for their use case.
  6. Disclose interests and make the method inspectable. Explain the publisher’s commercial relationship to the products. If code is public, link it, while avoiding any implication that publication alone proves neutrality.

For buyers, the same framework is a useful checklist for deciding whether a vendor comparison answers their question. The best benchmark is not necessarily the one with the most elaborate score; it is the one whose conditions and trade-offs are clear enough to relate to the buyer’s own deployment.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.