Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To measure AI share of voice, fix a prompt library and a competitor set, run the same prompts on ChatGPT, Gemini and Perplexity several times, keep every raw answer and citation, and report results separately for each engine with the denominator stated. Python handles collection, classification and arithmetic well. It cannot give you the engines’ internal ranking data, so every figure you produce describes sampled answers you collected on the dates you collected them, not a census of everything the engines might say.

Define what “share of voice” means before you calculate it

The phrase is used loosely. It can mean how often your brand appears in AI answers, your portion of all brand mentions across a competitor set, or another ratio a vendor or analyst has defined. Each answers a different question, so the formula has to be fixed before you collect any data. There is no universal industry standard for this metric. Vendor tools define their own, and your report should label the one you use.

Metric Numerator Denominator Question it answers Watch for
Mention rate Measured answers that name the target brand All successfully measured answers for that engine and prompt set How often the brand appears at all Alias errors and ordinary words that double as brand names
Own-domain citation rate Answers with at least one cited URL on the brand’s own domain or a subdomain All successfully measured answers How often the answer links to the brand’s own pages A third-party article that mentions the brand is a mention source, not an own-domain citation
Recommendation rate Answers that explicitly advocate the brand All successfully measured answers How often the answer picks or endorses the brand Needs written annotation rules; automated labels are not ground truth
Share of all brand mentions The target brand’s mention count across answers Total mentions of every tracked brand across the same answers The brand’s relative prominence inside the competitor set Changes when competitors are added or removed, so fix the set before collection

Keep these outcomes in separate fields. A mention says the brand was named. A citation says a source was linked. A recommendation says the answer advocated a choice. An answer can name a brand without citing it, and can cite a page without naming the brand in the answer text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When several brands can appear in one answer, a per-brand answer share is counted independently for each brand. Those shares can add up to more than 100%. The SourceWatch API documentation uses this per-brand definition: each brand’s share is its mentions divided by the number of answers measured. That is one defensible definition, not the only one.

Set up the inputs

Category, brands and aliases

Write down the business category in one sentence, the target brand, and a fixed list of competitors. Then map every spelling variant to the brand so that a missed alias does not look like a missed mention. Include product names, abbreviations and common misspellings, and record the mapping in version control next to the prompt set.

Prompt library

Build prompts from questions real customers ask, such as “What is the best tool for…” or “Which option should a small team choose for…”. For each prompt, record its wording, category, intended audience, locale and collection date. Yext describes a prompt-library approach with competitor comparisons, which is a useful model for this stage. Do not change wording or add prompts in the middle of a comparison. If the prompt set has to change, start a new measurement period and do not splice the two together.

A practical configuration file keeps all of this in one place. The example below uses placeholder brand names and a sample category:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
{
  "category": "project management software for small agencies",
  "target_brand": "ExampleCRM",
  "aliases": {
    "ExampleCRM": ["ExampleCRM", "Example CRM", "ExampleCRM Inc"],
    "RivalSuite": ["RivalSuite", "Rival Suite"]
  },
  "engines": ["ChatGPT", "Gemini", "Perplexity"],
  "runs_per_prompt": 5,
  "prompt_file": "prompts.csv"
}

Collect answers from each engine

Run the same prompt set on every engine you compare. Store each answer as its own record, including failed attempts. The fields below are a practical minimum:

  • engine and model_or_interface, the model or interface name as you can observe it
  • prompt_id, prompt_text and run_id
  • collected_at_utc, an ISO 8601 timestamp
  • answer_text, the raw answer, unedited
  • citation_urls, every URL the interface shows
  • retrieval_used, true or false when the interface indicates whether it searched the web, and null when it does not indicate this
  • recommendations and brand_mentions, filled in during classification
  • collection_status, either “ok” or “failed”, with an error message for failures
from dataclasses import dataclass, field, asdict
from typing import Optional
import json

@dataclass
class AnswerRecord:
    engine: str
    model_or_interface: str
    prompt_id: str
    prompt_text: str
    run_id: int
    collected_at_utc: str
    answer_text: str
    citation_urls: list = field(default_factory=list)
    retrieval_used: Optional[bool] = None
    collection_status: str = "ok"
    error: Optional[str] = None
    brand_mentions: list = field(default_factory=list)
    recommendations: list = field(default_factory=list)

def save_records(records, path):
    with open(path, "w", encoding="utf-8") as f:
        for r in records:
            f.write(json.dumps(asdict(r), ensure_ascii=False) + "n")

The code requires Python 3.10 or later for the modern type syntax used above. Store the raw corpus as JSON Lines so that a person can open any answer and check how it was classified.

Know what each engine interface is

Record which surface you queried. A consumer app, a web interface and an API can differ in model, search behavior and citation display. The sources reviewed for this guide do not establish how closely API output matches what a consumer app shows, so keep the surfaces separate and do not generalize between them.

Engine What to record What the sources establish
ChatGPT Model or interface name, and any web-search indicator the interface shows A 2026 paper by Ronald Sielinski studied OpenAI’s SearchGPT, not the ChatGPT app; results do not transfer automatically
Gemini Model or interface name, and any search or grounding indicator shown Google’s Search Console generative AI report covers Google Search and Discover surfaces; it is not a measure of the Gemini app
Perplexity Interface name, and whether the answer displays sources The 2026 paper by Ronald Sielinski studied Perplexity Search, a search-based product

Record failures as failures

A failed call is not a zero. If a run times out or returns an error, store it with collection_status set to “failed”. Decide and document the policy for failures before you calculate anything. The usual choice is to exclude failed runs from the denominator and report the failure count next to every rate. A brand that disappears because a provider returned errors looks very different from a brand that was absent from successful answers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plan the run count

Planned collections equal engines × prompts × runs per prompt. Three engines, 40 prompts and 5 runs per prompt give 600 collections. The documentation of one open-source Python collector says cost scales with that product and recommends a small run count first to validate the configuration before increasing sampling. Start small, check that the answers are being captured and parsed correctly, then scale.

Classify each answer

Brand detection with word boundaries avoids the most common false positives, such as matching a brand inside a longer word. The function below uses the alias mapping from your configuration:

import re

def brands_in_text(text, aliases):
    found = set()
    for brand, names in aliases.items():
        for name in names:
            pattern = r"(?<!w)" + re.escape(name) + r"(?!w)"
            if re.search(pattern, text, flags=re.IGNORECASE):
                found.add(brand)
                break
    return found

Automated matching will still make mistakes. A brand name that is also a common word, or a product name that overlaps a competitor’s product line, needs manual review. Draw a random sample of answers, label mentions, recommendations and citations by hand, and compare the labels with the script’s output. Report the agreement rate you observed. Treat automated sentiment labels, if you use them, as a rough signal rather than ground truth.

Recommendations and citation ownership need explicit rules. Write them down. For example: an answer counts as a recommendation only when it tells the reader to choose or use the brand, not when it lists it among options. An answer counts as an own-domain citation only when a cited URL’s host is the brand’s domain or a subdomain of it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Calculate the metrics

For a single engine, the mention rate is the number of successful answers that name the brand divided by the number of successful answers. The function below returns both the rate and the number of answers it is based on, so the sample size is never separated from the result:

def mention_rate(records, brand, aliases, engine=None):
    ok = [r for r in records
          if r.collection_status == "ok" and (engine is None or r.engine == engine)]
    if not ok:
        return None, 0
    hits = sum(1 for r in ok if brand in brands_in_text(r.answer_text, aliases))
    return hits / len(ok), len(ok)

Apply the same logic to recommendation and own-domain citation rates, with their own numerators and the same measured-answer denominator. Then compute the share of all brand mentions separately if you need a relative prominence figure.

A worked example shows why the denominator matters. Suppose 10 successful Perplexity answers name ExampleCRM in 6 answers and RivalSuite in 5. ExampleCRM’s answer share is 60% and RivalSuite’s is 50%, a combined 110%, because one answer can name both. The share of all brand mentions for ExampleCRM is 6 ÷ 11, or about 55%. Both figures are correct, but they answer different questions, so label them differently.

Read the numbers with sampling error in mind

Repeated runs are samples, not a complete picture of what an engine can say. A 2026 paper by Ronald Sielinski, which studied Perplexity Search, OpenAI’s SearchGPT and Google Gemini, found substantial variability across repeated submissions and argued that single-run visibility figures can look more precise than they are. The paper’s finding is methodological and tied to the platforms, topics and sampling it describes. It does not supply a benchmark you can apply to your category.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A simple way to show uncertainty is a Wilson 95% interval for each rate. Take 10 measured answers in which the brand appears in 5. The observed rate is 50%, but the Wilson interval runs from about 24% to about 76%. With only 10 answers, a change from 50% to 60% is well inside the noise. Report the number of answers next to each rate, and avoid claiming a meaningful change unless the intervals clearly separate across measurement periods.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Report results by engine first

Keep engine-level figures as the primary output. A single blended score can hide the fact that one engine mentions a brand often while another rarely does. If you publish a combined figure, state how engines and prompts are weighted. A simple unweighted average is misleading when engines have different numbers of successful runs or when prompt counts differ.

A report that others can audit should include:

  • The category, target brand, competitor set and alias mapping, with the date they were fixed
  • The full prompt set, with IDs, categories, audiences and locales
  • For each engine, the planned runs, successful answers, failed runs and the failure policy
  • Mention, own-domain citation and recommendation rates, each with its numerator, denominator and interval
  • The share-of-voice formula used, stated in words and in the report header
  • The collection dates and the interface or model identifiers observed
  • The results of the manual validation sample, including agreement and known error types

What Google’s own tools cover

Google Search Central’s guide “Google’s Guide to Optimizing for Generative AI Features on Google Search” states: “You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities).” The same guidance says site owners should keep following foundational SEO practices. Search Console offers a generative AI performance report covering visibility in Google Search and Discover generative AI features.

That report is first-party measurement for Google surfaces. It does not measure ChatGPT or Perplexity, and it is not a measure of the Gemini app. Google also says third-party tools do not have access to its internal ranking or AI systems, so any commercial monitoring product is collecting answers the same way you can, not reading hidden Google metrics.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build or buy

You can build a Python collector with the structure above, or use a managed monitoring product. Whichever you choose, compare options on the same criteria:

  • Coverage of all three engines and the interfaces you need
  • Control over prompts, competitor sets and the number of repeat runs
  • Export of raw answers and citation URLs, so classifications can be audited
  • Separate reporting of mentions, citations and recommendations
  • Per-engine reporting rather than only blended scores
  • Transparent handling of failed runs and documented denominators
  • Reporting and collaboration features your team needs

Category examples include Yext, which describes a prompt-library approach with competitor comparisons, and SourceWatch, whose API documentation describes visibility and share-of-voice outputs. These show that managed tools exist for this workflow. They do not establish a vendor ranking. Confirm current engine coverage, pricing and terms directly with any vendor before you rely on it, because these details change.

Sources and limits of this guide

This guide describes a method. It does not report a market benchmark for how often brands appear in ChatGPT, Gemini or Perplexity, and it does not report results from a hands-on comparison of the engines. The sources used were Google Search Central’s generative AI guidance, the Yext and SourceWatch product documentation, the documentation of an open-source Python collector, and the 2026 Sielinski paper. Engine interfaces, models, API availability and vendor features change, so check each against the current product documentation before you run a measurement.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.