Measure LLM brand visibility by running a fixed library of real buyer prompts across every AI surface your audience uses, repeating those runs over a defined time window, and coding mentions, recommendations, citations, position, competitors, accuracy and tone separately. Publish the prompts, competitors, platforms, sample size, dates and formulas with every score. Without that context, a visibility number cannot be interpreted or reproduced.
What LLM brand visibility actually measures
LLM visibility is not one outcome. An answer may mention your company without linking to it, cite your page without recommending you, or recommend you while relying entirely on third-party sources. Track those events as separate variables.
| Signal | Operational definition | Important qualification |
|---|---|---|
| Mention rate | Percentage of prompt runs in which the brand appears at least once. | State whether repeated mentions in one answer count once or multiple times. Ahrefs Brand Radar counts one mention per response. |
| Citation rate | Percentage of answers that cite a page on your domain, or your identified source domain. | Only calculate where the engine exposes sources; otherwise report the metric as unavailable. |
| Recommendation rate | Percentage of answers that actively suggest the brand as a solution. | A name in a background list is not a recommendation. |
| Position | Order in a recommendation list or answer. | Directional: ordering can change between runs. |
| Share of voice | Your brand mentions divided by the denominator you define for the frozen competitor set. | Inner Labs uses the share of all brand mentions; Ahrefs defines AI share of voice as an impressions-weighted percentage. These figures are not interchangeable. |
| Impressions or demand | A modeled estimate of demand associated with prompts where your brand appears. | AI platforms do not publish prompt-level volumes. Ahrefs sums Google search volumes, so label this as an estimate rather than actual AI query count. |
| Accuracy and tone | Human-reviewed labels for factual errors, outdated claims, sentiment and material omissions. | Automated labels should be checked by a person when decisions depend on them. |
Design a prompt set that represents buyers
Cover intent clusters
Build prompts from questions customers really ask, not only keyword lists. Include category discovery (for example, “best invoice automation for a small agency”), comparisons (“Brand A vs Brand B”), trust and diligence (“is Brand A reliable?”), and local or language-specific questions when geography matters. Inner Labs gives examples such as “best X in India,” “X vs Y,” and “is X reliable.”
Freeze prompts and competitors
Choose a competitor list before the measurement window. Relative share changes when the list changes. If you add a competitor, keep a common subset so the new period can still be compared with the old one, and record the change date.
#1 Best Overall
Record prompt metadata
- Prompt text and intent cluster
- Target product, category, location and language
- Tracked brands and domains
- Expected answer type (list, comparison, explanation or local result)
Choose engines, surfaces and run conditions
Track the platforms your audience uses, such as ChatGPT, Gemini, Claude and Perplexity. Track Google AI Overviews or AI Mode separately because their retrieval and citation behavior differs from chat interfaces. One platform is not a proxy for all platforms.
Define a schedule (for example, weekly or monthly), a run count per prompt, and comparable session conditions. Record model name or release when visible, date and time window, locale, language, logged-in state, tools or browsing mode, and temperature or randomness controls when the interface exposes them. A single response is an anecdote; repeated runs provide an estimate of a variable system.
Run a repeatable measurement procedure
- Prepare the registry. Assign each prompt an ID and store its frozen text, intent, location, language and competitor set.
- Execute runs. Run every prompt on every selected surface for the same window. Use the same session settings and capture the complete answer and visible source list.
- Archive raw evidence. Save timestamp, engine, model, prompt ID, answer text, cited URLs, screenshots or exported HTML, and any error state. Never overwrite an earlier run.
- Code outcomes. Mark each tracked brand as mentioned, recommended, cited, found-but-not-cited (when the product exposes that distinction), position, sentiment and factual concerns.
- Review exceptions. Have a human verify negative, surprising or potentially inaccurate claims before publishing a trend.
- Aggregate only after coding. Calculate metrics by platform, model, intent, geography and language before producing an overall figure.
Use explicit formulas
Publish the denominator with every result. For a frozen set of N runs, mention rate is:
mention rate = runs containing at least one brand mention / N
Recommendation and citation rates use the same structure, replacing the event. If you count brand mentions for share of voice, define:
share of voice = your counted mentions / counted mentions for all tracked brands
If you use an impressions-weighted vendor metric, identify it as such and document the weighting. Ahrefs Brand Radar describes impressions as summed Google search volumes for prompts where a brand appears and distinguishes cited sources from pages found during retrieval but not cited. Do not compare that number directly with a run-based share.
Build a practical data model
A spreadsheet works for a small study; a database or append-only files are safer for recurring measurement. Use one row per prompt run and fields such as:
Free tools Windows power users keep installed
One-click scans. No signup required.
run_id, timestamp, platform, model and session settings- prompt ID, exact prompt text, intent, geography and language
- raw answer and source list
- one Boolean column per brand for mention and recommendation
- citation domain, found-but-not-cited status and position
- sentiment, accuracy flag, reviewer and review note
Keep raw answers immutable. Store coding corrections as a new version so a later audit can reconstruct the original result.
Manual logging versus commercial platforms
| Approach | Strengths | Trade-offs and checks |
|---|---|---|
| Manual prompt log | Direct inspection, complete control over prompts, competitors and sample. | Ongoing labor; you must build exports, deduplication, model-change notes and quality review. |
| Ahrefs Brand Radar | Documents mentions, citations, found-in pages, impressions and AI share of voice. | Definitions are vendor-specific; verify current engine coverage, limits and pricing before purchase. |
| Yext Scout | Describes prompt-based visibility scoring and competitor tracking across named AI platforms, useful for local visibility. | Coverage and score construction are vendor claims; verify current details and retain raw answers where possible. |
Compare any tool on engine and surface coverage, prompt freezing, run count and session controls, citation versus retrieval visibility, formulas and denominators, raw-response export, competitor and location segmentation, model-change handling, human review and total cost. The available descriptions do not establish an independent accuracy or price comparison.
Rank #3
Interpret trends without overclaiming
- Report platform-level results before an aggregate; an overall gain can hide a loss on one important surface.
- Show sample size, date window, prompt and competitor sets, formula and whether demand figures are measured or estimated.
- Segment by category, product, geography and language when those slices affect buying decisions.
- Annotate model releases, interface changes and prompt-set edits.
- Treat position and sentiment as directional. Investigate material negative or inaccurate statements manually.
- Do not equate visibility with traffic, conversions or revenue. Those require separate attribution evidence.
Common failure modes and fixes
A score moves after prompts changed
Cause: the denominator changed. Fix: rerun the prior period on the common prompt and competitor set, then publish the revised scope.
One surprising answer is treated as a trend
Cause: stochastic output. Fix: increase repeated runs and report the run count and dates.
Citation rate is blank on a chat platform
Cause: the interface exposes no sources. Fix: mark citation as unavailable; do not infer it from a mention.
Vendor share-of-voice numbers disagree
Cause: incompatible denominators or impression weighting. Fix: publish each formula and compare only like-for-like metrics.
Automated sentiment labels look wrong
Cause: sarcasm, context or factual nuance. Fix: route negative or consequential cases to human review and retain the correction.
Rank #4
Archive rendered answers reliably
For browser-based collection, save the raw response first, then capture the rendered page with a timestamp and run ID. Check that consent dialogs, newsletter overlays and chat controls are not hiding the answer, and record failures instead of silently dropping them.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteOr skip the browser setup
ScreenshotNeo can archive a clean rendering when your measurement workflow needs visual evidence. Its API accepts one GET request and returns PNG, JPEG, WebP or PDF. The service accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
See the ScreenshotNeo API documentation for options such as full-page or selector capture, device and retina settings, waits, custom headers, cookies, blocking rules, caching and signed webhooks.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
How many prompts do I need?
There is no universal minimum. Choose enough prompts to cover your buyer intents and report the sample size; increase it when segmenting by platform, geography or language would otherwise leave very small cells.
Should I count a repeated brand name twice?
Decide before analysis. A response-level mention metric counts a response once, even if the name appears repeatedly; a token-level metric answers a different question and must be labeled separately.
Can visibility prove that SEO work generated revenue?
No. Visibility measures what answers show. Connect it to visits, leads and revenue only with separate, controlled attribution data.
Frequently Asked Questions
How often should an LLM visibility study be rerun?
Use a schedule that matches model and content-change risk, then keep the same prompts and session conditions within each comparison window.
What should be published with a visibility score?
Publish the exact prompt and competitor sets, engines and surfaces, run count, dates, formula, denominator, segmentation and whether each figure is measured or estimated.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

