"Are we showing up in ChatGPT?" is a reasonable question and a hard one to answer. Type your category into an assistant, see your brand in the list, and it feels like a win. Type it again an hour later, see a different list, and the win is gone.
The problem is not that the question is unanswerable. It is that the answer is a rate, not a yes or no. This post defines four metrics that turn a pile of screenshots into something you can track, explains how to sample prompts so the numbers mean something, and shows how to read the trends without fooling yourself.
Why one-off checks mislead
Language models generate text by sampling from probabilities. Even with identical wording, two runs of the same prompt can differ in which brands they name, in what order, and how they describe each one. Several other things add variation:
- The engine. ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews each use different models, retrieval systems and source preferences.
- Search grounding. When an assistant searches the web, the pages it finds can change from day to day.
- Location and language. Answers can shift by country and by the language of the prompt.
- Personalization and settings. Logged-in sessions, memory features and custom instructions can change what a user sees.
- Wording. "Best CRM for startups" and "what CRM should a ten-person startup use" can produce different shortlists.
So a single check is a sample of size one. If you see your brand, you may have been lucky. If you do not, you may have been unlucky. Neither tells you the rate. Run the prompt repeatedly and count.
A useful mental model: you are not asking "did we appear?" You are estimating "how often do we appear?" That is a measurement problem, and measurement needs a sample.
The four core metrics
These are the definitions we use in Citeflare. You can compute them yourself in a spreadsheet if you log each run.
Visibility score
The percentage of completed runs in a time window where your brand is mentioned at all.
If you tracked 20 prompts, ran each on three engines and took 60 completed runs, and your brand appeared in 21 of them, your visibility score is 35%.
It is the headline number and the simplest to explain. Its limits: it treats a first-place recommendation and a passing mention in a long list the same way. That is why you need the next two metrics.
Share of voice
Your mentions divided by all mentions, counting your brand and your tracked competitors, in the window.
If, across those runs, your brand was mentioned 21 times and your competitors were mentioned 63 times combined, your share of voice is 21 out of 84, or 25%.
Share of voice is relative. Your visibility score can stay flat while your share falls because a competitor started appearing more. It also depends on who you track, so keep the competitor list stable when you compare periods and note when you change it.
Average position
The mean position of your mentions, where 1 means you were named first.
Position matters because assistants often present an ordered shortlist, and people give more weight to the first names. If you are mentioned in 10 runs, first in four, second in three and third in three, your average position is (4×1 + 3×2 + 3×3) ÷ 10 = 1.9.
Two cautions. First, position is only defined when you are mentioned, so always read it alongside visibility score. Second, ordering in a generated list is not always a strict ranking. Treat it as a signal of prominence, not a leaderboard.
Sentiment
How the answer describes you: positive, neutral or negative.
Being mentioned is not enough if the mention is "popular but expensive and hard to set up." Track the count of mentions in each sentiment bucket and read the actual snippets for the negative ones. They often point at a specific fact, such as an outdated limitation or a pricing change, that you can correct at the source.
Sentiment classification is itself a judgment call, whether done by a person or a model. Spot-check it. A handful of misclassified snippets can mislead you on small samples.
Quick reference
| Metric | Question it answers | Watch out for |
|---|---|---|
| Visibility score | How often are we named? | Treats all mentions equally |
| Share of voice | How do we compare with competitors? | Depends on who you track |
| Average position | How prominent are we when named? | Meaningless without visibility |
| Sentiment | How are we described? | Classification needs spot checks |
A fifth thing worth tracking next to these is citations: which URLs the answers cite. It is not a score, but it tells you why the numbers move. See our guide on how assistants choose brands for why sources matter so much.
Sampling prompts well
Metrics are only as good as the prompts behind them. A bad prompt set produces confident nonsense.
Build the set from real buyer questions
- Use the questions customers and sales calls actually produce. Support tickets, sales notes and search queries are good sources.
- Cover the funnel: problem-aware ("how do I reduce churn"), solution-aware ("best churn analytics tools") and brand-aware ("X vs Y," "is X worth it").
- Include comparison and "alternatives to" prompts. They tend to name brands explicitly.
Keep branded and non-branded separate
Prompts that include your name will mention you almost by definition. Report branded and non-branded results separately, or your visibility score will look healthier than your discoverability is.
Size and repetition
There is no magic number, but the logic is straightforward:
- More prompts reduce the effect of any single oddly worded one.
- More runs per prompt reduce the effect of randomness in a single answer.
- More engines show you where your gaps are concentrated.
Start with a set you can sustain, such as 20 to 50 prompts, and run them on a schedule rather than ad hoc. Consistency over time matters more than a huge one-time sample.
Keep the set stable, and change it deliberately
If you change prompts every week, you cannot tell whether your numbers moved because of the world or because of your prompts. Freeze a core set for trend lines. Add new prompts as a separate group, and note the date you added them.
Reading trends without fooling yourself
- Look at windows, not days. A daily number bounces. Use a rolling window, such as 7 or 28 days, and compare like with like.
- Do not react to one point. A single bad day is noise until it persists across a window.
- Segment before concluding. A flat overall score can hide a rise on one engine and a fall on another. Break results down by engine, by topic and by country.
- Pair outcomes with causes. When visibility moves, check what changed: new citations, a competitor's new page, your own publishing, crawler activity.
- Beware small samples. If a topic has only three prompts, one run flipping changes the percentage noticeably. Report the count next to the percentage.
- Annotate events. Record launches, content updates and site changes on the timeline so you can connect them to shifts later.
A simple weekly routine
- Check overall visibility score and share of voice against the previous window.
- Look at which prompts lost or gained a mention.
- Open the citations for the largest changes.
- Pick one or two concrete fixes, such as updating a page or filling a content gap.
- Re-measure after the next window, not the next day.
What these numbers cannot tell you
They do not tell you traffic or revenue directly. Many AI answers never produce a click, and attribution is imperfect. Treat visibility as a leading indicator of being considered, and connect it to your own analytics, such as referral traffic from AI sources and branded search, as best you can. Be honest about the uncertainty.
How Citeflare helps
Citeflare runs your tracked prompts across ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews on a schedule, and computes visibility score, share of voice, average position and sentiment for you using the definitions above. You can slice by engine and topic, track competitors side by side, and drill into the cited sources behind each change. See how prompt tracking works or create an account.