How to check what ChatGPT says about your brand
To check what ChatGPT says about your brand, write the questions your customers actually ask, run each one several times in fresh sessions without personal memory, and record whether your brand is mentioned, recommended and described correctly, plus which pages are cited. Repeat on Claude, Gemini and Perplexity, in each language, and re-check weekly.
Step 1: Write the questions your customers ask
Do not start with your brand name. Most people ask about a need or a category, and the interesting question is whether you appear at all. Write 10 to 20 prompts across these categories:
- Discovery: "What are the best compostable coffee capsules?"
- Comparison: "[Your brand] vs [competitor]: which is better for an office?"
- Evaluation: "Is [your brand] reliable?", "What are the downsides of [your brand]?"
- Price: "How much does [your brand] cost?"
- Support: "How do I cancel a [your brand] subscription?"
- Current events: "What happened with [your brand] recently?"
Phrase them the way a person would type them into a chat, not as search keywords. If you sell in several countries, write each prompt in each language: answers, competitors and sources often differ by language.
Step 2: Prepare clean sessions
Consumer AI apps can personalise answers with memory, previous chats, custom instructions and location. To see what a new customer would see, use a fresh chat each time, turn off memory or use a temporary chat, and do not use an account where you have talked about your own brand. Note which model the app says it is using, and whether web search was used: an answer that searched the web and one that relied only on the model's training can differ a lot.
Step 3: Run each prompt several times
AI answers are not deterministic. Ask the same question twice and you may get a different list of brands. Run each prompt at least five times, each in a new session, and record every run separately. This is the step most people skip, and the reason most one-off checks mislead.
Step 4: Record the same fields for every answer
A spreadsheet with one row per answer works well:
| Field | What to record |
|---|---|
| Prompt, language, engine, date | So results can be grouped later |
| Mentioned? | Yes or no, including partial names and descriptions ("the Italian capsule brand") |
| Position | First, in a list, or only in passing |
| Recommended? | Whether the answer explicitly recommends your brand |
| Competitors named | Which ones, and in which order |
| Tone towards you | Positive, neutral or negative, about your brand specifically |
| Factual claims | Prices, features, dates: are they correct, outdated or wrong? |
| Citations | The URLs the answer cites |
Step 5: Repeat on the other assistants
ChatGPT is one engine among several. Claude, Gemini and Perplexity use different models and different retrieval, so the same prompt can produce a different picture. Perplexity, for example, always searches the web and cites sources prominently; other assistants mix what the model learned in training with what they retrieve. Run the same prompts, the same number of times, on each engine you care about.
Step 6: Read the results as rates, with their uncertainty
For each engine and language, compute your visibility rate (mentions divided by answers) and your recommendation rate (recommendations divided by mentions). Then be honest about how much a small sample can tell you. If you are mentioned in 6 of 10 answers, the 95% confidence interval runs from roughly 31% to 83%: a colleague who gets 4 of 10 next week has not proven that anything changed. With 40 answers, 24 mentions gives an interval of roughly 45% to 74%. More samples, narrower intervals.
Step 7: Follow the citations
Open the pages the answers cite. Are they your pages, a competitor's, a comparison article, a forum thread, a review site? If an answer quotes an old price, the cited page often shows where it came from, and sometimes it is your own outdated page. Note also that citations are not the whole story: engines read more than they cite, so a page can shape an answer without appearing in it.
Step 8: Re-check on a schedule
Answers drift as engines update their models and as new pages are published. Re-run the same prompt set weekly or monthly, in the same conditions, and compare rates, not single answers.
What are the limits of a manual check?
- Consumer apps are not controlled environments. Model versions, personalisation and location change without notice. Official APIs give more repeatable conditions, but their answers can also differ from the apps.
- It does not scale. 15 prompts, 4 engines, 2 languages and 5 runs is 600 answers per check, and each must be read and coded consistently.
- Tone is hard to judge consistently. Two people will label the same answer differently unless they agree on rules first.
- Citations are only part of the picture. Finding the uncited pages that shaped an answer takes a separate investigation.
How can this be automated?
This method is exactly what Tonecast automates. It proposes the prompts by category and language, samples ChatGPT, Claude, Gemini and Perplexity every week through their official APIs, records mentions, recommendations, tone, claims and citations for every answer, reports each rate with its 95% confidence interval, and builds the source graph of cited and formative pages behind the answers. You can see the result on fictitious brands in the public demo, or read how it works.