tonecast
Methodology

How Tonecast produces its numbers

Tonecast samples AI assistants repeatedly through their official APIs, reports every rate with a 95% confidence interval, measures its models against published quality targets on labelled data, versions every score so it can be recomputed, and shows an early signal only when it passes a significance test corrected for multiple comparisons. Each campaign's methodology card states all of this.

Last updated

Which principles guide the methodology?

How are the AI assistants queried?

Tonecast queries ChatGPT (OpenAI API), Claude (Anthropic API), Gemini (Google API) and Perplexity (Sonar API) through their official APIs, with web search or grounding enabled, so that answers reflect what the engines retrieve today and come with citations. Google AI Overviews is on the way. Tonecast does not scrape the consumer apps, and hard spending caps are checked before every call.

Stated limitation: answers obtained through an API can differ from what a person sees in a consumer app, which may add personal history, memory, location or a different model version. Tonecast measures the engines under controlled, repeatable conditions; it does not claim to reproduce any individual user's screen.

Why are prompts sampled several times?

AI answers are not deterministic: the same prompt on the same engine can name different brands on different runs. A single answer is an anecdote. Tonecast therefore runs every tracked prompt several times a month on every engine (6 runs on Essential, 8 on the other plans), every week, in every language tracked. A sample is one run of one prompt on one engine; every sample is stored in full and can be audited.

How are confidence intervals calculated?

Rates such as the AI visibility rate and the recommendation rate are proportions, reported with a 95% Wilson score interval, which behaves well with small samples and rates near 0% or 100%. Tone in AI is an average on a −100 to +100 scale, reported with a 95% interval on the mean. As an illustration, a brand mentioned in 6 of 10 answers has a visibility rate of 60% with an interval of roughly 31% to 83%; at 24 of 40 answers the interval narrows to roughly 45% to 74%. The interval is shown next to every rate, so a change within the interval is not presented as news.

What quality targets do the models have to meet?

Tonecast's models decide whether a page or an answer is about the subject, what tone it takes towards it, and whether the subject is mentioned in an AI answer. Each model is measured on a labelled benchmark of real pages and answers in Italian, English and French. These are the release targets:

What is measuredMetricTarget
Tone (sentiment towards the subject), Italian, English, FrenchMacro F1≥ 0.80
Tone, other supported languagesMacro F1≥ 0.75
Relevance (is this page about the subject?)Precision / recall≥ 0.90 / ≥ 0.80
Mention detection in AI answersPrecision / recall≥ 0.90 / ≥ 0.80
Tone in AI (tone of the answers)Macro F1≥ 0.80

These are targets, not claims of achieved accuracy. The measured values for the languages of each campaign are published on its methodology card, next to the targets, and updated at every model release. A model trained on labelled data is never scored on the same data it was trained on: its published figure comes from data it has not seen. Labels decided by people are reported apart from labels produced with machine assistance.

What happens when a model improves?

Every score records the model name and version that produced it. When a better model is released, Tonecast re-scores the stored answers and pages; the new score supersedes the old one, which is kept. The same applies to influence weights and lead estimates. Your corrections (for example "this page is not about us" or "this tone is wrong") are stored separately, applied to the indicators immediately, and become part of the labelled data that future models are measured on. A correction never edits a score.

How certain is the source graph?

The Cited tier is certain: those pages were cited in the sampled answers. The Formative tier is an estimate: pages the engines are likely to draw on, found through the search results for the prompts, entity pages and targeted searches, and weighted by an influence weight whose components are stored and shown. The Owned tier is declared by you. The methodology card states the coverage of the Cited tier and the estimated nature of the Formative tier. Every page carries the reason it was included.

When does an early signal become active?

The early signal tests whether the sources move before the answers, separately for each engine. Tonecast computes the cross-correlation between each source-side series (Formative Sentiment Index, claim prevalence, visibility precursors) and each answer-side series (Tone in AI, AI visibility rate, presence of a claim) at different lags. A signal becomes active only when all of the following hold:

In plain words: when many pairs of series are tested, some will look correlated by chance. The correction raises the bar so that a signal is shown only when it stands out from what chance alone would produce. Until then the state is observing; after a long enough period without a lead it becomes none. An active signal is a statistical pattern observed in this campaign, not a guarantee about the future.

What is on a campaign's methodology card?

What are the limitations?