How Tonecast produces its numbers
Tonecast samples AI assistants repeatedly through their official APIs, reports every rate with a 95% confidence interval, measures its models against published quality targets on labelled data, versions every score so it can be recomputed, and shows an early signal only when it passes a significance test corrected for multiple comparisons. Each campaign's methodology card states all of this.
Which principles guide the methodology?
- Measured, not assumed. Model accuracy is measured on labelled data before it is shown; engine costs are measured per call.
- Every number opens. Every rate can be opened down to the answers and pages that compose it.
- Uncertainty is shown. Rates carry confidence intervals; estimated quantities are labelled as estimates.
- Nothing is overwritten. Scores, weights and lead estimates are superseded by newer versions, never edited in place.
How are the AI assistants queried?
Tonecast queries ChatGPT (OpenAI API), Claude (Anthropic API), Gemini (Google API) and Perplexity (Sonar API) through their official APIs, with web search or grounding enabled, so that answers reflect what the engines retrieve today and come with citations. Google AI Overviews is on the way. Tonecast does not scrape the consumer apps, and hard spending caps are checked before every call.
Stated limitation: answers obtained through an API can differ from what a person sees in a consumer app, which may add personal history, memory, location or a different model version. Tonecast measures the engines under controlled, repeatable conditions; it does not claim to reproduce any individual user's screen.
Why are prompts sampled several times?
AI answers are not deterministic: the same prompt on the same engine can name different brands on different runs. A single answer is an anecdote. Tonecast therefore runs every tracked prompt several times a month on every engine (6 runs on Essential, 8 on the other plans), every week, in every language tracked. A sample is one run of one prompt on one engine; every sample is stored in full and can be audited.
How are confidence intervals calculated?
Rates such as the AI visibility rate and the recommendation rate are proportions, reported with a 95% Wilson score interval, which behaves well with small samples and rates near 0% or 100%. Tone in AI is an average on a −100 to +100 scale, reported with a 95% interval on the mean. As an illustration, a brand mentioned in 6 of 10 answers has a visibility rate of 60% with an interval of roughly 31% to 83%; at 24 of 40 answers the interval narrows to roughly 45% to 74%. The interval is shown next to every rate, so a change within the interval is not presented as news.
What quality targets do the models have to meet?
Tonecast's models decide whether a page or an answer is about the subject, what tone it takes towards it, and whether the subject is mentioned in an AI answer. Each model is measured on a labelled benchmark of real pages and answers in Italian, English and French. These are the release targets:
| What is measured | Metric | Target |
|---|---|---|
| Tone (sentiment towards the subject), Italian, English, French | Macro F1 | ≥ 0.80 |
| Tone, other supported languages | Macro F1 | ≥ 0.75 |
| Relevance (is this page about the subject?) | Precision / recall | ≥ 0.90 / ≥ 0.80 |
| Mention detection in AI answers | Precision / recall | ≥ 0.90 / ≥ 0.80 |
| Tone in AI (tone of the answers) | Macro F1 | ≥ 0.80 |
These are targets, not claims of achieved accuracy. The measured values for the languages of each campaign are published on its methodology card, next to the targets, and updated at every model release. A model trained on labelled data is never scored on the same data it was trained on: its published figure comes from data it has not seen. Labels decided by people are reported apart from labels produced with machine assistance.
What happens when a model improves?
Every score records the model name and version that produced it. When a better model is released, Tonecast re-scores the stored answers and pages; the new score supersedes the old one, which is kept. The same applies to influence weights and lead estimates. Your corrections (for example "this page is not about us" or "this tone is wrong") are stored separately, applied to the indicators immediately, and become part of the labelled data that future models are measured on. A correction never edits a score.
How certain is the source graph?
The Cited tier is certain: those pages were cited in the sampled answers. The Formative tier is an estimate: pages the engines are likely to draw on, found through the search results for the prompts, entity pages and targeted searches, and weighted by an influence weight whose components are stored and shown. The Owned tier is declared by you. The methodology card states the coverage of the Cited tier and the estimated nature of the Formative tier. Every page carries the reason it was included.
When does an early signal become active?
The early signal tests whether the sources move before the answers, separately for each engine. Tonecast computes the cross-correlation between each source-side series (Formative Sentiment Index, claim prevalence, visibility precursors) and each answer-side series (Tone in AI, AI visibility rate, presence of a claim) at different lags. A signal becomes active only when all of the following hold:
- the best lag is positive, meaning the sources move first;
- the correlation is at least 0.30;
- a permutation test gives p below 0.05;
- after the Benjamini–Hochberg correction across all the pairs tested for that engine, the adjusted value (q) is also below 0.05;
- there are at least six data points, which in practice means six to eight weeks of data.
In plain words: when many pairs of series are tested, some will look correlated by chance. The correction raises the bar so that a signal is shown only when it stands out from what chance alone would produce. Until then the state is observing; after a long enough period without a lead it becomes none. An active signal is a statistical pattern observed in this campaign, not a guarantee about the future.
What is on a campaign's methodology card?
- the prompts tracked and their languages;
- the engines queried and the number of answers per engine, including failed calls;
- the sampling window and cadence;
- visibility and tone per engine, with their 95% intervals;
- the measured accuracy of the models for the campaign's languages, against the targets above;
- the state of the early signal and the rule it follows;
- the number of human corrections recorded on the campaign.
What are the limitations?
- API versus app. API answers can differ from consumer apps.
- No archive of past answers. Answers are sampled from the day a campaign starts; only the source graph can be reconstructed for earlier months.
- Formative sources are estimated. Engines do not disclose everything they read; the Formative tier is a weighted estimate, not a list of confirmed inputs.
- Correlation is not causation. An active early signal shows that sources moved first in this campaign; it does not prove that they caused the change.
- Engines change. Engines change their retrieval behaviour and citation patterns over time; weights are re-estimated every week for this reason.
- Language coverage. Quality targets are defined first for Italian, English and French; other languages have a lower target and are stated as such.