AI visibility tracking tools: evaluate them as measurement instruments
AI visibility tracking tools sample LLM answers. Get the sample-size math, a 90-minute trial protocol, per-sample cost model, and a 100-point scorecard.
Article proof

An AI visibility tracking tool re-asks a fixed list of prompts across assistants such as ChatGPT, Perplexity, Gemini and Google's AI answers on a schedule, extracts the brands and source links from each response, and rolls the results up into mention rates and share of voice. That makes it a sampling instrument, not a reporting interface. The question that separates a useful one from an expensive chart is whether the sample it takes is large enough to detect the change you actually intend to act on.
Most guides for this keyword rank vendors. This one gives you the arithmetic, a trial protocol, and a scorecard instead, for two reasons: a ranking written today describes a category that reshuffles faster than the article gets updated, and the sampling design behind the number matters more than the logo above it.
What the tool is doing when it reports "32% share of voice"
Every product in this category runs the same four steps, whatever the marketing calls them:
- A prompt set. Somewhere between 20 and several hundred questions, either uploaded by you or suggested by the vendor.
- Scheduled execution. Each prompt is sent to each selected model, at some frequency, from some location, usually logged out.
- Extraction. The answer text is parsed for brand names and cited URLs — normally by a second model or a named-entity matcher.
- Aggregation. Mentions are counted, divided by runs, and compared against a competitor list you defined.
Two consequences follow, and both get skipped in vendor demos. First, the percentage on your dashboard is an estimate drawn from a sample, so it carries a margin of error that almost no interface prints. Second, none of this measures traffic. A mention rate says something about how a model answers a question; it says nothing about how many people asked it. Those are separate measurements, and mixing them is the most common reporting mistake in this category. If you need the traffic side too, the measurement plan is a different exercise — the AI visibility tracking for SaaS: a practical measurement guide walks through mentions, position, sentiment and sources as separate series.
Your dashboard number has a confidence interval nobody prints
Treat a mention rate as a proportion measured from *n* answer samples. At a 30% baseline — a realistic level for a brand that is present but not dominant in a category — the 95% interval around that estimate looks like this:
| Answer samples (n) | Standard error | 95% interval around a 30% mention rate |
|---|---|---|
| 30 | 8.4 pp | ±16.4 pp (roughly 14% – 46%) |
| 100 | 4.6 pp | ±9.0 pp (roughly 21% – 39%) |
| 400 | 2.3 pp | ±4.5 pp (roughly 26% – 35%) |
| 1,000 | 1.4 pp | ±2.8 pp (roughly 27% – 33%) |
Now the version that matters for decisions. To distinguish two periods — say, before and after a content push — you need enough samples in *each* period. Detecting a five-point move around a 30% baseline, at 95% confidence with 80% power, takes roughly 1,320 answer samples per model per period. Below that, a five-point swing in share of voice is a coin flip with a logo on it.
One honest caveat: repeated runs of the same prompt are not fully independent observations, because the prompt wording is fixed and models are correlated with themselves. Real precision is somewhat worse than the table suggests. Use these numbers as a floor, not a promise.
Sample size decides which questions you can even ask
Translate the arithmetic into three configurations. "Smallest trustworthy change" is the month-over-month delta you could defend at 95% confidence with 80% power, again assuming a baseline near 30%.
| Setup | Prompts | Models | Runs per prompt per month | Samples per model | Total samples | Smallest trustworthy monthly change |
|---|---|---|---|---|---|---|
| Spot check | 20 | 2 | 4 (weekly) | 80 | 160 | ~20 points |
| Standard | 50 | 4 | 30 (daily) | 1,500 | 6,000 | ~5 points |
| Deep | 150 | 5 | 30 (daily) | 4,500 | 22,500 | ~3 points |
The decision rule falls out of the table. A weekly spot check across 20 prompts is a qualitative instrument: use it to read *which* competitors and sources appear, never to declare that you gained four points. If you plan to report share of voice to anyone who makes budget decisions, you need something close to the Standard row — roughly 50 tracked prompts sampled daily per model — and you should say the confidence interval out loud when you present it.
This also settles a pricing question before you look at any plan page. A tier that caps you at 25 prompts with weekly refreshes cannot produce a number that supports a five-point claim, regardless of how the dashboard renders it.
Four kinds of tool, four different blind spots
| Category | What it measures well | What it cannot see | Fits when | Breaks when |
|---|---|---|---|---|
| Dedicated AI visibility platforms | Prompt-level mentions, citations, competitor sets, source domains | Real sessions and revenue; anything outside the prompt list you defined | AI answers are a named channel with an owner and a budget | The prompt set is small, vendor-generated, or never revisited |
| AI modules inside broader SEO suites | AI mentions sitting beside rankings and keyword data in one workflow | Depth of prompt-level detail; per-answer raw text is often unavailable | You already run the suite and want one reporting surface | You need to audit raw answers or re-baseline a changed metric |
| DIY scripts against provider APIs | Exactly the prompts, locales and cadence you specify; full raw output | Nothing you do not build — entity matching, dedupe, trend storage are yours | You have engineering time and unusual coverage needs | Maintenance is nobody's explicit job |
| Server logs and referral analytics | Actual crawler hits and actual sessions arriving from assistants | Whether you were mentioned at all in answers that produced no click | You need the revenue-side counterpart to mention data | You treat it as a substitute for answer sampling — it is the other half |
The fourth row deserves a note, because it is the half most buyers skip. Assistant crawlers identify themselves with user agents such as GPTBot or PerplexityBot; the exact strings change and get renamed, so pull the current list from each provider's published crawler documentation rather than from a blog post. On the referral side, check whether your analytics keeps assistant hostnames as their own referral source or folds them into direct — verify that in your own property before you quote any number from it.
The metrics worth tracking, and two that will mislead you
Four series are worth storing every month:
- Presence rate — the share of runs where your brand appears at all, per prompt and per model. This is the base measurement everything else derives from.
- Citation rate — the share of runs where a URL you own is linked, not merely where your name is typed. Mention and citation diverge sharply, and only one of them can send a session.
- Share of voice against a fixed competitor list — fixed being the operative word. If the list changes, the series breaks. Freeze it, version it, and note the date whenever you edit it.
- Source composition — which domains the model cites when answering your prompts. This is the most actionable output of the whole category, because it tells you which third-party pages to earn a place on.
Two metrics reliably mislead. The first is any composite "AI visibility score": the formula is vendor-defined, unpublished in most cases, and revised without announcement, so your trend line can move because the arithmetic changed rather than because the model did. Keep raw mention counts in your own storage so you can rebuild an index yourself if a vendor's number jumps overnight. The second is position within the answer. Where a brand appears in a paragraph is unstable between runs and has no established relationship to what a reader does next — track presence and citation instead, and let position stay a curiosity.
The 90-minute test to run inside the free trial
Run this before the trial expires. Each step has a threshold, and each threshold is one I apply as a working rule rather than an industry standard — adjust them, but write them down before you look at the results.
1. The repeat test (25 minutes). Pick 10 prompts. Run all 10 against one model five times inside a single hour. Count how many prompts returned an identical brand set on all five runs. *Rule: if fewer than 6 of 10 are stable, treat every week-over-week move under 10 points on that model as noise, whatever the dashboard implies.* This test alone tells you more about the tool's usable resolution than any feature list.
2. The locale test (15 minutes). Run five prompts once per market you sell into. If the tool offers no locale or language control, its numbers describe one unnamed market — usually US English — and you should stop treating them as global.
3. The entity test (20 minutes). Seed three spellings of your own brand ("Acme", "Acme.io", "Acme Software") plus one competitor whose name collides with a common word. Check two things: whether your variants merge into one entity, and whether an unrelated company with a similar name inflates someone's count. Ambiguous brand names are where extraction quality either holds up or quietly fails.
4. The export test (10 minutes). Export one week of raw answers, not aggregates. If the only export is the chart or a CSV of percentages, you cannot audit an anomaly, cannot re-baseline when a definition changes, and cannot leave without losing your history. A tracker that will not hand back raw answer text is selling you a number you can never check.
5. The prompt ownership test (20 minutes). Upload your own list. Then ask where the vendor's suggested prompts come from — keyword tools, model generation, aggregated customer data. "Suggested prompts" that nobody can trace are the weakest input in the whole pipeline, because your entire measurement inherits their bias.
Price the tool per answer sample, not per month
Monthly fees are not comparable across vendors, because the units differ. Convert everything to one figure:
``` samples per month = prompts × models × runs per prompt per month cost per sample = monthly fee ÷ samples per month ```
Worked on the Standard row above — 50 prompts × 4 models × 30 runs = 6,000 samples — an illustrative $200 monthly plan lands at about $0.033 per answer sample. Do the same division for every plan on your shortlist, including the tier above the one you were quoted. Plans that look 40% cheaper often halve the refresh frequency, which raises cost per sample while quietly destroying the precision you needed in the first place.
The same arithmetic settles build versus buy. Assume a blended $0.01 per answer in model calls — substitute your provider's current published rate, since these move — and 6,000 samples cost roughly $60 in API spend. Add maintenance at an assumed loaded $80 per hour and the break-even against that illustrative $200 plan sits at about 1.75 hours of upkeep per month. Prompt rotation, extraction fixes when a model changes its output format, and storage each eat into that quickly.
Decision rule: if you cannot honestly commit to holding your own harness under two hours of maintenance per month, buy. If you have unusual locale coverage, a private model, or a compliance requirement that forbids sending prompts to a third party, build — and budget the hours explicitly rather than hoping they disappear.
A 100-point scorecard for the shortlist
Score each candidate out of 100. The weights reflect what changes the quality of the number, not what looks good in a demo.
| Criterion | Weight | What earns the points |
|---|---|---|
| Sampling design and transparency | 25 | Published run frequency, documented location and session state, per-prompt sample counts visible, confidence or variance shown |
| Metric definitions and entity resolution | 20 | Written definitions of mention vs citation, brand-variant merging you can inspect and correct, competitor list you control |
| Data portability | 15 | Raw answer export, dated API access, no lock-in on history |
| Model and locale coverage | 15 | The assistants and markets you sell into, with per-locale reporting rather than a blended average |
| Source and citation attribution | 15 | Which domains and URLs get cited per prompt, exportable as a working list |
| Workflow fit | 10 | Alerts on real thresholds, integration with your publishing pipeline, roles and permissions |
Two cut-offs make this usable. Anything under 60 total is a dashboard rather than an instrument. Anything under 15 of the available 25 on sampling design should be rejected outright even with a high total — a beautifully integrated tool built on an unreadable sample just distributes a bad number faster.
Where AI visibility tracking quietly goes wrong
Five failure modes account for most of the wasted spend I have seen in this category, and each has an early symptom.
Prompt sets that flatter you. Someone seeds prompts using your own category framing, you appear in 90% of answers, and the tool becomes a mirror. *Early symptom:* presence rates above 80% in month one. *Counter:* no prompt may contain your brand name unless it lives in a separate branded bucket reported on its own, and at least 70% of the set should be phrased the way a buyer who has never heard of you would phrase it.
Optimising against your own list. You wrote the prompts, then wrote content aimed at them, then measured with the prompts. *Counter:* lock 20% of the prompt set as a control group and never show it to whoever produces the content.
Composite scores that shift underneath you. Covered above; the tell is a step change in the index with no matching change in raw counts.
Mistaking mention rate for demand. A 40% presence rate on a question nobody asks is worth less than 5% on a question asked immediately before a purchase decision. *Counter:* tag each prompt by distance from purchase and report a weighted presence figure alongside the raw one.
No owner for the follow-up. The tool reports that three review sites and one comparison page supply most citations in your category — and then nothing happens, because publishing capacity was never part of the plan. This is the failure mode that turns a subscription into a screenshot habit. If the response to a finding has to queue behind everything else, connect the tracker to whatever produces pages; the AI SEO automation for SaaS: build a scalable content pipeline guide covers where those handoffs usually break.
Checking AI visibility without a subscription
A manual protocol gives you the qualitative half of the picture at zero licence cost. It will not give you a defensible five-point trend — see the sample-size table — but it will tell you which competitors and which source domains dominate your category's answers.
- Write 20 prompts a buyer would actually type. Store them in a sheet with one row per prompt and one column per week.
- Run each prompt in two assistants, logged out, in a fresh private window, once per week. At roughly 30 seconds per run, 40 runs take about 20 minutes; coding the answers into "mentioned / cited / absent" adds another 15 to 20. Budget 40 minutes per week.
- Record three fields per run: was your brand named, was a URL you own linked, and which three domains were cited.
- Separately, pull assistant crawler hits from your server logs and assistant referral hostnames from your analytics. These two feeds cost nothing and cover the half that prompt sampling cannot see.
After four weeks you will have 160 observations and a ranked list of the domains models cite in your category. That list is the useful output. Chasing a percentage at this sample size is not — and knowing that is exactly what the arithmetic above buys you.
When not to buy one at all
Skip the purchase, at least for this quarter, in four situations. If you have fewer than about 30 indexable pages addressing the category, you are measuring an absence you already understand — build the corpus first. If nobody owns the follow-up work, the subscription produces charts and no changes. If your brand name collides with a common noun or another company in the same market, extraction will be unreliable at any price, so spend the budget on disambiguation — consistent naming, distinct brand strings, structured data — before you spend it on tracking. And if you are pre-launch, there is nothing to sample.
The same discipline applies to the rest of your stack; the criteria in How to evaluate SEO automation tools for your workflow transfer to this category almost unchanged.
Questions buyers ask
Which AI visibility tracking tool should I choose?
Choose on sampling design first, then portability, then coverage. Run the 90-minute trial protocol above against two or three candidates in the same week with the same prompt list, score them on the 100-point scorecard, and pick the one that clears 60 with at least 15 of 25 on sampling. Published rankings in this category go stale quickly and rarely disclose the sample sizes behind their claims — verify current features and terms on each vendor's own documentation before committing.
How do I track AI visibility?
Fix a prompt list, run it against your chosen assistants on a schedule, and record three things per run: whether your brand was named, whether a page you own was cited, and which domains were cited instead. Store raw counts, not just percentages. Then pair that with server-log crawler hits and referral data so you can see both the answer side and the session side.
How can I check AI visibility for free?
Use the manual protocol in the section above — 20 prompts, two assistants, weekly, roughly 40 minutes — plus your existing server logs and analytics. Some vendors offer limited free tiers; check the current terms on their own pricing pages rather than in comparison articles, since tier limits change often and a low prompt cap will not support trend claims anyway.
Do these tools work outside English-language markets?
Test it rather than assume it. Run five prompts in each target language and market during the trial, and check whether the tool reports per locale or blends everything into one average. A blended number across markets with different competitive sets is not interpretable, and locale control is one of the more common gaps in this category.
How often should I re-examine the prompt set?
Quarterly, with a written changelog, and never in the same month you are trying to read a trend. Every prompt added or removed breaks comparability, so batch the edits, date them, and keep the previous set running in parallel for one cycle if a decision depends on the series.