Guide

How to Measure AI Share of Voice Without Paid Tools

The short answer

How do I measure AI visibility without paying for a tool?

Run a fixed prompt battery by hand: 20–30 questions your buyers actually ask, run monthly across 4–5 answer engines (ChatGPT, Perplexity, Google AI Overviews, Gemini, Copilot), with every answer scored in a spreadsheet as cited, mentioned, or absent. Share of voice is the percentage of prompts where your brand appears. The whole apparatus is a prompt list, a calendar slot, and a spreadsheet — it costs nothing but an afternoon.

Share of voice in AI search — the share of relevant prompts where an answer engine mentions or cites your brand — does not require a paid tracker to measure. It requires a fixed list of prompts, a recurring calendar slot, and a spreadsheet. That is the entire apparatus; our own 3 production brands carry no paid visibility tool anywhere in their stacks [our data].

This page documents the method end to end, including its real weaknesses. The weaknesses matter: this is sampling, not surveillance, and pretending otherwise is how the niche's worst data gets manufactured.

What counts as AI share of voice?

Share of voice is a percentage: prompts where your brand appears, divided by total prompts in the battery, per engine, per month. If ChatGPT surfaces your brand on 6 of 25 prompts, that is 24% for the month.

Keep two rates, not one. Cited means the engine linked your page — there is a path for a user to reach you. Mentioned means your brand appeared with no link — reputation without a click path. The two move independently, and conflating them hides real changes. Pew's browsing study is the reason clicks alone cannot be your metric: users clicked a cited source in just 1% of visits that included an AI summary (March 2025 data, Pew Research Center) — most of what an assistant says about you never shows up in analytics, which is also why GA4 referral tracking can only ever see a fraction of your visibility.

How do you build the prompt battery?

Build it from real demand, not from what you wish people asked. Draw on three feeds: Search Console queries the brand already earns impressions for — the feed our own GSC mining pipeline exists to surface [our data] — questions that come up on sales calls, and the comparison questions competitors get asked.

A battery of 20–30 prompts per brand is the practical range — small enough to run by hand in an afternoon, large enough that one flipped answer moves the percentage by only 3–5 points instead of swinging it. Structure it in four bands:

BandPromptsExample shape
Category questions~10"how does [category] work", "best way to [job-to-be-done]"
Comparison questions~6"[option A] vs [option B]", "alternatives to [incumbent]"
Brand questions~4"what is [brand]", "is [brand] legit"
Money questions~6"how much does [service] cost", "is [service] worth it"

Then freeze it. The frozen battery is the method: if the prompts drift month to month, the trend line measures your prompt-writing, not your visibility. New prompts go in a separate, dated section.

How do you run and score a monthly pass?

Run every prompt in each engine — ChatGPT, Perplexity, Google for AI Overviews, Gemini, Copilot — in a fresh chat, on a noted date, and screenshot anything that mentions your brand or names a competitor. For Google, the AI Overview either appears on the query or it does not; both are data. Google's own documentation is clear that no special markup earns you a slot there — eligibility is ordinary indexing and snippet eligibility (Google, AI Features) — so what you are sampling is the output of normal publishing work.

The spreadsheet carries one row per prompt per engine: date, engine, prompt, result (cited / mentioned / absent / wrong), linked URL if any, competitor named if any, and a one-line note. The "wrong" category earns its column — an engine confidently misstating your pricing or product is a finding that routes straight into the correction workflow.

Score trends, not runs. Answer engines are non-deterministic — the same prompt can return different answers on the same day — so a single flipped answer means little, and three consecutive months of the same flip means a lot.

What are the honest limits of the manual method?

Four limits, none fatal, all worth stating:

  • It is sampling. A monthly pass is 1 sample of a system that shifts daily. You will miss short-lived changes; you will still catch durable ones.
  • Non-determinism adds noise. Expect single-prompt flicker. The frozen battery and monthly cadence are what average it out.
  • Some answers are frozen in training. Kevin Indig's 2026 analysis found 24% of ChatGPT answers are generated without fetching any live page (Growth Memo, 2026). A share of what assistants say about you comes from training data, and no edit you publish this quarter reaches it quickly.
  • Personalization and location leak in. Clean sessions reduce this; they do not eliminate it. Note the conditions and keep them constant.

These limits apply equally to paid trackers — they sample the same non-deterministic systems, just more often. Frequency buys smoother trend lines, not access to ground truth.

When should you pay for a tool instead?

When the battery outgrows the afternoon. Managing many brands, needing daily alerting, or reporting to stakeholders who want dashboards are all real reasons to automate, and the paid tier exists for them — the honest comparison of what Profound, Peec, and Otterly charge and deliver is in our AI visibility tools pricing breakdown.

Against our own interest as people who sell measurement-heavy builds: most single-site operators should not buy a tracker yet, and should not hire anyone to run this method either. The spreadsheet version takes one afternoon a month, produces evidence you can act on, and tells you within a quarter whether AI visibility is even a material channel for your niche. Spend the tool budget on the thing the battery measures — pages worth citing, which is the entire subject of the operator's guide to generative engine optimization.

Frequently asked questions

How do I track AI citations for free?

Run a fixed battery of 20–30 buyer questions through the major answer engines monthly and log every mention or citation in a spreadsheet. Share of voice is the percentage of prompts where you appear. No paid tool is involved at any step — our own brands use none [our data].

How many prompts do I need to measure AI visibility?

Enough to see movement without drowning in busywork — 20–30 per brand is the practical range. Below roughly 20 the month-to-month percentages get too noisy to read; beyond about 30, manual runs stop being sustainable and consistency slips before insight improves.

Why do I get different answers to the same prompt?

Answer engines are non-deterministic: the same prompt can produce different sources and phrasing across runs, sessions, and days. That is why the method scores monthly trends across a 20–30 prompt battery rather than treating any single answer as the truth.

Are paid AI visibility tools worth it instead?

They automate the same sampling at daily or weekly frequency, which matters once you manage many brands or need alerting. For 1 site, a monthly manual battery usually surfaces the same wins and losses. Our own 3 production brands run on free measurement, with no paid tracker anywhere [our data].

Sources

  1. The State of AI Search Optimization 2026Growth Memo (Kevin Indig)
  2. AI Features and Your WebsiteGoogle
  3. Google users are less likely to click on links when an AI summary appears in the resultsPew Research Center