Wolf field guide · Updated August 9, 2026
How to monitor what AI says about your brand
AI brand monitoring measures whether ChatGPT, Perplexity, Gemini, Google AI search, and Copilot mention your brand, recommend competitors, describe you accurately, and cite sources you can influence. This guide gives you a defensible manual workflow, 25 buyer prompts, practical formulas, and a free tracking template.
Free CSV and Markdown files. No email required.
Takeaway 01
Build prompts around buyer jobs, not only your brand name.
Takeaway 02
Keep mentions, recommendations, and citations as separate signals.
Takeaway 03
Compare stable cohorts over time; never treat one answer as market truth.
Definition
What is AI brand monitoring?
AI brand monitoring is the repeated observation of generated answers for a controlled set of questions that matter to your buyers. It tells you whether an AI system names your brand, how it frames the brand, which alternatives it prefers, and which web sources appear alongside the answer.
It is not a database of everything people privately ask an assistant. Most platforms do not publish complete prompt-volume or impression logs to brands. Monitoring therefore uses a representative panel of prompts: a deliberate sample based on customer research, search demand, sales objections, category language, and real buying situations.
The terminology matters. A July 2026 DataForSEO snapshot shared by its founder found far more AI-assistant demand for the job phrase “brand monitoring” than for the emerging “GEO” vocabulary, while Google search demand favored “generative engine optimization.” It was a one-country, one-language, one-month snapshot—not a universal market study—but the practical lesson is sound: explain the job in customer language and use GEO or AI visibility where readers are researching the discipline.
Do not mix the datasets
Traditional vs AI brand monitoring
| Question | Traditional monitoring | AI brand monitoring |
|---|---|---|
| What is collected? | Published posts, news, reviews, and conversations | Generated answers and their available sources |
| How is it found? | Keyword matching or platform firehoses | Repeatable prompts run against selected engines |
| Primary unit | A public mention | One prompt × engine × market observation |
| Useful metrics | Volume, reach, engagement, sentiment | Mention rate, share of voice, relative position, citations, accuracy |
| Main limitation | Coverage and access restrictions | No complete view of private prompt demand; answers vary |
Measurement model
Six signals worth recording
01
Mention rate
The percentage of successful answers in your prompt set that mention your brand. Formula: answers mentioning your brand ÷ all successful answers × 100.
02
Competitive share of voice
Your brand mentions divided by all tracked-brand mentions in the same answer cohort. Define the competitor set before comparing periods.
03
Relative position
Where your brand appears among named alternatives. A first recommendation and an incidental final mention should not be treated as equal evidence.
04
Citation presence
Whether the engine links to your domain. Keep this separate from mentions because an answer can name you without citing you, or cite a third party that discusses you.
05
Source landscape
The domains and exact URLs used to support the answer. Group them into owned, editorial, community, directory, review, and competitor sources.
06
Description accuracy
Whether the answer gets your category, audience, pricing model, capabilities, availability, and limitations right. Record the error instead of hiding it inside a score.
A useful visibility formula
Mention rate = brand-positive observations / successful observations × 100
“Brand-positive” here means the brand was actually detected, not that the tone was positive.
A useful competitive formula
Share of voice = your mentions / all tracked-brand mentions × 100
Publish the prompt cohort, engines, markets, and competitor set beside the number.
Five surfaces, five contexts
Do not merge every engine into one unexplained score
The engines do not expose sources in exactly the same way. Official documentation says ChatGPT search can show inline citations and a Sources panel; Perplexity describes its answers as web-searched and cited; Gemini says sources are available for some responses, not all; and Microsoft says Copilot can generate a Bing query and expose a Sources button when web search is used. Google also says existing SEO fundamentals remain relevant to its AI features in Search.
ChatGPT
Record whether search was used, every inline source, and the full answer.
Perplexity
Capture cited domains and exact URLs as a first-class part of the observation.
Gemini
Record sources when shown and mark source availability explicitly when absent.
Google AI search
Keep AI Overview or AI Mode context attached to the underlying query and market.
Copilot
Record the answer, the Sources panel, and regional context when web search is involved.
Primary references: OpenAI on ChatGPT search, Perplexity on cited answers, Google on Gemini sources, Google Search AI-feature guidance, and Microsoft on Copilot web search. Product behavior changes; verify the current interfaces when you run the audit.
Manual workflow
Build a baseline in seven steps
- 01
Define the decision
Choose the buyer decision you want to observe: discovering a category, comparing products, checking trust, evaluating price, or resolving an objection.
- 02
Collect customer language
Use sales calls, support tickets, on-site search, reviews, communities, and keyword research. Preserve the phrasing customers use instead of converting every question into industry jargon.
- 03
Create a balanced prompt cohort
Include unbranded discovery prompts, competitor comparisons, branded validation questions, commercial constraints, and regional variants. Ten good prompts beat one hundred vague prompts.
- 04
Lock the test context
Document engine, experience or mode, account state where relevant, country, language, date, and prompt text. Avoid follow-up context unless conversational journeys are the object of the test.
- 05
Save raw evidence
Store the complete answer, brand mentions in order, citations, exact URLs, failures, and notes. A metric without the answer behind it is hard to diagnose or defend.
- 06
Calculate by cohort
Report successful observations, mention rate, share of voice, position, citations, and accuracy by engine and market before creating an aggregate view.
- 07
Turn gaps into experiments
Prioritize incorrect facts, missing buyer questions, recurring competitor sources, and credible third-party pages. Make one trackable change, then compare future runs without claiming causation too early.
Starter prompt library
25 prompts to adapt—not copy blindly
Replace the bracketed variables with real customer language. Keep at least half of the cohort unbranded so you can see which brands an assistant volunteers. Prompts that force your brand into the question answer a different research question: how the system describes you once it already knows who to discuss.
Discovery
- 1. What are the best [category] tools for [audience]?
- 2. Which [category] products should a small team consider?
- 3. Recommend an affordable [category] solution for [use case].
- 4. What is the easiest way to solve [job to be done]?
- 5. Which companies are known for [desired outcome]?
Comparison
- 1. Compare the leading [category] tools for [audience].
- 2. What are the best alternatives to [competitor]?
- 3. [Brand] vs [competitor]: which is better for [use case]?
- 4. Which [category] tool has the best value for occasional use?
- 5. What should I evaluate before choosing a [category] platform?
Trust and objections
- 1. Is [brand] a legitimate option for [use case]?
- 2. What are the strengths and limitations of [brand]?
- 3. Is [brand] suitable for a company in [country]?
- 4. Does [brand] support [required feature or market]?
- 5. What do customers say about [brand]?
Commercial intent
- 1. How much does [category] software cost for a small brand?
- 2. Which [category] tools offer pay-as-you-go pricing?
- 3. What is the most affordable way to track [outcome]?
- 4. Which [category] platform is best without an annual contract?
- 5. Recommend a [category] tool under [budget].
Local and situational
- 1. What are the best [category] options in [country or language]?
- 2. Recommend a [category] service for a [industry] company.
- 3. Which [category] tool works for an agency managing clients?
- 4. What should a founder use to monitor [outcome] weekly?
- 5. Which solution is best for [specific constraint]?
Research standard
What weak monitoring gets wrong
One prompt is not a benchmark
The same question can produce different answers. Use a stable panel and retain successful-run counts. If the decision is important, sample repeated runs rather than manufacturing certainty from one response.
A mention is not a recommendation
Separate a brand that is recommended, neutrally listed, criticized, cited as a source, or mentioned incidentally. These outcomes have different commercial meanings.
A citation is not automatically endorsement
A cited page may support one fact while the answer recommends another company. Read the surrounding claim and classify why the source appears.
Prompt volume is usually modeled, not observed
Google keyword volume and assistant prompt demand are different datasets. Search keywords are useful evidence for language and intent, but they are not a direct log of private AI conversations.
Visibility is not revenue attribution
Use monitoring to understand representation and competitive exposure. Measure visits and conversions with analytics, and add self-reported attribution when buyers may discover you in AI but navigate later through search or direct traffic.
Free starter kit
Run your first AI brand-monitoring audit
The CSV includes fields for prompts, engines, markets, mentions, relative position, competitors, citations, source URLs, accuracy, and notes. The weekly checklist keeps setup, collection, analysis, and action consistent.
Manual vs automated
When a spreadsheet stops being enough
Manual monitoring is the right way to learn the problem. It forces you to inspect actual answers, improve the prompt cohort, and understand which evidence matters. Keep it manual while you are testing fewer than roughly 10–20 prompts, one market, and a small number of engines.
Automation becomes useful when copying answers consumes the time you should spend diagnosing gaps. Wolf is a pay-as-you-go AI brand monitoring tool built to run defined prompts across ChatGPT, Perplexity, Gemini, Google AI search, and Microsoft Copilot; retain raw-answer evidence; detect brands and relative mention order; and collect available citations and source URLs across 100+ regional and language targets.
Wolf does not claim access to every private assistant query or promise a permanent “AI rank.” One check is one prompt against one engine and one target. A complete five-engine sweep currently uses six credits, failed checks are not charged, and paid credits do not expire.
The first 100 users receive 500 credits free.
Frequently asked questions