# AI visibility tools compared

What Profound, Peec, Otterly, Scrunch, and the rest of the AI visibility tracking category actually measure, and the questions to ask before you buy one.

Source: https://martenfield.com/guides/ai-visibility-tools-compared/
Published: 2026-07-31

---
Most of this category launched inside eighteen months, raised heavily, and now
sells the same headline: find out whether AI mentions your brand. The
differences that matter are underneath that sentence.

This is a functional comparison. We have deliberately not published a pricing
table, because published prices in this category change faster than a guide can
be maintained, and a stale price is worse than no price. Check current pricing
on each vendor's own page before you shortlist.

## The four things these tools actually do

Every product in the category is some combination of four jobs. Almost none of
them do all four well, and the marketing rarely tells you which one you are
buying.

**Prompt running.** Sending a set of questions to one or more engines on a
schedule. The differentiators are how many engines, how many runs per prompt,
and whether you control the prompt set or accept a generated one.

**Answer parsing.** Turning a generated answer into structured records: which
brands were named, in what order, with what sentiment, and which URLs were
cited. This is where the quality difference lives, and it is the hardest part to
evaluate from a demo.

**Change tracking.** Keeping the history so you can show that something moved.
Sounds trivial. It is the reason most spreadsheet approaches collapse in month
three.

**Recommendation.** Telling you what to do about it. Treat this part with
suspicion across every vendor in the category, including the ones with the
largest funding rounds. Nobody has a validated causal model linking a specific
site change to a specific citation outcome, because the engines do not publish
ranking behaviour and change it without notice.

## The questions that separate them

Ask these of any vendor, including us. The answers are more diagnostic than any
feature list.

### How many runs per prompt, and do you report the variance?

Generated answers are not deterministic. Ask the same model the same question
three times and you will frequently get three different vendor lists. A tool
reporting one run per prompt is reporting a coin flip and calling it a metric.

If a vendor cannot tell you their run count, they are either running once or
they have not thought about it. Both are disqualifying if you intend to
contract on the number.

### API or consumer interface?

Most tools query models through APIs, because it is cheaper, faster, and
scriptable. Your buyers use the consumer apps, which have different retrieval
behaviour, different system prompts, personalisation, and often a different
model version.

The gap is real and it is not always small. There is no fully clean answer
here, since scraping consumer interfaces at volume is fragile and against most
terms of service. What matters is that the vendor tells you which one they are
measuring, rather than letting you assume it is the one your buyer uses.

### Do you store raw responses, or only the parsed result?

This is the question that reveals whether a vendor expects to be doing this in
two years.

Parsers are wrong sometimes. Brand names collide with common words, competitors
get missed, a citation gets attributed to the wrong sentence. If the vendor
stored only the parsed output, a parser fix cannot be applied backwards, and
your history is permanently built on the old mistake. If they stored the raw
answer, they re-parse and your history corrects itself.

### Which engines, and does the list include Google AI Overviews?

Most tools cover ChatGPT, Perplexity, and Gemini. Coverage of Google AI
Overviews is the common gap, because it needs a search results provider rather
than a model API, and vendor claims about the fidelity of that data conflict
with each other.

Ask specifically how AI Overviews data is obtained. If the answer is vague, the
coverage is probably thin.

### Can I export the underlying rows?

You want one row per run, per query, per engine, per brand, with the mention
type, the cited URL, the position, and the timestamp. If a tool only exports its
own charts, you cannot audit it, and you cannot take your history with you when
you leave.

## Where the named tools sit

Broad strokes, and check each vendor's current documentation before relying on
this, because the category is moving quickly.

**Profound** is the enterprise end. Heavily funded, broad engine coverage, built
for large brands with a team to run it. The depth is real and so is the price
floor.

**Peec** grew fast in the mid-market on a cleaner, more focused product.
Strongest when you have a defined brand and competitor set and want the tracking
to be straightforward rather than exhaustive.

**Otterly** is the low-cost entry point. Genuinely useful for answering "am I
mentioned at all", and honest about being a monitoring tool rather than a
strategy platform.

**Scrunch** sits toward the enterprise end with an emphasis on brand presence
across answers rather than pure citation counting.

The category also has a long tail of tools that are a scheduled prompt runner
plus a dashboard. There is nothing wrong with that, and it is worth knowing that
is what you are paying for.

## What none of them can tell you

No tool in this category can tell you what a citation is worth to you.

Attribution from generated answers is unreliable. Referral data is incomplete,
many answers produce no click at all, and the buyer who reads your name in an
answer and searches for you directly the following week arrives as organic
brand search. Any vendor quoting you pipeline attributed to AI answers is
modelling, not measuring.

That is not an argument against measuring citation share. It is an argument for
knowing what the number is: a leading indicator of whether you are in the
consideration set when a buyer asks a model, which is a real thing to care about
and a different thing from revenue.

## If you are choosing today

Write your own prompt set first, before you look at any product. Twenty to fifty
questions your buyers actually ask, in their words, including the ones where you
would expect a competitor to win. That list is the asset. The tool is
interchangeable, and most of them will happily generate a prompt set for you
that measures the questions you are already winning.

Then run it manually once, across the engines you care about, three runs each,
logged out. It takes an afternoon. You will learn more from reading twenty raw
answers than from any dashboard, and you will know what you are asking a vendor
to automate.

## Questions

### Do I need an AI visibility tool at all?

If you only want to know whether you are mentioned, you can answer that yourself in an afternoon with a spreadsheet and a query list. A tool earns its cost when you need the same measurement repeated on a schedule, across several engines, with the history kept so you can prove something changed.

### What is the difference between mention tracking and citation tracking?

A mention is the model naming you in the text of an answer. A citation is the model linking a source it drew from. They move independently. You can be mentioned constantly without ever being cited, which tells you the model knows of you from training data but is not currently retrieving you.

### Why do two tools give different numbers for the same brand?

Because they are asking different questions, at different times, from different locations, with different prompt sets, and often through different interfaces. Generated answers vary between runs. Any tool reporting a single figure without stating run count, date, and prompt set is reporting an anecdote.