TECHNICAL DOCUMENT
Purpose: This document explains (1) how Semrush’s AI Visibility Toolkit measures prompt/topic “volume,” and (2) the realistic options for acquiring the underlying datasets yourself — including vendors, product types, procurement process, and cost expectations.
Scope: Focused on the volume/demand signal (e.g., an “AI Volume” of 541 for a topic), not on general LLM training data.
Last updated: July 2026
1. Executive Summary
Semrush does not measure prompt volume by counting exact questions. Instead, it licenses third-party behavioral (clickstream) data, feeds it into its own machine-learning models, clusters individual prompts into topics, and outputs an estimated demand figure per topic. The raw underlying data is proprietary and licensed — there is no open-source equivalent.
If you want the same capability, you have two broad paths:
- Rent the finished pipeline — pay for a tool (Semrush, Profound, etc.) that has already done the collection and modeling.
- Build your own — license raw clickstream data from a data vendor and build the topic-clustering and volume-modeling layer yourself.
The right path depends entirely on whether you need insights (rent) or a proprietary product/data asset (build).
2. How Semrush Measures Prompt Volume
2.1 The core technique
Semrush’s “AI Volume” is defined as the estimated number of search queries a topic receives on a selected AI platform. The key word is estimated — it is a modeled figure, not a raw count.
The reason for this modeling approach: individual prompts are too specific and unique to measure directly. A single exact prompt might be typed only once, so counting exact strings is meaningless. Semrush therefore measures at the topic level rather than the prompt level.
2.2 The pipeline (step by step)
|
Step |
What happens |
Output |
|
1. Source data |
Semrush licenses third-party data on real AI interactions (clickstream) plus Google’s keyword dataset for AI Overviews. |
Raw prompt/behavior signals |
|
2. Normalize |
Duplicates removed, phrasing simplified, original intent/semantics preserved. |
Cleaned prompt set |
|
3. Cluster |
An ML model groups related prompts moving in the same “semantic direction” into Topics. |
Topic groupings |
|
4. Estimate demand |
Third-party interaction data combined with Semrush’s ML models to estimate topic engagement. |
AI Volume (e.g., 541) |
|
5. Detect presence |
A separate 289M+ prompt/response database identifies where a brand is mentioned or cited. |
Mentions / Citations |
2.3 Two distinct data pipelines (important distinction)
Semrush runs two separate systems that are easy to conflate:
- Volume pipeline → answers “how much demand exists?” → built from licensed clickstream + ML modeling → produces AI Volume.
- Presence pipeline → answers “where does the brand appear?” → built from a 289M+ prompt/response corpus captured across ChatGPT, Gemini, Google AI Overviews, and AI Mode → produces Mentions, Citations, Cited Pages.
The 289M prompt corpus is NOT the volume base. It is the presence/visibility base.
2.4 What the numbers actually mean
- AI Volume (e.g., 541): a modeled, probabilistic estimate of query demand for an entire topic cluster — not a literal count of one prompt, and not an officially labeled “per month” figure.
- Prompts (e.g., 10.3K): the total count of distinct prompts being tracked for the domain.
- Your Mentions (e.g., 18): how many of those queries the brand actually appeared in.
2.5 Key caveats to communicate
- The metric is officially framed as a directional signal, not an exact count.
- Semrush does not explicitly stamp “per month” on AI Volume (monthly refers to the data refresh cadence, not the volume window).
- Treat AI Volume as best used for relative comparison between topics (e.g., Go Karts 541 vs. Charleston 60), not as a precise absolute number.
3. What “Prompt Data” Actually Means (Two Different Products)
Before purchasing anything, you must decide which of two fundamentally different data products you need. They are sold by different vendors at very different prices.
Option A — Raw prompt content
The actual text of questions people typed, plus responses. This is training-data-style: a large corpus of prompts/responses. Useful for building/fine-tuning models or analyzing phrasing — not for measuring demand.
Option B — Prompt volume / behavioral data
Clickstream data showing how many real users ask what, tied to what they subsequently clicked. This is the product that produces a volume-per-topic signal, and it is what Semrush licenses. If your goal is to replicate the “541” style metric, this is your product.
Rule of thumb: Volume cannot be self-generated. You cannot see other people’s queries by running your own prompts. Demand data requires a panel/clickstream source that only a handful of companies own.
4. Acquisition Options
4.1 Decision framework
|
Your goal |
Recommended path |
|
Get volume insights for marketing/strategy |
Rent a tool (Semrush / Profound). Do not buy raw data. |
|
Build your own AI-visibility product |
Buy clickstream (Option B) + build modeling layer. |
|
Train/fine-tune an AI model |
Buy raw prompt corpus (Option A) via data marketplace. |
|
One-off research / analysis |
Rent a tool or buy a small sample dataset. |
4.2 Path 1 — Rent the finished pipeline (lowest effort)
You pay a vendor who has already done collection + modeling. You get the metric, not the raw data.
|
Tool |
Notable for |
Entry pricing (approx.) |
|
Semrush AI Visibility |
Volume + presence inside a full SEO suite |
Included in Semrush plans |
|
Profound |
Leader in real-user prompt volume data |
~$99/mo (ChatGPT); depth in higher tiers |
|
Peec AI |
Clean mid-market monitoring |
~€199/mo Pro |
|
Otterly |
Cheapest credible monitoring entry |
Low-cost tier |
|
Sanbi.ai |
Lowest price point in the category |
~$37/mo (annual) |
If you only need the insight (not ownership of data), this path is dramatically cheaper than buying and modeling raw data yourself.
4.3 Path 2 — Buy clickstream / behavioral data (Option B)
This is the “same source” Semrush uses. Note: clickstream is a data category, not a company. You license a feed from a panel provider.
|
Vendor |
Strength / relevance |
|
Datos (a Semrush company) |
Very large clickstream dataset — tens of millions of users, 20B+ URLs, 185 countries; explicitly positioned for AI use cases. Closest to the Semrush source. |
|
SimilarWeb |
Large clickstream / digital-intelligence panel. |
|
Scrunch AI |
Opt-in panel joining clickstream to the same users’ ChatGPT / Claude / Gemini conversations — the exact prompt-to-behavior linkage needed for AI volume. |
|
MFour |
Combines validated survey + behavioral data across app, web, and location. |
What you receive: raw pageview/query-level logs. You still must build the topic-clustering and volume-estimation modeling yourself (the ML layer Semrush provides for you).
4.4 Path 3 — Buy raw prompt datasets (Option A)
If you want prompt text (e.g., for model training), use AI-data marketplaces:
|
Marketplace / Provider |
Notes |
|
Datarade |
Aggregator/marketplace — compare many LLM data providers on coverage, timeliness, pricing. Good first stop. |
|
Opendatabay |
Licensed marketplace for AI training/fine-tuning datasets; verified providers, no scraping. |
|
Nexdata |
Off-the-shelf + custom dataset collection/annotation. |
|
Shaip |
GDPR/HIPAA-compliant licensed datasets, free samples. |
5. How the Buying Process Works
5.1 Clickstream / behavioral data (Path 2)
There is no self-service checkout. The typical flow:
- Contact — use the vendor’s “Contact us / Request data” form.
- Scoping call — define geographies, time window, query volume, refresh cadence, and delivery format.
- Sample — vendor provides a sample extract for evaluation.
- Legal — sign a Data License Agreement (usually annual, enterprise contract).
- Delivery — feed delivered on the agreed cadence (a “rolling window,” e.g. trailing-month delivery, is negotiated into the contract).
5.2 Raw datasets (Path 3)
Closer to a real marketplace:
- Browse catalog on the marketplace.
- Request a free sample.
- License the dataset (contract-based, but generally more accessible than a full clickstream deal).
5.3 Cost expectations
|
Product |
Typical cost profile |
|
Rented tool (Path 1) |
~$37–$500+ / month |
|
Clickstream feed (Path 2) |
Five- to six-figure annual enterprise contract |
|
Raw dataset (Path 3) |
Varies widely; samples often free, full sets by quote |
6. Key Considerations Before Buying
- You are buying raw data, not a metric. Clickstream gives you logs; the topic-clustering and volume-modeling is a real engineering project you must staff.
- Volume requires a panel. It cannot be DIY-generated by running your own prompts — you can only self-build response monitoring (does my brand appear), never demand.
- Legal/compliance. Behavioral data carries privacy obligations (opt-in panels, GDPR, etc.). Confirm the vendor’s sourcing is compliant.
- No exact truth exists. All providers stress AI visibility is directional; personalization and rapid model change mean no vendor offers exact numbers.
- Refresh cadence matters. Decide whether you need daily, weekly, or monthly refresh, and price accordingly.
7. Recommendation
- If the goal is marketing/strategy insight: rent Profound or Semrush. Renting their finished pipeline costs a fraction of buying and modeling raw data.
- If the goal is a proprietary product or owned data asset: license clickstream from Datos, SimilarWeb, or Scrunch AI, and budget for the modeling layer on top.
Appendix A — Glossary
- Clickstream data: anonymized record of what real users browse/click/type, collected from opt-in panels (browser extensions, apps, ISPs). A category of data, not a company.
- AI Volume: Semrush’s modeled estimate of query demand for a topic.
- Topic clustering: grouping semantically related prompts into a single measurable unit.
- Presence data: signals of where a brand is mentioned/cited in AI answers (distinct from volume).
- Data License Agreement (DLA): the contract governing how licensed data may be used.
Appendix B — Vendor Quick Reference
Rent (tools): Semrush · Profound · Peec AI · Otterly · Sanbi.ai
Clickstream (behavioral): Datos · SimilarWeb · Scrunch AI · MFour
Raw datasets (marketplaces): Datarade · Opendatabay · Nexdata · Shaip
Note: Vendor pricing figures are approximate and change quickly. Treat them as directional and confirm on a scoping call before quoting them in any formal context.
