Ask five people what an AI visibility baseline is and you will get five different answers, and most of them are wrong, or at least incomplete. A baseline is not the first time you ran your prompts. It is not a single number you pull the week before a pitch and drop into a slide.
Treat it that way and you will spend the next six months arguing about whether a three-point swing is real progress or noise, with no data to settle it either way.
I use Peec to run this for every client, and the biggest mistake I see, mine included in the early days, is treating the baseline as an event instead of a process. Get this part wrong and everything you report afterwards is built on sand.
A baseline isn’t a day, it’s a window
I used to think the trick was picking the right moment: run the numbers the week before a pitch, or right before a campaign kicks off, so you have a clean before-and-after to compare. A few people in this space still teach it that way. I no longer think it holds up, and there is real research behind why.
AI search results are noisy in a way classic SEO rankings never were. A study built specifically to test this, Don’t Measure Once: Measuring Visibility in AI Search, found roughly 65% turnover in which sources get cited from one day to the next, for the exact same prompt. Aleyda Solis, whose AI search prompt library methodology I lean on for a lot of this, builds her own approach around exactly this instability, treating a single run as close to meaningless and testing repeatedly before drawing any conclusion. One run is not signal. It is coin-flip territory.
So the question was never which day to pick. No single day, whichever one you choose, is a baseline on its own. The same research puts a number on how long you need to run before the noise settles: standard error drops below 0.10 at around 10 days of rolling data, and below 0.05 at 24 days. Aleyda’s own protocol runs 3 to 5 executions per prompt before she calls anything a baseline.
Practically, that means running your full prompt set daily, or close to it, for two to three weeks before you call it a baseline, then averaging the results. The one thing that does matter is that the whole window sits before whatever change you want to measure, a campaign, a content push, a technical fix. Time the window’s end, not a single day’s start.
Build the prompt library first
You cannot baseline a prompt set that does not exist yet, and sizing it right matters more than people think. Too small and you are back to the noise problem above. Too large and you are burning tracking budget on prompts nobody in your market asks.
| Brand profile | Rough library size |
|---|---|
| Single product, loose persona segmentation | 30 to 60 prompts |
| Single product, strong persona segmentation | 50 to 100 prompts |
| Multi-product or multi-service brand | 100 to 250+ prompts |
| Enterprise or holdco with multiple verticals | 250+ prompts |
This sizing guidance is based on running AI visibility programs for clients of very different sizes, and it broadly lines up with the segmentation logic in Aleyda Solis’s prompt library methodology, even though the exact bands differ. A single-SKU brand chasing 250 prompts is wasting money. An enterprise brand trying to get away with 60 is going to end up with a baseline that only reflects its best-known product line.
Earlier in this series I covered organising prompts by funnel stage, awareness through post-purchase. That covers one axis. The one I did not go into there, because it deserved its own space, is segmenting by product or service line.
A single-product brand can get away with one library built around funnel stage alone. A multi-product or multi-service brand needs to think about both axes at once, funnel stage times product category. A skincare brand selling three product lines needs roughly three times the coverage of a single-SKU brand if it wants a baseline that represents each line, not just the one that happens to dominate what AI engines already know about the brand.
Don’t let branded prompts inflate your starting point
Here is a mistake I see constantly, agencies included. People build their baseline prompt set and it is heavy with prompts that already contain the brand name: “[Brand] reviews”, “is [Brand] good for X”, “alternatives to [Brand]”. Run those and your visibility number looks fantastic. It is also close to meaningless.
When you name the brand in the prompt, you are not testing whether AI engines discover and recommend you. You are testing whether they can read a name back to you, which they can, every time, for any brand with a Wikipedia page or a half-decent site. That is name recognition, not visibility, and it will make your baseline look stronger than your actual market position.
Your baseline prompt set should be dominated by non-branded and comparative prompts, the kind a real buyer would ask without already knowing you exist: “best vitamin C serum for sensitive skin”, “vitamin C serum vs niacinamide serum for dark spots”. Keep a handful of branded prompts in the mix, they still tell you something about reputation and sentiment, but they should not be the majority, and they should never be what you lead with when you report the baseline number.
Lock the prompt set before you pull data
Prompt-set stability sounds like a boring implementation detail. It isn’t. Tweak the wording of your prompts, add a few, drop a few, halfway through your baseline window, and you have broken the comparison before you even started. Every number after that point is measuring a different question, not a different market position.
Freeze the library once you have built and sized it properly. Run it as-is for the full baseline window. Resist the urge to “improve” a prompt mid-run just because it is not producing interesting results, that impulse alone tells you something, usually that the prompt is honest and you do not like the answer.
There are two legitimate reasons to touch the set later: a genuinely new product line launching, or a platform-level change, a new model version, a major update to how an engine retrieves sources. Aleyda’s guidance, which matches what I have seen with clients, is to re-baseline after platform changes rather than reading the before-and-after as your own performance shift. Outside of those two cases, leave it alone.
Running the baseline pull
Once the set is locked, the pull itself is mechanical.
- 1Run the full prompt library daily, or as close to daily as your tool allows, for two to three weeks minimum.
- 2Capture visibility score, share of voice, and citation rate for every run, not just an end-of-window snapshot.
- 3Break results down by topic, not just one blended number for the whole brand. I covered why in the metrics guide, a single average hides which product lines or topics are invisible.
- 4Record the exact competitor set you tracked against. Add or drop a competitor later and your share of voice will move for reasons that have nothing to do with your own performance.
- 5Average across the window and write that number down as day zero, with the date range attached, not a single date.
What good baseline data looks like
Once you have averaged across the window, a few signs tell you whether you have something usable.
The trend line should have settled, not still be swinging wildly from run to run by the time you reach the end of the window. Topic-level scores should vary, a flat identical score across every topic usually means you are not measuring anything differentiated, not that performance is genuinely uniform. Your competitor set should be documented and dated, so anyone reading the number six months from now knows exactly what it was measured against. And non-branded prompts should make up the clear majority of what you ran.
If any of those is missing, you do not have a baseline yet. You have a data pull, and there is a difference.
FAQ
How often should I refresh the baseline entirely?
Not on a calendar cadence. Only after one of the two triggers already covered, a real platform or model change, or a structural shift in your own prompt set such as a new product line. Refreshing on a schedule for its own sake resets your ability to see trends rather than protecting it.
Do I need a tool like Peec to do this properly?
You need something that can run a prompt set repeatedly and log the results. Manually querying ChatGPT by hand five hundred times over three weeks is not a real plan. Peec is what I use daily, but the actual requirement is the capability, repeatable runs, logging, topic tagging, not the specific vendor.
What if I don’t have a campaign or launch to measure against?
You do not need one. The baseline’s job is to give you a documented, dated starting point, not to prove a specific initiative worked. Even with nothing planned, three weeks of averaged, topic-tagged data on record means that whenever something does change, on your side or the platform’s, you have something real to compare it to.
Next steps
Skip the temptation to treat day one as done the moment you have a number. A single pull, however carefully timed, tells you almost nothing on its own. Run the window, lock the set, keep branded prompts to a minority, and you will have a baseline that survives an actual client conversation six months from now, not just the first slide.