AI Visibility for Shopify Brands: Why Prompt Selection Is the Decision That Makes or Breaks Your Tracking
Every brand tracking its presence in ChatGPT, Perplexity, or AI Overviews eventually asks the same question: are we actually visible, or are we just visible for the prompts we happened to pick? Most of the time, the honest answer is the second one and nobody notices, because the dashboard still shows green.
We've watched this play out with more than one Shopify brand we've worked with on AISEO. The team runs a handful of test prompts, sees the brand name show up in the response, and reports "strong AI visibility" up the chain. What they're actually measuring is whether the model recognizes the brand when asked about it directly which is closer to a spell-check than a visibility audit.
The branded-query trap
If your test prompt is "tell me about [your brand]," you will almost always get a good answer. The model has your name, it has whatever's been written about you, and it has no reason to withhold that information. This tells you the model has heard of you. It tells you nothing about whether you show up when a shopper hasn't already decided to look for you by name.
The prompts that actually matter are the ones a customer would type before they know your brand exists: category questions, comparison questions. That's where the real competition for citation happens, and it's where most brands have never actually tested themselves.
What prompt selection quietly determines
The set of prompts you choose to test isn't a neutral sampling decision; it's the thing that decides what your tracking data can even show you.
- Branded prompts measure recognition, not discoverability. Useful for confirming basic facts are correct, not for anything else.
- Category prompts (best options in a given product category for a stated need) measure whether you're competitive in the space a shopper actually starts their search in.
- Comparison prompts (how you stack up against a named alternative) measure whether your specific claims hold up, which is a different bar than being mentioned at all.
- Long-tail, specific prompts (a narrow use-case a shopper might type) measure whether your content is granular enough to be pulled for a query nobody wrote a dedicated page for.
A brand that only tests the first category will always look strong. A brand that tests all four will usually find the real gaps sitting in the last two.
Where prompt sets should come from
The instinct is to write test prompts from a brainstorm "what would someone ask about us?" That produces prompts shaped by how the brand thinks about itself, not how a shopper actually phrases a question. A more reliable source is the language a shopper has already used somewhere you can see it: search console queries, customer service transcripts, review text, and the actual questions sitting unanswered in your FAQ.
When we build a prompt-tracking set for a client's AISEO work, we pull directly from these sources rather than generating a list from assumption. It's slower to assemble, but it means the prompt set reflects real query language instead of brand-internal phrasing which is usually a more revealing test than anything a marketing team would think to ask.
Tracking correctness, not just presence
There's a second bias sitting underneath prompt selection: even with a good prompt set, most tracking stops at "were we mentioned," when the more useful question is "was what got said about us accurate." A model can cite your brand and still get the price wrong, misstate a material, or describe a discontinued product as current. That's a worse outcome than not being mentioned, because it's a wrong answer wearing your name.
A visibility check that only counts mentions will miss this entirely. One that spot-checks the actual claims in the response against your current PDP, not your memory of what it says catches it.
Building a prompt set that won't lie to you
A workable starting set doesn't need to be large. Somewhere around a dozen prompts, split across the four categories above, run consistently over time, will tell you more than fifty branded variations run once. The consistency matters more than the volume a prompt set that changes every time you test it can't show you whether things are getting better or worse, only whether today's answer happened to be good.
If NOIR & BLANCO's own AISEO audits have a starting discipline, it's this one: no client sees a "visibility score" built only from prompts that were guaranteed to work. The uncomfortable prompts the category and comparison ones are the ones that actually tell you where the work needs to go next.
Comments
Post a Comment