Key Takeaways
- SparkToro found AI brand recommendations change on almost every run
- Lists, ordering and result counts vary substantially between identical prompts
- Position within AI recommendations is therefore a weak metric
- Visibility percentage across many runs is more reliable
- Different prompts nonetheless surface similar core brands that is the durable signal
- Richer context genuinely narrows the candidate set the model draws from
- Ahrefs found AI Overviews reduce clicks by 58% on affected queries
- Gemini has passed 750 million monthly active users
- Much published material on this topic is vendor marketing rather than research
Richer business context produces different answers but SparkToro’s testing found AI brand recommendations change on almost every run regardless. Any test that ignores that is measuring noise.
Summary
Adding business context to an AI prompt changes the recommendations it returns but research from SparkToro found that brand recommendations from ChatGPT, Claude and Google AI vary substantially between runs of the same prompt. That variance means single-run comparisons cannot distinguish the effect of better context from ordinary output variation. Measuring visibility percentage across many runs is the more reliable method.
The Finding That Reframes the Question
SparkToro’s large-scale testing across ChatGPT, Claude and Google AI found brand recommendations change almost every run.
Not marginally. Lists change, order shuffles, and some brands appear and disappear between identical queries. Their conclusion was that ranking positions within AI recommendations are essentially meaningless as a metric.
Two practical consequences follow.
First, single-run experiments prove very little. If you give a model thin context, get one answer, then give it rich context and get a different answer, you have not demonstrated the effect of context. You have demonstrated that the model produced two outputs. Running the thin prompt ten times would likely produce ten variations too.
Second, there is a real signal underneath the noise. SparkToro also found that even with very different human prompts, models surface similar core brands. The specific list is unstable; the underlying set is not.
That is the thing worth measuring.
What This Means for Testing Context
If you want to know whether business context changes recommendations, the method matters more than the result.
A defensible approach:
- Run each prompt variant many times, not once. Ten runs minimum, more if the category is competitive
- Measure visibility percentage — how often a brand appears across all runs — rather than its position in any single one
- Hold everything else constant. Same model version, same session state, same day. Model updates alone can shift outputs
- Compare distributions, not answers. The question is whether the frequency of a brand appearing changed, not whether it appeared once
- Test across models. A context change that helps in one model may do nothing in another
The failure mode is confirmation bias with extra steps. Add context you believe should help, run once, see your brand appear, conclude the context worked. Run the original prompt nine more times and your brand may well appear in four of them anyway.
Why Context Should Change Recommendations
The mechanism is straightforward, and the effect is real even if measuring it is hard.
More context narrows the space of plausible answers. A prompt asking for the best project management tool has an enormous candidate set. The same prompt specifying company size, industry, budget, existing stack, compliance requirements and team structure has a much smaller one.
Richer context shifts a model from generic recall toward constrained matching. That is a genuine change in the task, not just a change in phrasing which is why context engineering matters regardless of measurement difficulty.
The practical implication for businesses is that being recommended for a specific use case is a different problem from being recommended generally. If a model consistently names you when the prompt includes your ideal customer’s constraints, that is more valuable than appearing in generic lists, and it is measured differently.
The Wider Context
AI Overviews are reducing clicks by 58% on affected queries, according to Ahrefs’ analysis — which is why this measurement question has commercial urgency rather than academic interest.
Google’s Gemini has passed 750 million monthly active users, up from 650 million the previous quarter. The audience being served AI recommendations is now large enough that recommendation visibility is a mainstream marketing concern rather than an early-adopter one.
A Warning About the Sources on This Topic
This subject is unusually saturated with vendor marketing presented as research.
Searching for material on AI recommendation visibility returns a large volume of press releases from agencies and platforms selling AI visibility services. Common patterns to watch for:
- Impressive proprietary metrics with no methodology one agency claims a “90.9% AI recommendation rate” across its own monitoring data
- Four-pillar or five-step frameworks that resolve into a service offering
- Newswire distribution on financial sites, which lends the appearance of business reporting to what is a marketing document
- Urgency framing that brands which delay will be “defined by external sources” and pay more later
None of that means the underlying services are worthless. It means the statistics they publish about themselves are not evidence.
The distinction to apply: SparkToro published a testing methodology and a finding that complicates the commercial pitch. A vendor publishing a success rate for its own clients has not done the same thing.
What Is Actually Worth Doing
- Establish a baseline across many runs before changing anything
- Track visibility percentage, not position
- Test the prompts your customers would actually write, including the constraints they would mention
- Check whether you appear for specific use cases, not just category queries
- Re-baseline after model updates, which can shift outputs independently of anything you did
- Treat single-run screenshots as anecdote, including your own
Conclusion
The interesting question is not whether business context changes AI recommendations. It does, for reasons that follow directly from how these systems work.
The interesting question is whether you can tell. And on SparkToro’s evidence, a single before-and-after comparison cannot — because running the same prompt twice already produces different answers.
Which makes the discipline more important than the insight. Baseline across many runs, measure frequency rather than position, and be suspicious of any result that arrived on the first attempt.
Frequently Asked Questions
Does adding business context change AI recommendations?
Yes, because more context narrows the set of plausible answers and shifts the model from general recall toward constrained matching. Measuring the size of that effect reliably requires many runs, not a single comparison.
How consistent are AI brand recommendations?
Not very. SparkToro’s testing across ChatGPT, Claude and Google AI found recommendations change on almost every run, with lists, order and result counts varying substantially between identical prompts.
How should I measure AI visibility then?
By tracking the percentage of runs in which your brand appears, across many repetitions of the same prompt, rather than by recording your position in any single response.
Why do models still recommend the same brands despite the variance?
SparkToro found that even very different prompts surface similar core brands. The specific list is unstable, but the underlying candidate set is comparatively durable which is the signal worth optimising toward.
How many runs do I need?
Enough to distinguish a real change from ordinary variation. Ten is a reasonable minimum for a simple comparison; competitive categories with large candidate sets need considerably more.
Should I test across multiple models?
Yes. A context change that improves visibility in one model may have no effect in another, since each has different training data, retrieval behaviour and response patterns.
How much are AI Overviews affecting traffic?
Ahrefs’ analysis found AI Overviews reduce clicks by 58% on affected queries. Other studies using different methodologies report varying figures, so the source and method should always accompany the number.
Can I trust vendor statistics on AI visibility?
Treat them cautiously. Much published material on this topic is press-release marketing containing proprietary metrics with no disclosed methodology. Prefer research that publishes its method and reports findings that complicate the seller’s pitch.



