Before Bowstone published a single finding from its weekly AI visibility data, we ran every prompt ten times.
Not across ten weeks. Ten times in a single session, on the same day, with nothing changing in the world — just to see how much the results varied on their own.
What we found changed how we report everything.
The Problem With AI Visibility Data That Nobody Talks About
AI model outputs are stochastic. Ask the same question twice and you may get a different answer — not because anything changed, but because of randomness built into how large language models generate text.
This creates a fundamental measurement problem. If you run a prompt once a week and report the result, you have no way of knowing whether a change in the data reflects something real — a franchise's brand momentum shifting, a player entering the cultural conversation — or just natural variance in an inherently noisy system.
Most AI visibility tools don't address this. They report a number. They might report week-over-week change. But they don't tell you how much that number moves when nothing is changing at all.
We wanted to know that number before we published anything.
What We Did
We selected 10 prompts from the Bowstone battery, spanning three categories:
- Franchise value prompts — questions asking which teams or franchises are most valuable, framed to elicit ranked responses
- Performance and ranking prompts — questions asking which teams or players are best right now, across specific leagues
- Open-ended recommendation prompts — questions asking which franchise or athlete a brand should partner with, or which league has the strongest brand
We ran each prompt ten times across Claude, GPT-4o, and Perplexity in a single session with a delay between calls. Nothing changed between runs — same prompts, same models, same day. The only variable was the randomness inherent in the models themselves.
300 total API calls. One day. Zero real-world events.
For each entity that appeared in the results, we recorded how many of the 10 runs it appeared in — its appearance rate — and, where applicable, its rank position across those runs. We then computed a stability score per prompt on a scale from 0 (completely random) to 1 (identical every time).
What We Found
The results split cleanly into two groups: prompts where the models largely agree, and prompts where they don't.
Stability scores by prompt category:
| Prompt Category | Stability Score |
|---|---|
| Franchise value — NFL | 0.95 |
| Franchise value — NBA | 0.88 |
| Player ranking — NHL | 0.83 |
| Team brand strength — MLB | 0.80 |
| Team performance — NBA | 0.77 |
| Franchise recommendation — national brand | 0.745 |
| Player ranking — NFL | 0.741 |
| Team recommendation — sponsorship | 0.631 |
| League brand strength — cross-sport | 0.543 |
| Athlete recommendation — global brand deal | 0.442 |
The stable end: Franchise value prompts are remarkably consistent. The top NFL franchise by value appeared at rank 1 in every single run across Claude and Perplexity. The top four franchises were identical across nearly every run. That's a publishable finding. When Bowstone reports which franchise leads AI visibility for brand value, that claim is grounded in near-perfect consistency across repeated measurement.
The volatile end: Open-ended recommendation prompts are a different story. The global brand ambassador prompt scored 0.44 — barely above random. Claude recommended the same athlete in 10 of 10 runs. GPT-4o recommended a different athlete in 6 of 10. Perplexity split its responses across two different athletes across its runs. There is no consensus answer. Reporting a weekly "top brand ambassador" ranking from a single run of this prompt would be publishing noise.
The pattern makes intuitive sense. There is broad consensus about which franchises are worth the most money — that information is well-documented and consistent across training data. There is far less consensus about which athlete a brand should hire right now — that's a judgment call, and models diverge accordingly. Stable prompts have stable answers. Volatile prompts reflect genuine model disagreement, not just measurement error.
The Retrieval vs. Parametric Split
One of the more striking findings was the divergence between Perplexity and the other models.
Perplexity is retrieval-augmented — it searches the web at query time and cites its sources. Claude and GPT-4o are primarily parametric — they draw on knowledge from their training data, which may be months or years old.
In the calibration data, this distinction shows up as systematic divergence. For a prompt about which NBA teams are best right now, Claude returned results reflecting the current NBA landscape. GPT-4o returned results reflecting a world before a major mid-season trade — listing a player who had since moved to a different team as a key reason to rank his former franchise first. Perplexity, drawing on live sources, had already updated.
This isn't a bug. It's a fundamental characteristic of how different models work. But it means that when you aggregate across models — as most AI visibility tools do — you're blending two different things: what's happening now and what was true months ago. Bowstone tracks all models separately precisely because that divergence is information.
What Changed After Calibration
The calibration study directly shaped how Bowstone presents its data:
Confidence indicators. Every entity in the Bowstone dashboard now displays a confidence indicator — a colored dot — reflecting the stability of the prompts that contributed to its score. Teal means high confidence (stability ≥ 0.80): week-over-week movement on this entity is likely a real signal. Yellow means medium confidence (0.60–0.79): large movements are probably real, small movements should be interpreted cautiously.
Different language for different prompt types. Bowstone's weekly newsletter draws on value and ranking prompts — the stable ones — for its headline findings. Open-ended recommendation prompts are used for qualitative context, not quantitative claims.
The noise floor is now known. Before calibration, a week-over-week score change could have been signal or noise. After calibration, we know which prompts are stable enough to treat small movements as meaningful and which require larger movements before they're worth reporting.
What This Means If You're Trying to Improve Your AI Visibility
The calibration data surfaces something practically useful for any sports organization tracking its AI presence.
Focus on the stable prompts first. Value prompts and ranking prompts are where AI models have the most consistent answers — and therefore where your brand's position is most durable. Appearing consistently in franchise value or top teams responses is a stronger signal than appearing in a volatile recommendation prompt.
Don't optimize for a single model. The divergence between Claude, GPT-4o, and Perplexity means that an AI visibility strategy focused on one platform is leaving exposure on the table — or misreading its own position. A franchise that appears consistently in Claude responses but rarely in Perplexity is in a different position than one with consistent cross-model presence.
Appearance rate matters more than mention count. A franchise appearing in 9 of 10 runs at rank 2 is in a stronger position than one appearing in 3 of 10 runs at rank 1. Single-run mention counts are noisy. Appearance rate — measured across repeated runs — is the more reliable metric.
The Broader Point
Bowstone's calibration study is not a feature. It's a prerequisite for honest reporting. The weekly index means more because we know what the baseline variance looks like — and we built the confidence system around it rather than pretending the noise doesn't exist.
The calibration will be repeated quarterly and after any major model release. The full methodology, including stability scores and confidence thresholds, is published at bowstone.ai/methodology.
Bowstone tracks 250+ sports entities weekly across Claude, GPT-4o, Gemini, and Perplexity. Subscribe to the weekly digest or explore the dashboard.