We Asked AI 2,349 Times What Business Card Tool Sales Reps Should Use: A 3-Month AI Visibility Experiment (2026)
A three-month benchmark running 2,349 queries across two major AI search engines revealed that brand visibility can vary by up to 23 percentage points between platforms—and the divergence is expanding. When prompts explicitly named a brand, mention rates hit 84.8%, but dropped to 48.
If you only want the bottom line: when you submit identical prompts to two different AI search engines, the probability of your brand being mentioned can differ by 23 percentage points. Even more striking, that gap widened from 9 points to 23 points over a 90-day window. If you only measure your presence on a single AI engine, your understanding of your search visibility is fundamentally flawed.
This report outlines our own operational benchmark. From June 10 to September 7, 2026, we tracked 43 fixed prompts submitted continuously across two distinct AI search environments, accumulating 2,349 recorded responses. For every query, we logged four data points: whether our brand was mentioned, where we ranked in the response list, whether our domain was cited as a reference source, and which competing brands or URLs appeared alongside us.
Plenty of articles discuss Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO) in the abstract. This article skips the theory. Here is the raw data, along with four realities the numbers forced us to confront.
How the Experiment Was Run
To make sense of the figures, the methodology needs to be transparent.
We assembled 43 core prompts divided into two distinct categories. The first group consisted of unbranded/categorical queries, simulating prospects who do not yet know which solutions exist (e.g., "What is the best business card management tool for sales reps in Taiwan?"). The second group consisted of branded queries, simulating prospects who have encountered a specific brand and want validation or comparison, where the brand name was explicitly included in the prompt.
Each prompt was systematically run across two environments: an integrated AI search overview feature embedded in a conventional search engine, and a standalone conversational AI product. The same batch of queries was executed repeatedly over a three-month span.
Every response was automatically parsed across four structured variables:
- Was our brand mentioned?
- If mentioned, what was our numerical position in the recommendation list?
- Did the output cite our website URL as a factual source?
- What other brands, URLs, and platforms appeared in the same response?
The dataset breaks down as follows:
| Metric | Count |
|---|---|
| Total Queries Run | 2,349 |
| Unique Prompt Templates | 43 |
| Timeframe | 2026-06-10 to 2026-09-07 |
| Records with Competitor Mentions | 2,253 |
| Records with Citation URLs | 2,311 |
Finding 1: The Gap Between Engines Is Wider Than Expected—and Growing
Looking at aggregate figures across the entire three months, the integrated search AI overview mentioned our brand in 63.2% of responses. The standalone conversational AI product mentioned us in 46.7% of responses. Across the exact same prompt set over the exact same period, there was a 16.5 percentage-point spread.
Breaking the data down month-by-month reveals an even more significant pattern:
| Month | Search AI Overview | Conversational AI Product | Gap (% pts) |
|---|---|---|---|
| 2026-06 | 52.1% | 43.2% | 8.9 |
| 2026-07 | 66.1% | 47.4% | 18.7 |
| 2026-08 | 70.7% | 48.8% | 21.9 |
| 2026-09 | 76.7% | 53.5% | 23.2 |
The strategic implication is critical for any team producing digital content: if you evaluate your AI visibility through only one tool, you will arrive at an overly optimistic or overly pessimistic conclusion, and your measurement error will compound over time.
While we cannot inspect proprietary engine weighting with absolute certainty, the most logical explanation lies in retrieval architecture. Search-integrated AI operates on active web indexes where traditional technical and topical SEO efforts propagate quickly into generative answers. Standalone conversational AI tools draw from narrower citation pools with different crawl schedules and weighting logic. The takeaway is practical: the performance variance across engines is real and substantial.
Finding 2: Being Known vs. Being Discovered Are Two Completely Different Things
This was the starkest finding in the dataset.
When we split the 43 prompts by query type, the difference in mention rates was dramatic:
| Prompt Type | Total Runs | Mention Rate |
|---|---|---|
| Branded Queries (Brand Name Included) | 422 | 84.8% |
| Unbranded / Categorical Queries | 1,862 | 48.2% |
If a buyer already knows your name and asks an AI engine about you, there is an 85% chance the engine will surface your product accurately. The model recognizes your entity and possesses sufficient context. However, when a prospective buyer who has never heard of you asks, "What software should I use for this workflow?", your brand surfaces less than half the time.
It is easy to misinterpret branded query data as a sign of strong top-of-funnel reach. If your monitoring relies primarily on prompts that mention your product, an 84.8% mention rate looks impressive. In reality, that only proves the model has indexed your entity—not that it will proactively recommend you. Unbranded discovery is what actually drives net-new customer acquisition.
Averaging branded and unbranded queries into a single composite visibility score obscures what matters. They must be tracked independently, and unbranded categorical queries should serve as your true acquisition metric.Finding 3: AI Citations Don't Come From Where You Think
We examined the 2,311 records that contained outbound citation URLs and ranked the referring domains by citation frequency.
| Rank | Domain | Times Cited | Source Type |
|---|---|---|---|
| 1 | Official Brand Website | 1,765 | Official Website |
| 2 | apps.apple.com | 1,356 | App Store |
| 3 | tw.my-best.com | 944 | Product Review Aggregator |
| 4 | Competitor Websites | 917 | Official Website |
| 5 | play.google.com | 762 | App Store |
| 6 | eventx.io | 687 | Industry Peer Website |
| 7 | www.facebook.com | 639 | Social Media Platform |
| 8 | kikinote.net | 549 | Tech Blog |
Combined, the Apple App Store and Google Play Store generated 2,118 citations—exceeding our own website's total. Regional product review aggregators and independent tech blogs accounted for roughly 1,500 citations. These authoritative third-party locations share one common reality: you cannot secure them simply by publishing more blog posts on your own domain.
App store authority depends on storefront metadata, category classification, release cadence, and review volume. Software review directories require direct outreach, inclusion campaigns, and profile optimization. Social platform mentions often stem from discussions outside your direct control.
This challenges a widespread assumption in digital marketing: that writing high-quality content on your own domain is sufficient for AI recommendation. In reality, generative models synthesizing product recommendations lean heavily on third-party aggregators and structured marketplace listings to corroborate facts. Your blog matters, but it is only one component of a much broader citation graph.
Operationally, an effective AEO workflow must include: auditing storefront profiles on major app marketplaces, ensuring active placement on reputable B2B software directories, and verifying that discussions on community platforms contain accurate, up-to-date product specifications.
Finding 4: AI Might Categorize You Somewhere You Never Expected
Our fourth finding emerged from analyzing the co-occurrence data: which tools, products, and categories were consistently recommended in the same breath as our solution?
Excluding raw web domains, the most common co-occurring product categories fell into three groups: business card scanners, event/trade-show software, and general-purpose CRM platforms.
The first two categories were predictable. The third was not. General-purpose CRMs appeared frequently alongside card-scanning tools. In the semantic space constructed by modern LLMs, "managing business cards" and "managing customer pipelines" do not exist as siloed software categories. The AI treats them as points along a single operational spectrum.
Event and exhibition software appeared just as consistently. This reflects the real-world workflow driving user queries: sales reps do not just collect cards in a vacuum; they return from multi-day industry expos with stacks of paper cards needing immediate triage. Consequently, the AI surfaces event capture tools alongside scanning and CRM solutions.
The takeaway is not simply mapping who your immediate competitors are. Generative models reveal how prospects actually frame their operational problems, and that framing is often broader than your internal product positioning. If an AI engine repeatedly bundles your product with an adjacent software category, it indicates buyers evaluate both tools within the exact same purchasing context. Rather than fighting that categorization, your positioning should address how your product solves that broader end-to-end scenario.
What We Changed Over Three Months
Data is only useful if it informs daily operations. Over the course of this benchmark, we adjusted our playbook in three specific ways:
- Dual-track measurement: We permanently separated our tracking into distinct branded and unbranded cohorts. All executive and acquisition reporting uses the unbranded categorical metrics. The reported visibility numbers are lower, but they reflect genuine prospect discovery.
- Third-party directory distribution: Our content production calendar previously focused almost entirely on our internal publication hub. We shifted resources to audit and maintain external software directories, regional review aggregators, and app marketplace profiles to provide clear data points for AI retrieval agents.
- Fixed-prompt recurring baselines: We retired ad-hoc testing in favor of a standardized, automated run executed at fixed monthly intervals using identical prompts. Trendlines over a 90-day period provide actionable intelligence; isolated one-off prompts do not.
Three Takeaways for Running Your Own Measurements
If your team plans to measure its own AI footprint, avoid the common pitfalls we encountered:
- Keep your prompt bank consistent: We found 40 to 50 targeted prompts to be the practical minimum required to establish statistically stable trends. If you change your phrasing every week, you are measuring prompt variance rather than engine visibility.
- Track at least two separate engines: As our data showed, performance across platforms can diverge by more than 20 percentage points. Benchmarking a single tool guarantees a distorted perspective.
- Differentiate mentions from citations: Being named in an AI answer ("Card2Gold is an option...") is not the same as having your domain cited as a source link. Citations drive direct referral traffic and confirm algorithmic trust. In our data, one engine consistently matched mentions with direct citations, while the other frequently mentioned products without generating outbound source links.
Final Thoughts
The figures in this report come from our internal tracking systems, measuring our own product's visibility within its primary target markets. That scope inherently involves parameters specific to our category: B2B sales workflows, mobile software, and specific regional markets like Taiwan.
However, the operational realities we uncovered apply across industries. AI search engines diverge significantly in citation behavior, and being indexed under your brand name does not mean you will be recommended to new buyers searching by category. You do not need to run 2,349 queries to understand those realities—but tracking the data will prove why they matter.
If you are tracking these dynamics for your own organization, the visibility tracking module inside Card2Gold runs on the exact same infrastructure used to generate this report.
Get started with Card2Gold — free
Scan cards, AI opportunity scoring, company background checks, CRM pipeline management
Try it free