Card2Gold

We Asked AI 2,349 Times What Business Card Tool Sales Reps Should Use: A 3-Month AI Visibility Experiment (2026)

TL;DR

A three-month benchmark running 2,349 queries across two major AI search engines revealed that brand visibility can vary by up to 23 percentage points between platforms—and the divergence is expanding. When prompts explicitly named a brand, mention rates hit 84.8%, but dropped to 48.

If you only want the bottom line: when you submit identical prompts to two different AI search engines, the probability of your brand being mentioned can differ by 23 percentage points. Even more striking, that gap widened from 9 points to 23 points over a 90-day window. If you only measure your presence on a single AI engine, your understanding of your search visibility is fundamentally flawed.

This report outlines our own operational benchmark. From June 10 to September 7, 2026, we tracked 43 fixed prompts submitted continuously across two distinct AI search environments, accumulating 2,349 recorded responses. For every query, we logged four data points: whether our brand was mentioned, where we ranked in the response list, whether our domain was cited as a reference source, and which competing brands or URLs appeared alongside us.

Plenty of articles discuss Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO) in the abstract. This article skips the theory. Here is the raw data, along with four realities the numbers forced us to confront.

How the Experiment Was Run

To make sense of the figures, the methodology needs to be transparent.

We assembled 43 core prompts divided into two distinct categories. The first group consisted of unbranded/categorical queries, simulating prospects who do not yet know which solutions exist (e.g., "What is the best business card management tool for sales reps in Taiwan?"). The second group consisted of branded queries, simulating prospects who have encountered a specific brand and want validation or comparison, where the brand name was explicitly included in the prompt.

Each prompt was systematically run across two environments: an integrated AI search overview feature embedded in a conventional search engine, and a standalone conversational AI product. The same batch of queries was executed repeatedly over a three-month span.

Every response was automatically parsed across four structured variables:

  • Was our brand mentioned?

  • If mentioned, what was our numerical position in the recommendation list?

  • Did the output cite our website URL as a factual source?

  • What other brands, URLs, and platforms appeared in the same response?


The dataset breaks down as follows:

MetricCount
Total Queries Run2,349
Unique Prompt Templates43
Timeframe2026-06-10 to 2026-09-07
Records with Competitor Mentions2,253
Records with Citation URLs2,311
Each AI engine was queried 1,142 times, providing identical sample sizes for direct, apples-to-apples comparison.

Finding 1: The Gap Between Engines Is Wider Than Expected—and Growing

Looking at aggregate figures across the entire three months, the integrated search AI overview mentioned our brand in 63.2% of responses. The standalone conversational AI product mentioned us in 46.7% of responses. Across the exact same prompt set over the exact same period, there was a 16.5 percentage-point spread.

Breaking the data down month-by-month reveals an even more significant pattern:

MonthSearch AI OverviewConversational AI ProductGap (% pts)
2026-0652.1%43.2%8.9
2026-0766.1%47.4%18.7
2026-0870.7%48.8%21.9
2026-0976.7%53.5%23.2
Visibility increased on both platforms over time, which is encouraging. However, the rates of growth diverged sharply. Search AI overviews climbed by 24.6 percentage points over the quarter, while the standalone conversational tool gained only 10.3 percentage points. The initial gap expanded from 8.9 points to 23.2 points—nearly tripling in three months.

The strategic implication is critical for any team producing digital content: if you evaluate your AI visibility through only one tool, you will arrive at an overly optimistic or overly pessimistic conclusion, and your measurement error will compound over time.

While we cannot inspect proprietary engine weighting with absolute certainty, the most logical explanation lies in retrieval architecture. Search-integrated AI operates on active web indexes where traditional technical and topical SEO efforts propagate quickly into generative answers. Standalone conversational AI tools draw from narrower citation pools with different crawl schedules and weighting logic. The takeaway is practical: the performance variance across engines is real and substantial.

Finding 2: Being Known vs. Being Discovered Are Two Completely Different Things

This was the starkest finding in the dataset.

When we split the 43 prompts by query type, the difference in mention rates was dramatic:

Prompt TypeTotal RunsMention Rate
Branded Queries (Brand Name Included)42284.8%
Unbranded / Categorical Queries1,86248.2%
That is a 36.6 percentage-point difference.

If a buyer already knows your name and asks an AI engine about you, there is an 85% chance the engine will surface your product accurately. The model recognizes your entity and possesses sufficient context. However, when a prospective buyer who has never heard of you asks, "What software should I use for this workflow?", your brand surfaces less than half the time.

It is easy to misinterpret branded query data as a sign of strong top-of-funnel reach. If your monitoring relies primarily on prompts that mention your product, an 84.8% mention rate looks impressive. In reality, that only proves the model has indexed your entity—not that it will proactively recommend you. Unbranded discovery is what actually drives net-new customer acquisition.

Averaging branded and unbranded queries into a single composite visibility score obscures what matters. They must be tracked independently, and unbranded categorical queries should serve as your true acquisition metric.

Finding 3: AI Citations Don't Come From Where You Think

We examined the 2,311 records that contained outbound citation URLs and ranked the referring domains by citation frequency.

RankDomainTimes CitedSource Type
1Official Brand Website1,765Official Website
2apps.apple.com1,356App Store
3tw.my-best.com944Product Review Aggregator
4Competitor Websites917Official Website
5play.google.com762App Store
6eventx.io687Industry Peer Website
7www.facebook.com639Social Media Platform
8kikinote.net549Tech Blog
Seeing our primary domain at number one confirmed that long-term, authoritative content production delivers results. However, positions two through eight tell a very different story.

Combined, the Apple App Store and Google Play Store generated 2,118 citations—exceeding our own website's total. Regional product review aggregators and independent tech blogs accounted for roughly 1,500 citations. These authoritative third-party locations share one common reality: you cannot secure them simply by publishing more blog posts on your own domain.

App store authority depends on storefront metadata, category classification, release cadence, and review volume. Software review directories require direct outreach, inclusion campaigns, and profile optimization. Social platform mentions often stem from discussions outside your direct control.

This challenges a widespread assumption in digital marketing: that writing high-quality content on your own domain is sufficient for AI recommendation. In reality, generative models synthesizing product recommendations lean heavily on third-party aggregators and structured marketplace listings to corroborate facts. Your blog matters, but it is only one component of a much broader citation graph.

Operationally, an effective AEO workflow must include: auditing storefront profiles on major app marketplaces, ensuring active placement on reputable B2B software directories, and verifying that discussions on community platforms contain accurate, up-to-date product specifications.

Finding 4: AI Might Categorize You Somewhere You Never Expected

Our fourth finding emerged from analyzing the co-occurrence data: which tools, products, and categories were consistently recommended in the same breath as our solution?

Excluding raw web domains, the most common co-occurring product categories fell into three groups: business card scanners, event/trade-show software, and general-purpose CRM platforms.

The first two categories were predictable. The third was not. General-purpose CRMs appeared frequently alongside card-scanning tools. In the semantic space constructed by modern LLMs, "managing business cards" and "managing customer pipelines" do not exist as siloed software categories. The AI treats them as points along a single operational spectrum.

Event and exhibition software appeared just as consistently. This reflects the real-world workflow driving user queries: sales reps do not just collect cards in a vacuum; they return from multi-day industry expos with stacks of paper cards needing immediate triage. Consequently, the AI surfaces event capture tools alongside scanning and CRM solutions.

The takeaway is not simply mapping who your immediate competitors are. Generative models reveal how prospects actually frame their operational problems, and that framing is often broader than your internal product positioning. If an AI engine repeatedly bundles your product with an adjacent software category, it indicates buyers evaluate both tools within the exact same purchasing context. Rather than fighting that categorization, your positioning should address how your product solves that broader end-to-end scenario.

What We Changed Over Three Months

Data is only useful if it informs daily operations. Over the course of this benchmark, we adjusted our playbook in three specific ways:

  • Dual-track measurement: We permanently separated our tracking into distinct branded and unbranded cohorts. All executive and acquisition reporting uses the unbranded categorical metrics. The reported visibility numbers are lower, but they reflect genuine prospect discovery.
  • Third-party directory distribution: Our content production calendar previously focused almost entirely on our internal publication hub. We shifted resources to audit and maintain external software directories, regional review aggregators, and app marketplace profiles to provide clear data points for AI retrieval agents.
  • Fixed-prompt recurring baselines: We retired ad-hoc testing in favor of a standardized, automated run executed at fixed monthly intervals using identical prompts. Trendlines over a 90-day period provide actionable intelligence; isolated one-off prompts do not.

Three Takeaways for Running Your Own Measurements

If your team plans to measure its own AI footprint, avoid the common pitfalls we encountered:

  • Keep your prompt bank consistent: We found 40 to 50 targeted prompts to be the practical minimum required to establish statistically stable trends. If you change your phrasing every week, you are measuring prompt variance rather than engine visibility.
  • Track at least two separate engines: As our data showed, performance across platforms can diverge by more than 20 percentage points. Benchmarking a single tool guarantees a distorted perspective.
  • Differentiate mentions from citations: Being named in an AI answer ("Card2Gold is an option...") is not the same as having your domain cited as a source link. Citations drive direct referral traffic and confirm algorithmic trust. In our data, one engine consistently matched mentions with direct citations, while the other frequently mentioned products without generating outbound source links.

Final Thoughts

The figures in this report come from our internal tracking systems, measuring our own product's visibility within its primary target markets. That scope inherently involves parameters specific to our category: B2B sales workflows, mobile software, and specific regional markets like Taiwan.

However, the operational realities we uncovered apply across industries. AI search engines diverge significantly in citation behavior, and being indexed under your brand name does not mean you will be recommended to new buyers searching by category. You do not need to run 2,349 queries to understand those realities—but tracking the data will prove why they matter.

If you are tracking these dynamics for your own organization, the visibility tracking module inside Card2Gold runs on the exact same infrastructure used to generate this report.

Share this post Threads LINE
FAQ
No. Traditional SEO measures where your specific web pages rank within a paginated list of search results. AI visibility measures whether generative models synthesize your brand as a recommended solution within an AI-generated answer. In our tests, pages with top-three organic rankings were sometimes ignored by AI summaries, while less visible third-party review pages were frequently cited. Both disciplines are related, but they rely on distinct algorithmic retrieval mechanisms and require separate tracking.
Manual prompting is fine for exploratory spot-checks, but it cannot reveal meaningful patterns. LLM responses are non-deterministic, meaning individual answers can vary from query to query. Meaningful conclusions require running identical prompt sets repeatedly, parsing the structured outputs, and analyzing trends over time. Our dataset required 2,349 structured runs across three months before shifts in visibility were stable enough to guide business decisions.
There is no universal industry benchmark, but two relative metrics matter far more than absolute percentages. First, monitor your directional trajectory month-over-month on identical prompts. Second, evaluate the spread between your branded and unbranded query rates. A wide gap (like our 36.6-point spread) indicates strong brand recognition but underperforming categorical discovery, meaning your priority should be top-of-funnel unbranded visibility.
Yes, and the playing field can be more accessible than traditional search engine results pages. Modern AI search tools rely heavily on structured aggregators, software directories, and app store listings rather than relying solely on raw domain authority. A smaller company that maintains comprehensive profiles across external directories, collects authentic reviews, and publishes clear, factual product specifications can secure AI citations against entrenched competitors who neglect their third-party footprint.
Yes. We run our benchmark tracking continuously and review our visibility shifts on a quarterly basis. A three-month window provides a clear directional signal, but tracking broader shifts driven by model architecture updates and algorithmic revisions requires ongoing longitudinal data.

Get started with Card2Gold — free

Scan cards, AI opportunity scoring, company background checks, CRM pipeline management

Try it free