A board member asks the CMO whether the bank shows up when business owners ask ChatGPT where to open a small business checking account. Someone on the team types the question, takes a screenshot of an answer that names the bank, and the question seems settled. A week later a colleague asks the same question and gets a different list.
Each screenshot shows a real answer, but one answer is not a measurement. AI answers vary from one run to the next, differ by platform, and change when the customer mentions a place or a need. The bank needs to know how often it appears across many answers, on each platform, for each product it cares about, and whether what those answers say about it is correct.

Start from the questions customers ask, and keep them fixed. For each priority product, write the questions a customer would ask at each stage of choosing, from understanding the options to comparing providers and checking a specific bank before applying. Phrase them the way customers do, including the versions that name a city, a type of business or a life situation, since answers change when those details appear. Once the set is agreed, do not edit questions in place. Add or retire them with a date, so month-to-month changes reflect the answers rather than the questions.
Run the set on each platform separately. An index of banks built by 5W, a communications firm, covered 31,500 prompts across five engines and 75 institutions, and found Google AI Overviews the most volatile of them week to week. Its author, the firm's founder, recommends tracking results monthly across engines, so report each platform on its own.
Run each question more than once. Research by SparkToro with Gumshoe, an AI visibility vendor, repeated 12 recommendation prompts nearly 3,000 times and found that the same list of recommendations came back in fewer than 1 in 100 runs. How often a brand appeared across many runs was more stable than where it ranked in any single answer. The study used no banking questions, but the lesson carries over. Report visibility as a rate across repeated runs, and do not read much into position within one answer.
Record the fields in the table below for every answer.
| Field | What to record |
|---|---|
| Named | Whether the bank appears in the answer text |
| Linked | Whether the answer links to the bank's own site, and to which page |
| Described | Whether rates, fees, eligibility and product names match the bank's current terms |
| Alongside | Which other banks, fintechs and comparison sites appear in the same answer |
| Sources | Which domains the answer cites, split between the bank's site and third parties |
| Context | Platform, mode, date and any location or detail in the question |
The bank's own analytics add a second, partial view. Google counts traffic from AI Overviews and AI Mode in the Search Console Performance report, within the Web search type, and since August 2026 a separate Generative AI performance report also shows impressions of a site's pages in those features by page, country and device. Bing Webmaster Tools' AI Performance report shows how often a site's pages are cited in Copilot answers. Google Analytics 4 now has a default AI Assistant channel for visits from assistants such as ChatGPT, Gemini and Copilot, which excludes Google's own AI Overviews and AI Mode. OpenAI says ChatGPT adds utm_source=chatgpt.com to referral links for sites that allow its search crawler. Perplexity documents that its answers link to pages, but not how those visits identify themselves. None of these numbers captures an answer that shaped a customer's choice without a click, which is why the answers themselves remain the primary measure.
A small set of indicators, each defined in writing and read by product and platform, works better than a long dashboard or a single score for the whole bank. Tools define citations differently. Ahrefs counts a citation when an answer links to the brand's pages, while Peec AI counts a source citation when the brand's content is used even if the brand is not named. A bank that compares numbers across tools or vendors without fixing the definitions first will be comparing different things.
| Indicator | What it tells the CMO | How to calculate it |
|---|---|---|
| Mention rate | How often the bank is part of the answer at all | Answers naming the bank, divided by answers in the set, per product and platform |
| Share of voice | The bank's presence against the names customers see instead | The bank's mentions, divided by mentions of all tracked banks, fintechs and comparison sites in the same answers |
| Link share | Whether the bank's site is used as a source when the bank is named | Answers naming the bank that link to its site, divided by answers naming the bank |
| Accuracy rate | Whether what customers read about the bank is correct | Answers describing the bank's terms correctly, divided by answers that describe them at all |
| Source mix | Who is shaping the answer for each product | Cited domains grouped into the bank's site, publishers, comparison sites, reviews and forums |
| AI referral outcomes | What the visits that do arrive go on to do | Sessions from AI assistant referrals and their applications or qualified inquiries, from the bank's analytics |

Link share is the indicator to read against mention rate. When it is low for questions about the bank itself, the gap usually points to the bank's own pages. For comparison questions, it points to the outside sources.
Accuracy deserves a place next to visibility, and it is easy to leave out. An answer that names the bank with last year's rate, or with a fee the account no longer charges, should go to the fix list rather than the visibility report. It is also the indicator the compliance team will care about most, because it shows where outdated or wrong information about the bank's products is circulating.
For leadership, report mention rate, share of voice, link share and accuracy by product and platform each month. Report business outcomes each quarter once there is enough data to assess them.
Tools and agencies do different jobs, and a bank may need a tool, an agency or simply a well-run internal process. A tool runs a set of questions on a schedule and stores the answers. An agency or internal team decides which questions matter, interprets the answers for each product, and works on pages and outside sources to improve those results. That is why the question set should be settled before choosing a tool.
Use the following questions to evaluate whether a tool meets the bank's needs.
Agencies are worth judging on the same basis as any other measurement partner. Ask to see the question set they would use for your products, how they would define each indicator, and an example of the monthly readout, before discussing anything else.
There is no single best tool for every bank, and none of the major tools lists a banking-specific tracking feature. Conductor has a finance industry page, but it describes general compliance and security capabilities rather than banking-specific AI tracking. The table below lists options in alphabetical order, with what each maker says it tracks. It is a starting point for evaluation and not a ranking.
| Tool | Maker | Platforms the maker lists | What it reports |
|---|---|---|---|
| AI Catalyst | BrightEdge | ChatGPT, Perplexity, Google AI Overviews, Google AI Mode, Gemini and others | Mentions and citations reported separately, sentiment, prompts mapped to buying stages |
| AI Search Performance | Conductor | ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, Grok, Microsoft Copilot, Perplexity | Brand mentions, website citations, share of voice, sentiment, AI referral traffic joined to analytics data |
| AI Visibility Toolkit | Semrush | Google AI Overviews, AI Mode, ChatGPT, Perplexity, Gemini | Visibility score, share of voice, sentiment, prompt tracking |
| AI Visibility Tracker | SE Ranking | AI Overviews, AI Mode, ChatGPT, Gemini, Perplexity | Brand mentions, linked mentions, position in answers, competitor comparison |
| Brand Radar | Ahrefs | AI Overviews, AI Mode, ChatGPT, Perplexity, Gemini, Copilot, with Claude through custom prompts | Mentions, citations, AI share of voice, estimated impressions, custom prompt tracking |
| Otterly.AI | OtterlyAI | ChatGPT, Google AI Overviews, Google AI Mode, Gemini, Perplexity, Copilot, Claude | Brand mentions, website citations, prompt research, competitor tracking |
| Peec AI | Peec AI | ChatGPT, Gemini and AI Mode named on its site, with Perplexity and DeepSeek also mentioned | Share of answers mentioning the brand, sources used and cited, rankings |
| Profound | Profound | ChatGPT, Perplexity, Claude, Gemini, Microsoft Copilot, DeepSeek, Google AI Overviews | Visibility by platform, citation analytics, prompt volumes, sentiment |
Product pages checked on 30 September 2026. Coverage and features change often, so confirm them in a trial with the bank's own questions before committing.
For ChatGPT and Perplexity specifically, the choice matters less than the setup. Load the same question set into the selected tool and confirm how it collects answers from ChatGPT and Perplexity. In the first month, manually check a sample of questions to see whether the tool's answers match what a customer sees.
Measurement is where CB/I Digital usually begins its work with a bank. We build the question set around its priority products and markets, run it on each AI platform every month, and report mentions, links, accuracy and sources by product. The readout identifies gaps in the bank's own pages, outside sources that need attention, or products that do not yet compete.
The SEO & AI Search services we offer include the monthly measurement and the fixes that follow. If your team is setting up AI visibility reporting, send us your question set and we will review it with you.
How many questions should a bank's measurement set include?
Enough to cover each priority product at each stage of choosing, and few enough to review every month. Start narrow with the products that matter most and widen the set once the monthly readout is running.
Can a bank measure AI visibility without buying a tool?
Yes, at a small scale. A team can run a short question set by hand on each platform, record the same fields every time and repeat it monthly. The effort grows quickly with more products, markets and repeated runs, which is usually the point at which a tool is worth pricing.
How long before the numbers mean something?
Treat the first month as a baseline rather than a result. Running the set weekly for the first four weeks shows how much the answers move on their own, which tells the team how large a change has to be before it counts. After that, a monthly run is usually enough, with a change in the question set recorded whenever the bank launches or retires a product.
Copyright © 2026 CB/I DIGITAL INC. All rights reserved | Privacy Policy