Skip to main content
Back to Blog

Keyword Research for AI Search: Combine Search Demand With Buyer Prompts

A two-axis keyword research method that combines search demand with conversational buyer prompts, cited sources, commercial fit, and AI visibility gaps.

20 min read
Keyword Research for AI Search: Combine Search Demand With Buyer Prompts

Keyword research gives you a vocabulary of demand. It does not give you a complete map of the decisions behind that demand.

Someone might search Google for “AI visibility tools” and then ask ChatGPT:

Which AI visibility platform is realistic for a three-person SaaS marketing team that needs cited-source analysis, competitor tracking, multi-engine monitoring, and direct CMS publishing but cannot afford an enterprise contract?

The keyword identifies the category. The prompt exposes the operating constraints, buying stage, comparison criteria, and likely page format.

You need both.

Keyword research for AI search is a two-axis process. The first axis measures discoverability through search queries, rankings, impressions, trends, and competitive pages. The second captures conversational demand through buyer prompts, follow-up questions, model answers, brand mentions, and cited sources.

BeVisible is designed around that combined workflow: monitor the buyer prompts and sources shaping AI answers, compare them with search and commercial evidence, then turn a validated content gap into an article that can be reviewed, scheduled, published, and measured again.

The output is not two separate calendars. It is one prioritized set of topic clusters, each connected to a buyer job and the page most capable of serving it.

Why keywords alone miss part of the demand

Keywords remain useful because search engines expose valuable signals:

  • approximate demand and seasonality;
  • the language people use;
  • current ranking pages and formats;
  • the queries already showing your pages;
  • related questions and entities;
  • paid-search value as a directional commercial signal;
  • the authority and intent of competing results.

Google Search Console's Performance report shows queries, pages, impressions, clicks, CTR, and average position for your own property. Google recommends focusing more on trends in impressions and clicks than on average position alone. Google Trends lets you compare relative interest among terms, topics, regions, and periods.

But neither source gives you a clean list of everything buyers ask an AI assistant. Conversational prompts can:

  • combine several constraints in one request;
  • ask for a personalized shortlist;
  • introduce context in follow-up turns;
  • compare products by a narrow implementation detail;
  • request a recommendation rather than a document;
  • ask the model to synthesize several sources;
  • resolve the journey without a conventional click.

The mistake is not using keywords. The mistake is treating search volume as the only observable demand.

The two-axis model

Score every cluster on two independent axes before combining them.

Two-axis demand map combining search demand with conversational prompt demand to guide content investment

Axis A: search demand

Measure:

  • impressions and clicks from Search Console;
  • estimated query volume from your keyword data source;
  • trend direction and seasonality;
  • SERP intent and dominant page type;
  • current ranking or absence;
  • competition and authority gap;
  • paid value, with caution;
  • existing page and cannibalization risk.

Axis B: prompt demand and answer visibility

Measure:

  • frequency in customer, sales, support, community, or monitored prompt sources;
  • buying-stage relevance;
  • constraint richness;
  • frequency of competitor mentions;
  • frequency and concentration of cited sources;
  • own-brand mention rate;
  • owned-domain citation rate;
  • factual or positioning gaps;
  • whether one canonical page can answer the prompt cluster.

Do not normalize the two axes too early. A topic with low measurable search volume and high prompt relevance may deserve a precise product or comparison page. A high-volume topic with no product fit may not deserve content at all.

Start with the business question

Before exporting keywords, define:

  • target segment;
  • product and problem scope;
  • geography and language;
  • buying stages;
  • primary conversion;
  • category and direct competitors;
  • high-value capabilities;
  • topics the business can support with real evidence;
  • topics the site will not pursue.

For BeVisible, a useful scope is not “AI.” It is the set of buyer questions around monitoring AI answers, brand and competitor mentions, source citations, provider differences, visibility changes over time, and publishing articles from validated gaps for lean B2B teams.

That boundary filters out enormous volumes of adjacent demand with no credible connection to the product.

Step 1: build a traditional keyword baseline

Use multiple sources because each sees a different slice.

Search Console: existing demand

Export 12–16 months of query and page data when available. Keep:

  • query;
  • page;
  • clicks;
  • impressions;
  • CTR;
  • average position;
  • country;
  • device;
  • date or comparison period.

Separate branded from non-branded queries. Search Console now provides a branded/non-branded filter for eligible properties, but Google notes that the classification can be imperfect and is unavailable in some cases. Preserve your own deterministic rule if you need consistency across systems.

Remember the data limitations. Search Console omits some anonymized queries, truncates table rows, and aggregates differently by property and page. It is an excellent first-party performance source, not a census of all search demand.

Keyword database: uncovered demand

Use a reputable keyword database to expand:

  • category terms;
  • problems;
  • capabilities;
  • comparisons and alternatives;
  • jobs to be done;
  • integrations;
  • audiences;
  • industries;
  • questions;
  • pricing and purchase modifiers.

Treat search volume and difficulty as estimates. The exact number matters less than the decision it changes. A term estimated at 90 searches and one at 110 usually belong in the same practical band.

Google Trends: direction and vocabulary

Compare terms and topics across the same region and period. Google Trends allows comparisons among terms or broader topics and can reveal whether the market uses “answer engine optimization,” “AI SEO,” “GEO,” or another label.

Trends is relative, not absolute volume. Use it to understand direction, seasonality, and vocabulary, not to calculate a traffic forecast by multiplying chart values.

SERP review: current intent

For every promising cluster, inspect the current results:

  • result types;
  • page formats;
  • freshness;
  • dominant and mixed intents;
  • recurring questions;
  • authoritative sources;
  • product and publisher concentration;
  • information competitors repeat;
  • meaningful gaps;
  • whether AI features appear.

This is where a keyword becomes a page hypothesis. “Content refresh” might surface definitions, checklists, agency pages, and tools. The page you create depends on the intent you can satisfy, not the phrase alone.

Step 2: collect buyer prompts

Prompt data can come from observed behavior or deliberate research. Label the source.

First-party prompt sources

  • sales call questions;
  • support conversations;
  • onboarding responses;
  • site-search logs;
  • chatbot questions;
  • demo requests;
  • customer interviews;
  • community discussions you own or can ethically analyze.

These sources are valuable because they preserve the buyer's language and constraints. They are also sensitive. Remove personal data, respect consent and retention rules, and do not publish raw customer prompts without permission.

Monitored prompt portfolio

Create a stable set of representative prompts and run them across the AI surfaces your buyers use.

BeVisible currently monitors prompts such as:

  • “AI visibility software for small marketing teams”;
  • “AI answer tracking for startup founders”;
  • “GEO platform for content teams”;
  • “tools that show which sources AI cites”;
  • “how to turn AI visibility gaps into publishable articles”;
  • “AI visibility tools with direct CMS publishing.”

The useful part is not the list itself. It is the attached response evidence: provider, answer, own and competitor mentions, cited URLs, and run date.

Prompt expansion workshop

For every high-value job, create variants along controlled dimensions:

DimensionExamples
Company stagestartup, growth-stage, enterprise
Teamfounder, content team, agency, product marketing
Constraintlow budget, no annual contract, EU data needs, small team
Capabilitycitations, sentiment, competitor tracking, historical comparison, publishing
PlatformChatGPT, Perplexity, Gemini, AI Overviews, AI Mode
Workflowmonitor, compare, diagnose, create, review, publish, remeasure
Decisionbest, compare, alternative, worth it, how to choose

Do not generate hundreds of permutations and call them demand. Keep variants that represent a plausible decision and can be served by a meaningful answer.

Govern the prompt portfolio

A prompt list becomes a measurement system only when it is versioned. Give every prompt a stable ID and store its exact wording, buyer job, funnel stage, audience, geography, language, source, creation date, and the engines on which it should run. Record why it entered the portfolio. A prompt suggested by a sales call has a different evidentiary basis from one generated by a language model.

Split the portfolio into two sets:

  • Core measurement prompts stay stable so that changes over time remain interpretable. Use these for headline reporting.
  • Exploration prompts test new language, constraints, and categories. Promote them to the core set only after review.

Prompt portfolio governance board separating stable core measurement prompts from exploration prompts and recording version, buyer job, engine, and run date

When wording changes, create a new version instead of silently editing the old one. A small phrase can change the answer's intent, source mix, or product set. Keep the former version long enough to run an overlap test if the prompt matters to a trend line.

This governance also protects reporting from a common mistake: adding many favorable prompts and then interpreting a higher visibility score as market improvement. Report the numerator, denominator, prompt-set version, engines, and run dates alongside any percentage. If the portfolio changed, separate the effect of the new sample from changes on the stable core.

Step 3: cluster prompts by buyer job

Keyword clustering often relies on lexical similarity or shared SERPs. Prompt clustering needs one more rule: would the same page satisfy the decision?

Consider:

  • “best AI visibility tools for SaaS”;
  • “AI visibility software for small marketing teams”;
  • “AI search monitoring platforms for marketing teams.”

They may belong to one commercial cluster if the page can compare tools by team fit, engines, citations, workflow, and cost.

Now compare:

  • “how to get cited in Perplexity answers”;
  • “best AI visibility tools for SaaS.”

They share entities but not jobs. One asks for a method; the other asks for a product decision. Combining them produces a page that satisfies neither cleanly.

Cluster fields

For each cluster, store:

  • cluster name;
  • buyer job;
  • buying stage;
  • representative keywords;
  • representative prompts;
  • required constraints;
  • dominant search intent;
  • likely page format;
  • existing canonical URL;
  • source of demand;
  • search metrics;
  • prompt and citation metrics;
  • evidence available;
  • commercial relevance;
  • next action.

Step 4: map keywords to prompts

The map reveals four useful states.

Search demandPrompt demandInterpretationTypical action
HighHighEstablished category or problem with conversational buying activityBuild or strengthen a pillar/commercial page
HighLow or unknownSearch demand exists, but your prompt set may not represent the journeyValidate with first-party research before scaling
LowHighA narrow or emerging decision is visible in conversationsCreate targeted documentation, comparison, or workflow content if commercially relevant
LowLowWeak evidenceDeprioritize unless strategically necessary

Example keyword-to-prompt map

Keyword clusterSearch intentBuyer promptPrompt intentPage decision
AI SEOUnderstand category“How do I improve AI search visibility?”Operational overviewAI SEO pillar
SEO content writing toolsCompare products“Which content tool handles research, source verification, and CMS handoff?”Constrained purchaseTool comparison
Keyword researchLearn method“How do I find buyer prompts keyword tools miss?”Research workflowThis guide
Content refresh SEOUpdate a page“Why did an article lose rankings and AI citations?”Diagnosis and executionRefresh playbook
AI visibility contentComplete workflow“How do I turn a missing citation into content?”End-to-end SOPWorkflow guide

The cluster is useful because each page has a separate job. The keyword and prompt axes converge into a coherent internal-link system rather than five variations of “AI SEO.”

Worked cluster: AI visibility software for small teams

Suppose a first-party prompt monitor repeatedly evaluates “AI visibility software for small marketing teams.” Treat that as a starting point, not proof that it deserves a page.

First, identify the decision inside the prompt. “Small marketing teams” implies constraints such as setup time, price, reporting overhead, engine coverage, and whether one person can operate the monitor-to-publish workflow. Expand only into variants that change the buying decision: for example, a startup founder, an agency managing several clients, or a team that needs direct publishing to its existing CMS. A cosmetic rewrite like “best software for a small marketing department” adds little.

Next, map those decisions to search evidence. Search Console may reveal impressions for product-category and comparison terms even when it never shows the conversational wording. Trends and third-party tools can indicate whether the category is growing, but they do not validate every modifier. Sales calls, onboarding notes, and support questions can confirm whether the constraints are real.

Then inspect the generated answers. Record which products are mentioned, which URLs are cited, and the role each citation plays. If official product pages establish features while independent comparisons drive the shortlist, your content and distribution plan need both roles. If engines use materially different source sets, do not treat one answer as the universal ranking.

Finally, choose the smallest page that resolves the decision. A narrowly evidenced comparison for small teams may be justified. If the evidence is thin, enrich an existing comparison with a clearly labeled small-team section instead. The goal is not one page per prompt; it is one authoritative answer per distinct buyer decision.

Step 5: analyze cited sources

Prompt monitoring becomes research when you inspect the sources used in answers.

For each prompt cluster, aggregate:

  • cited domains;
  • cited URLs;
  • citation frequency;
  • citation position;
  • page type;
  • source recency;
  • brand mentioned in the answer;
  • brand associated with the cited page;
  • whether the citation supports the generated claim.

Look for source roles

Sources often play different roles:

  • definition: explains the category;
  • evidence: supplies a statistic or test result;
  • comparison: evaluates vendors;
  • product fact: verifies a feature or price;
  • experience: provides reviews or discussion;
  • procedure: supplies steps;
  • authority: corroborates a claim.

The action depends on the role. If independent comparisons drive recommendations, an owned article may not substitute for third-party corroboration. If the answer cites an outdated product page, updated documentation may work better than outreach.

Grade source quality separately from frequency

Frequent citation does not automatically make a source reliable. Review provenance, fit to the claim, publication or verification date, disclosed methodology, commercial incentives, accessibility, URL stability, and corroboration by independent sources. A vendor page can be the best source for its current feature set and a poor source for an industry-wide performance claim. A survey can look authoritative while hiding an unrepresentative sample.

Store this assessment next to citation frequency. It helps distinguish a content gap from a source-quality problem and prevents the team from copying whatever happens to be cited most often.

Measure concentration

If one domain appears repeatedly, investigate it. If each engine uses a different source set, design for source diversity rather than chasing a single publisher.

In BeVisible's recent sample for “AI visibility software for small marketing teams,” source sets varied materially across ChatGPT, AI Overviews, Gemini, and AI Mode. That is a warning against optimizing for one manually selected answer.

Step 6: score commercial fit and visibility gaps

Use a score that preserves its components.

Suggested 100-point model

ComponentWeightScoring question
ICP relevance20Does this cluster belong to a problem the product genuinely solves?
Buying-stage value15Is the reader approaching a decision or meaningful next step?
Search demand15Do first- and third-party search signals justify attention?
Prompt evidence15Does the cluster recur in observed or monitored buyer language?
Visibility gap15Are competitors or other sources repeatedly present while you are absent?
Evidence advantage10Can you contribute proof, expertise, data, or a useful tool?
Authority feasibility5Can the site credibly compete or earn distribution?
Production fit5Can the team maintain the page's facts and scope?

Score 0–5 for each component, multiply by weight, and retain the notes behind the score.

Apply penalties after the base score

  • minus 20 for strong cannibalization risk;
  • minus 15 when required evidence is unavailable;
  • minus 15 when the only rationale is estimated traffic;
  • minus 10 when volatile facts cannot be maintained;
  • minus 10 when the page would sit outside the site's expertise;
  • minus 5–20 for legal, privacy, or brand risk.

The penalty model stops a high-volume phrase from overpowering an obvious quality problem.

Step 7: choose create, refresh, consolidate, outreach, or no action

Keyword research should end in a decision, not a spreadsheet of phrases.

Decision tree routing an observed content opportunity to refresh, consolidate, create, outreach, technical repair, or no action

Create

Use when no page satisfies the buyer job and the team has sufficient evidence.

Refresh

Use when the canonical page exists but the query mix, facts, format, or citations have changed. Follow the content refresh for AI search playbook.

Consolidate

Use when multiple URLs split one intent. Merge the strongest unique material, redirect obsolete URLs, and repair internal links.

Outreach or distribution

Use when third-party sources shape the answer and your product needs accurate inclusion or corroboration.

Technical fix

Use when the page cannot be crawled, indexed, rendered, or attributed correctly.

No action

Use when demand is irrelevant, evidence is weak, the maintenance burden is too high, or another page already handles the job.

“No action” is a successful research outcome. It protects the domain from publishing work it cannot defend.

Prioritization worksheet

Copy this schema into a spreadsheet or database:

FieldExample
ClusterAI visibility software for small teams
Buyer jobSelect a practical monitoring platform
StageCommercial evaluation
KeywordsAI visibility software, AI search monitoring tools
PromptsAI visibility software for small marketing teams
Existing URLNone or canonical comparison URL
Search demand bandMedium
TrendGrowing / stable / declining / unknown
Prompt sourceMonitored portfolio + sales calls
Engines observedChatGPT, Perplexity, Gemini, AIO, AI Mode
Recurring competitorsNamed vendors
Recurring sourcesURLs/domains from response evidence
Own visibilityBaseline rate and period
Evidence advantageFirst-party prompt/citation data and workflow test
ActionCreate comparison
ScoreComponent total, not only final number
RisksProduct facts change; comparison maintenance
OwnerProduct marketing
Review cadenceMonthly facts, quarterly full review

Three practical examples

Example 1: high keyword volume, weak product fit

“AI writing generator” has measurable demand, but a visibility-monitoring product does not need a generic generator article unless it can offer a distinct workflow or tool. The score should fall on ICP relevance and evidence advantage even if volume is attractive.

Decision: no action or a tightly scoped supporting mention, not a pillar.

Example 2: low measured volume, strong prompt value

“AI visibility tools with direct CMS publishing” may have limited conventional volume, but it expresses a real workflow constraint near purchase. If repeated in prompts or sales conversations, it may justify a comparison or capability page.

Decision: targeted commercial page or product documentation.

Example 3: search and prompt demand overlap

“Content refresh SEO” has a clear search-learning intent. A conversational version asks how to recover both rankings and AI citations. The shared job is updating a decaying page, but the AI layer adds source and mention measurement.

Decision: one complete refresh playbook, not separate “SEO refresh” and “AI refresh” posts that mostly overlap.

Research mistakes that create bad content

Treating every prompt as a keyword

Long prompts contain context that should shape evidence and format. Collapsing them into a two-word phrase removes the useful part.

Treating generated prompt variants as demand

An LLM can create 1,000 plausible questions. That proves linguistic possibility, not buyer frequency. Label synthetic expansion and validate it against observed sources.

Combining branded and organic performance

Branded prompts measure representation and verification. Organic prompts measure discovery. Mixing them inflates performance and obscures the opportunity.

Using search volume as precision

Volume is modeled, rounded, segmented, and tool-dependent. Use bands and decision thresholds, not false exactness.

Ignoring Search Console limitations

Query tables are not complete, anonymized rows are omitted, and property/page aggregation changes totals. Document exports and filters.

Copying the cited pages

Cited sources show what the system used, not a template you should duplicate. Your page needs better evidence, clearer scope, a more useful decision, or a different source role.

Publishing clusters without canonical ownership

If two planned pages answer the same buyer job, merge them before drafting. The cheapest cannibalization fix happens in the research sheet.

A monthly research cadence

Week 1: collect

  • export Search Console query/page data;
  • update keyword and trend signals;
  • ingest first-party buyer questions;
  • run the stable monitored prompt set;
  • store answers and citations.

Week 2: cluster and inspect

  • group by buyer job;
  • map keywords to prompts;
  • inspect recurring cited sources;
  • find existing canonical pages;
  • flag facts or intents that changed.

Week 3: prioritize

  • score commercial fit, demand, gaps, and evidence;
  • apply overlap and maintenance penalties;
  • choose create, refresh, consolidate, outreach, technical, or no action;
  • assign owners.

Week 4: brief and measure

  • create evidence-backed briefs;
  • capture baselines;
  • review the previous month's published actions;
  • adjust the portfolio, not the scoring rules, based on what you learned.

The output that matters

The end product of keyword research for AI search is not a larger keyword list. It is a smaller, better-defended production queue.

Each item should tell you:

  • which buyer job matters;
  • how search and prompt evidence support it;
  • which sources currently shape the answer;
  • what canonical page should exist;
  • what evidence will make the page distinct;
  • whether the action is content at all;
  • how success will be measured.

Keep keyword demand and prompt demand separate long enough to see their disagreement. Then combine them at the page-decision layer.

That is how you find topics that can be discovered in search, useful in an AI answer, credible for your domain, and valuable to the business at the same time.

Last reviewed: July 21, 2026.