U.S. Growth

How to get your brand cited by ChatGPT, Gemini and Perplexity

A practical AEO and GEO checklist: the structured data, content shape and authority signals that get a brand into AI answers, not just onto page one of Google.

Ignite Consulting · Updated Apr 12, 2026 · 28 min read

A customer no longer types a keyword and scans ten blue links. They ask a full question, and a model answers in a paragraph with two or three sources hanging off the side. If your brand is one of those sources, you are in the consideration set before the customer has clicked anything. If you are not, you are invisible at the exact moment the decision is forming. Getting cited is the new ranking, and it does not follow the same rules.

The uncomfortable part: ranking well on Google is necessary but nowhere near sufficient. SeoClarity's analysis of 432,000 keywords found that 97% of AI Overviews cite at least one source from the top 20 organic results, not the top 10, which tells you the eligibility pool is wider and the selection logic is different (SeoClarity). At the same time, Averi's review of 680 million AI citations found only an 11% domain overlap between ChatGPT and Perplexity (Averi). A single page rarely wins every engine. You earn citations the way you earn trust: by being verifiable, specific and structured enough that a machine can lift you cleanly into an answer.

This guide is long on purpose. Getting cited by AI engines is not a five-minute checklist, and the people selling it as one are either confused or hoping you are. What follows is the full picture: how the engines actually choose sources, how to shape content and structured data so a model can quote you without guessing, how to build the off-site authority that makes a claim read as fact rather than marketing, a step-by-step playbook you can run, the mistakes that quietly sink most efforts, the metrics worth watching, and a long FAQ. Read it once to understand the terrain, then keep it open as a working reference.

One framing helps before we start. For most of the last two decades, the unit of visibility was a position: where you sat in a ranked list. The customer did the synthesis. They scanned the list, opened a few tabs, compared, and decided. The new unit of visibility is a sentence inside an answer the model already synthesized. The customer reads a paragraph, sees two or three sources, and forms an opinion before clicking anything. That shift moves the work earlier in the funnel and raises the bar on what counts as good content. A page can be number one on Google and never appear in the answer that sits above the links. A page can rank on the second page of Google and be the sentence a model quotes. The list and the answer are now two different games played on the same web, and this guide is about the second one.

First, understand what "getting cited" actually means

There is a persistent myth that you can somehow pay, prompt, or trick a model into naming your brand. You cannot. When ChatGPT, Gemini, or Perplexity produces an answer, it is doing one of two things, and usually both at once. It is recalling your brand as an entity it learned about during training, and it is retrieving fresh pages in real time and reading them before it writes. A citation is the visible tip of those two processes meeting: the model decided your brand belongs in the answer, and it found a source it trusts enough to point at.

That distinction matters because it splits your work into two jobs that look different and are funded differently. The first job is retrieval readiness: making sure the right bots can reach your pages, parse them, and lift a clean answer. The second job is entity authority: making your brand a well-attested, consistently-described entity that the model already associates with the topic before anyone hits enter. Most brands obsess over the first and neglect the second. The brands that win do both, and they treat the second as the harder, slower, more valuable one.

Retrieval versus memory, and why both matter

Retrieval is the live layer. Perplexity in particular runs a search and reads results on almost every query, so a page published this morning can appear in an answer this afternoon. ChatGPT cites fewer sources per answer but pulls each one in more deeply, weaving it into the prose rather than listing it (AuthorityTech). Memory is the trained layer: what the model absorbed about your category and the players in it during pre-training. You influence memory slowly, through years of consistent third-party description, and you influence retrieval quickly, through crawlable, quotable, current pages. A real program works both clocks at once.

Citation is a trust decision, not a popularity contest

Engines do not cite the loudest brand. They cite the source that lets them answer confidently and defend the answer. That is why a calm, specific, well-sourced page beats a glossy page full of adjectives, and why a single number with context ("cut onboarding from 14 days to 3") beats a paragraph of "industry-leading" filler. You are not optimizing for clicks here. You are optimizing for a machine's willingness to repeat you and attach its credibility to yours. Everything below is downstream of that one idea.

Know what each engine actually pulls from

The engines do not share a brain, and they do not share a source pool. Semrush's study of 30 million cited sources found Reddit was the single most-cited domain across ChatGPT, Gemini, Perplexity, Google AI Mode and AI Overviews, with YouTube, LinkedIn, Wikipedia and Forbes filling out the top tier (Semrush). Underneath that headline, the platforms diverge hard. Wikipedia carries an outsized share of ChatGPT's sourcing, while Reddit dominates Perplexity's (PR Newswire).

That has direct consequences for where you invest. If you sell into a category where customers ask Perplexity, your unpaid mentions on Reddit, LinkedIn and authoritative third-party sites matter as much as your own pages. If your customers lean on ChatGPT, a clean, well-sourced Wikipedia presence and citations on reference-grade sites do more lifting. The work is not to game one engine. It is to make your brand the obvious, repeated, well-attested answer across the surfaces each engine trusts.

Illustrative engine profiles (typical tendencies, not fixed rules)
EngineRetrieval styleLeans heavily onWhat wins citations
ChatGPT SearchSelective, fewer sources, deeper integrationWikipedia, reference-grade sites, established mediaEntity clarity, authoritative third-party corroboration
PerplexityLive retrieval on most queries, many sourcesReddit, YouTube, LinkedIn, fresh pagesFreshness, real community discussion, clear answers
Google AI Overviews / AI ModeTied to its index, top-20 eligibilityStrong organic pages, structured dataSchema, depth, matching the query intent precisely
GeminiGoogle ecosystem, Google-Extended corpusIndexed pages, YouTube, Google propertiesCrawl access via Google-Extended, clear structure

Treat that table as a map of tendencies, not a rulebook. The shares shift month to month and category to category, and any single number you read in a study is a snapshot. The durable takeaway is the shape: different engines weight different surfaces, so a brand that lives only on its own domain is betting everything on a narrow slice of how AI answers get built. The fix is not more pages. It is presence across the surfaces these engines actually read.

Why category matters more than the averages

Aggregate citation studies are useful for orientation and dangerous for strategy. A B2B software buyer researching a compliance tool behaves nothing like a consumer comparing running shoes. The compliance buyer's answers may lean on analyst write-ups, documentation, and a handful of trade publications. The shoe shopper's answers may lean on Reddit threads, YouTube reviews, and retailer pages. Before you decide where to invest, run your own customer questions through the engines and watch which domains actually show up for you. That ground truth beats any industry average, and it is the first step of the playbook later in this guide.

The two jobs in detail: entity authority and retrieval readiness

Earlier we split the work into entity authority and retrieval readiness. It is worth spending real time on the distinction, because almost every wasted GEO budget comes from confusing the two or doing only one. They run on different clocks, cost different things, and fail in different ways. A program that treats them as a single undifferentiated "AI optimization" line item tends to overspend on the fast, visible job and underspend on the slow, valuable one.

Entity authority: becoming a thing the model already knows

An entity is the model's internal idea of who you are: a named company, in a named category, with a set of attributes it has seen described consistently across many independent sources. When a model has a strong entity for you, it can name you in an answer without retrieving a single page, because it already "knows" you belong in that category. When it has a weak or absent entity, no amount of on-page polish will conjure you into an answer about your space, because the model has nothing to recall. Entity authority is built slowly, through years of consistent third-party description, and it is the closest thing to a moat in this discipline. It is also the part most teams skip, because it cannot be shipped from the content management system and it does not produce a screenshot you can show a manager next week.

The practical signals that build an entity are mundane and cumulative. A consistent legal name and brand name everywhere you appear. The same category description on your site, your profiles, your press, and the directories that feed knowledge graphs. Independent coverage that repeats the same facts about you. A presence on the reference-grade surfaces the models lean on. None of these is a trick. Each is a small vote that, repeated enough times by enough independent sources, becomes the model's settled understanding of you. The brands that get cited as authorities are the ones that became authorities first, on the open web, and then let the engines notice.

Retrieval readiness: being liftable right now

Retrieval readiness is the fast job. It is everything that determines whether, when a model runs a live search to answer a question this minute, it can reach your page, read it, and lift a clean answer. Crawl access, server-rendered content, answer-first structure, accurate schema, fresh dates: all of it serves the single goal of being easy to quote on demand. This job rewards speed and precision, and it pays off in days to weeks on retrieval-heavy engines. It is satisfying because the wins are visible. It is also a trap if you treat it as the whole job, because a perfectly liftable page about a brand the model has no entity for will still lose to a competitor the model already recognizes.

The right mental model is a relay. Entity authority gets you considered. Retrieval readiness gets you quoted. A strong entity with poor retrieval readiness shows up described from memory, sometimes inaccurately, and rarely with a link to you. Strong retrieval readiness with a weak entity gets crawled and ignored. You need both batons to finish the race, and you fund them differently: retrieval readiness as a sprint you can mostly complete in a quarter, entity authority as a standing program that compounds for years.

Make the machine reachable: crawl access and rendering

Many brands lose at the first gate without realizing it. If your robots.txt, a CDN firewall, or an over-eager bot-management rule blocks the AI crawlers, nothing downstream matters. The page can be brilliant and perfectly structured, and the engine will never see it. AI bot traffic climbed more than 300% between early 2025 and early 2026, which means crawl-budget and access decisions you made for old-school SEO may quietly be excluding the agents that now decide your AI visibility (Cubitrek).

The user-agents to allow on purpose

There is a meaningful difference between training crawlers and live-retrieval agents. Training crawlers gather data to improve future models; you can reasonably choose to allow or block them based on your stance on training. Live-retrieval agents are the ones fetching pages to answer a question right now, and blocking those is self-defeating if you want to be cited. At minimum, decide deliberately about these:

  • OpenAI: GPTBot (training), OAI-SearchBot (ChatGPT search results), ChatGPT-User (live user-triggered fetches). Allow the search and user agents even if you debate the training one.
  • Perplexity: PerplexityBot, and its user-triggered fetcher. Perplexity's freshness bias makes this access especially valuable.
  • Google: Google-Extended, which governs whether your content feeds Gemini and AI Overviews. Blocking it does not improve normal Search ranking; it just removes you from the AI surface.
  • Anthropic and others: their search and user agents follow the same logic. The category of bot matters more than the brand.

The rendering trap most teams miss

Allowing a bot in is not enough if the bot cannot read what loads. A common and costly mistake: the page renders its real content through client-side JavaScript, and a retrieval crawler that does not execute JS sees an empty shell. Your beautifully written answer is technically on the page, but invisible to the machine that would have cited it. Put the core answer, the headings, and the key facts in server-rendered HTML so the content exists before any script runs. Test this the lazy way by viewing the raw page source, or fetching the URL with JavaScript disabled, and asking whether the answer is still there in plain text.

Freshness as an access issue

Live-retrieval engines lean toward recent material, and stale pages quietly lose citation share over time. Treat your highest-value pages as living documents: revisit them on a schedule, update the data and the dates honestly, and resubmit or simply keep them crawlable so the engines re-fetch. This is not about churning out new posts. It is about keeping your best, most-cited pages current so they stay eligible.

Shape content the way a model reads it

Models do not reward clever prose. They reward passages they can extract and quote without ambiguity. The pages that get cited tend to answer one question per section, lead with the answer, then support it. The most useful single move is to write a direct, self-contained answer in the first 40 to 60 words under a heading that matches how a person would phrase the question out loud.

Answer-first structure, with the question as the heading

Write your H2s and H3s as the questions a customer actually asks, not as internal marketing labels. "How does an overseas factory get found in ChatGPT by US buyers" is a heading a model can match to a query. "Our advantages" is not. Under each heading, open with a conclusion that stands on its own without the surrounding paragraphs, then give the reasoning, then an example. The inverted pyramid that journalists use is also the shape an engine prefers, because it can lift the top of the pyramid and trust it in isolation.

Name entities explicitly and avoid pronouns

Models build associations from named entities. Write the brand name, the product name, the place, and the person in full rather than leaning on "we," "our solution," or "it." Every time you replace a pronoun with the actual entity name, you make the relationship between your brand and the topic legible to the machine. This feels repetitive to a human reader and reads as clarity to a model, so strike a balance: be explicit at the start of sections and at the points you most want quoted.

The content types that earn citations

A few patterns earn citations far more than others. Early GEO research from a Princeton-led team found that adding cited sources and statistics each lifted AI visibility meaningfully, and that inserting quotations from named experts lifted it further still, because models treat quotes and attributions as proxies for credibility (GEO research summary). In practice, prioritize:

  • Original data and proprietary research. Across all engines, first-party numbers, surveys and benchmarks are the highest-leverage content type, because there is nowhere else to pull the figure from. You become the primary source, and primary sources get cited by name.
  • Bottom-funnel pages, not generic explainers. Pricing pages, comparison pages and real case studies outperform thin "what is" and "how to" filler for driving AI-referred traffic. Those are also the queries closest to a purchase.
  • Crisp definitions and direct answers. A two-sentence definition at the top of a section is far more quotable than the same point buried in paragraph six.
  • Plain numbers with context. "Cut onboarding time from 14 days to 3" is liftable. "Dramatically faster onboarding" is not.
  • Named-expert quotes. A short, attributed quote from a real person with a real title gives the model a credibility signal it can carry into the answer.

Reddit's dominance also tells you something about format: models trust content that reads like a real person solving a real problem. A genuinely useful answer in a community thread, written by someone who actually knows the space, can outrank a polished corporate page in an AI answer. That is earned presence, not a posting trick, and the moment it becomes a posting trick the platforms and the models both start discounting it.

Liftable versus unliftable phrasing (same claim, different odds of being quoted)
Unliftable (vague)Liftable (specific)Why the model prefers it
"Our onboarding is dramatically faster.""We cut median onboarding from 14 days to 3."A concrete number with a baseline can be quoted and defended.
"We serve many leading brands.""We support 40+ DTC brands shipping to the US and EU."A bounded, checkable claim reads as fact, not puffery.
"GEO is important for visibility.""GEO is the practice of earning citations in AI-generated answers."A clean definition is the most quotable unit on the page.
"Pricing is competitive and flexible.""Plans start at a fixed monthly retainer with no long-term lock-in."Specific structure answers the real customer question directly.

Write for the snippet, then for the page

A useful discipline when drafting any page you want cited: write the quotable sentence first, in isolation, before you write the section around it. Imagine the model has room for exactly one sentence from your page in its answer. What would you want that sentence to be? Write it so it stands completely on its own, names the entity, contains a concrete fact, and answers the question cleanly. Then build the section to support and earn that sentence. This inverts the usual writing habit, where the strong line emerges somewhere in the middle after a warm-up. Models do not read warm-ups generously. They scan for the liftable unit, and if the best one is buried in paragraph four behind throat-clearing, a competitor who led with theirs gets the citation instead.

Match the question's phrasing, not just its topic

There is a difference between a page about a topic and a page that answers a question. A page titled "Cold Email Deliverability" is about a topic. A heading that reads "Why do my cold emails land in spam, and how do I fix it" is an answer the model can map to a real query. The closer your headings track the literal phrasing a person would type or speak, the easier the match. This is why turning your H2s and H3s into questions, covered earlier, is not a cosmetic choice. It is the hinge between being topically relevant and being the specific answer the engine reaches for. When you build your question set in the playbook below, harvest the exact wording from the engines themselves: the "people also ask" style follow-ups, the way Perplexity rephrases, the related questions ChatGPT volunteers. Those are the phrasings to write headings against.

A short worked example of reshaping a page

Take a generic services page that opens, "We are a full-service growth partner committed to delivering best-in-class results for ambitious brands." A model reads that and has nothing to lift. There is no entity relationship, no fact, no answer. Now reshape it. Lead with a heading phrased as the customer's question, "What does a US growth consultancy actually do for an overseas brand," then open with a self-contained answer: "A US growth consultancy helps an overseas brand become visible and credible to American customers across search, AI answers, earned media, and paid channels, then hands the brand a measurable pipeline to run." Follow with a concrete, bounded claim and a named example. The reshaped version is liftable in three places, names the entity and category explicitly, and reads as fact rather than mood. Nothing about it is dishonest or inflated. It is simply written so a machine can repeat it without guessing.

Make the machine-readable layer airtight

Structured data is how you tell an engine what a page is, who wrote it and why they are credible, without making it guess. Several 2026 analyses link schema to higher selection rates, with FAQ schema associated with materially higher citation weighting in ChatGPT and structured data tied to a meaningful lift in AI Overview selection (Sapt). Treat those figures as directional rather than guaranteed, but the principle holds: schema removes ambiguity, and ambiguity is what gets a page passed over.

The schema types that earn their keep for GEO

You do not need every schema type in the catalog. For AI visibility, a focused set does most of the work, applied accurately rather than stuffed:

  • Organization schema that binds your brand entity together: legal name, logo, and a sameAs array linking your verified LinkedIn, Wikipedia, Wikidata, and major profiles. This is how you tell the machine "all of these refer to the same company."
  • Article schema with author and publisher fields that map to a real, named expert, plus accurate datePublished and dateModified values that match what is on the page.
  • FAQPage schema that explicitly pairs questions with answers. This is the single highest-leverage type for AEO because it hands the model a ready-made question-and-answer unit.
  • Product schema for commerce pages, and BreadcrumbList for site structure, both of which help the engine understand where a page sits and what it represents.

E-E-A-T signals the engine can actually verify

Schema tells the machine what to think; visible signals let it check. Engines increasingly look at who is talking before they repeat it, so make the experience, expertise, authority, and trust legible on the page itself:

  • Author bios with real credentials, linked to a real profile, not a generic "admin" byline.
  • Clear dates and citations to primary sources, so a claim can be traced rather than taken on faith.
  • A genuine "About" footprint: a real company page, consistent contact details, and an entity description that matches what third parties say about you.
  • Consistency across the web, because the engine cross-checks. If your own site says one thing and your LinkedIn says another, the disagreement itself is a trust penalty.

On llms.txt: ship it, but do not believe the hype

Be honest about llms.txt. Roughly one in ten domains has adopted it, but as of early 2026 no major AI company has committed to reading it in production (Codersera). It is cheap to ship and useful for AI-assisted developer tools, so add it, but do not mistake it for a citation strategy. The page itself, its crawlability, and its off-site corroboration decide whether you get cited, not a courtesy file most engines do not yet request.

Key takeaways

  • Top-10 Google ranking is not the bar. The eligibility pool reaches the top 20, and each engine draws from a different source set.
  • Win surfaces, not a single page: your own content, plus earned presence on Reddit, LinkedIn, Wikipedia and authoritative third-party sites.
  • Lead every section with a direct, self-contained answer and a real number a model can lift verbatim.
  • Publish original data. First-party research is the one thing an engine cannot source anywhere else.
  • Get the machine-readable layer right: accurate schema, named expert authors, and crawl access for live-retrieval AI agents.

Build authority off your own site

This is where most brands stall, because it cannot be done from the CMS. Engines weigh corroboration: a claim repeated across independent, credible sources reads as fact, while a claim that lives only on your homepage reads as marketing. The single highest-return play here is earned media. A feature in a publication the model already trusts, a mention in an industry roundup, a credible review on a site like G2, each one adds an outside vote that the engine can cite instead of, or alongside, your own page.

Digital PR is the most reliable lever

There is a measurable gap between traditional rankings and AI visibility, and PR is the most reliable way to close it. Run digital PR with the explicit goal of getting your brand named in the third-party sources these engines lean on, rather than chasing links alone. When Wikipedia, Forbes, Reddit and a respected trade outlet all describe your company the same way, you have given every engine a consistent story to repeat. That consistency is the asset. A scattered set of mentions that describe you five different ways is far weaker than three mentions that all say the same clear thing. For the full argument on why earned coverage is the moat here, see our piece on digital PR as the GEO moat.

Reference-grade presence: Wikipedia and structured facts

For ChatGPT in particular, reference-grade sources do heavy lifting. If your brand genuinely meets notability standards, a neutral, well-sourced Wikipedia entry is one of the highest-impact single investments you can make, because it feeds the trained memory of multiple models. You cannot manufacture notability, and you should never pay someone to spin up a promotional entry that will be deleted. What you can do is earn the independent coverage that makes a legitimate entry possible, and ensure structured facts about your company on directories, Wikidata, and knowledge-graph sources are consistent and correct.

Community presence, earned the slow way

For Perplexity and any engine that leans on Reddit, genuine community presence carries weight. The only durable way to earn it is to actually participate: answer real questions in your space, helpfully, as a knowledgeable human, over time. Astroturfing and review manipulation get detected, violate platform rules, and backfire when an engine learns to discount the pattern. The brands that show up well in community-sourced answers are the ones whose customers and employees are genuinely active and genuinely useful.

Consistency is the hidden multiplier

Of all the off-site levers, consistency is the one teams underrate most. A model cross-checks. When your homepage, your LinkedIn, a trade article, and a directory all describe you the same way, in the same category, with the same claims, the agreement itself becomes a trust signal: independent sources corroborating each other is exactly what a citation engine is built to reward. When they disagree, the disagreement is a penalty, because the model cannot tell which version to repeat and tends to hedge or pick a competitor whose story is unambiguous. This means the highest-leverage cleanup is often not creating new content at all. It is auditing every surface that describes your company and making them say the same true thing. Pick one category description, one set of headline facts, one boilerplate, and propagate it everywhere a machine reads. The work is unglamorous and it moves description accuracy faster than almost anything else.

The build-versus-buy question for off-site authority

Off-site authority is the part clients most often want to shortcut, and the part where shortcuts most reliably fail. You cannot buy your way into a model's trained memory, and you cannot manufacture notability. What you can do is earn coverage on the surfaces your baseline shows actually feed the engines in your category, then make sure the facts those sources carry are consistent with everything else about you. For overseas brands in particular, this is the slow, unavoidable center of the work, and it overlaps heavily with the middleman and distribution decisions covered in overseas middleman traps: a brand that has outsourced its entire US presence to a third party often finds the model has no coherent entity for the brand itself, only for the intermediary.

How to monitor what AI actually says about your brand

You cannot improve what you do not measure, and AI visibility does not show up cleanly in the analytics dashboards built for classic search. The traffic an AI answer sends you is smaller and harder to attribute than organic clicks, and the most important outcome, being named accurately in an answer the customer never clicks through from, leaves almost no trace in your logs at all. So monitoring AI visibility starts not with a dashboard but with a disciplined habit: asking the engines your customer questions, on a schedule, and recording what they say.

The manual baseline anyone can run

The cheapest, most honest measurement is the one you can do yourself with no tooling. Take your customer question set, ask each question in ChatGPT, Gemini, and Perplexity, and record three things per answer: whether you are mentioned, whether the answer links to you, and whether the description is accurate. Note which competitors appear and which domains the engines cite. Do this once to set a baseline, then again every month, and the trend tells you whether the work is moving anything. This is unglamorous and it is the single most reliable signal you have, because it measures the actual outcome rather than a proxy. Tools can scale it, but they cannot replace the judgment of reading what the model says about you in plain language.

What the logs tell you that the answers do not

Your server logs hold a leading indicator the answers cannot give you: how often the AI crawlers fetch your pages. A rising count of GPTBot, OAI-SearchBot, PerplexityBot, and Google-Extended hits on a page is evidence the retrieval layer is reaching it, which is a precondition for any citation. If a page you care about is never being fetched by these agents, no amount of content work will get it cited, and the fix is access, not copy. Watching crawler behavior also catches regressions early: a CDN rule change or a robots.txt edit that quietly blocks an agent shows up as a sudden drop in fetches long before it shows up as lost citations.

Where dedicated tooling helps, and where it overpromises

A category of AI-visibility tools now scales the manual baseline: running large question sets across engines, tracking citation share over time, and flagging changes. They are genuinely useful for breadth and for catching drift you would miss by hand. Treat their numbers as directional, though, not absolute. Answers vary by session, by phrasing, by region, and by the model version live that day, so any single citation-share figure is a snapshot, not a constant. Use the tools to watch the trend and to widen coverage, and keep doing the manual reads for the questions that matter most, because reading the actual answer in context is the only way to catch a description that is technically a mention but subtly wrong or unflattering.

A step-by-step GEO playbook you can run

Here is the sequence we use, in order, because order matters. Do not skip to content production before you have fixed access and mapped the landscape; you will optimize pages no engine can read or for queries no customer asks.

  1. Build your question set. List the 20 to 50 questions your real customers ask at the point of decision: comparisons, "best X for Y," pricing, alternatives, and category definitions. These are your targets and your scoreboard.
  2. Run the baseline. Ask each question in ChatGPT, Gemini, and Perplexity. Record three things per answer: does it mention you, does it link to you, and is the description accurate. Note which competitors and which domains get cited instead.
  3. Fix crawl access. Audit robots.txt, CDN, and bot-management rules. Allow the live-retrieval agents on purpose. Confirm your core content is in server-rendered HTML, not hidden behind JavaScript.
  4. Tighten the machine-readable layer. Add or correct Organization, Article, FAQPage, and Product schema. Make author and publisher fields point to real, named experts. Get dates honest and consistent.
  5. Reshape your highest-value pages. Rewrite headings as customer questions, add answer-first openers, name entities explicitly, and insert real numbers, definitions, and attributed quotes.
  6. Publish something only you can publish. A first-party survey, a benchmark, a teardown, a price-transparency breakdown. Make yourself the primary source for one figure your category cares about.
  7. Close the off-site gap. Run digital PR aimed at the third-party surfaces the baseline showed your competitors winning. Pursue legitimate reference-grade presence and consistent structured facts.
  8. Re-run the question set monthly. Compare to the baseline. Track citation share, accuracy of description, and which new domains appear. Double down on what moved.

A 90-day cadence that actually fits a real team

In the first month, do the unglamorous foundation work: question set, baseline, crawl access, and schema. None of it produces a visible win, and all of it is load-bearing. In the second month, reshape the pages and ship the one piece of original data; this is when you start to see movement on retrieval-heavy engines like Perplexity. In the third month, push the off-site authority work, which is slower and compounds later, often showing up first as more accurate descriptions and then as more citations. Anyone promising visible citation gains in week one is selling the calendar, not the work.

Two short scenarios, start to citation

Scenario one: a US local service business

A regional HVAC company wants to show up when someone asks an AI assistant "who is the best emergency furnace repair near me" in their metro. The baseline shows the engines cite a couple of directories and one competitor, and never name the company. The fix is layered: clean, consistent business facts everywhere the engine cross-checks (their Google Business Profile, directories, and their own site all stating the same service area and hours), genuine reviews that describe specific outcomes, an FAQ page answering the exact emergency questions people ask, and a few local-press mentions. Three months in, the assistant starts naming them alongside the directory, because the engine now has a consistent, corroborated story about who serves that area. The mechanics of this overlap heavily with classic local search, which we cover in US local SEO and Google Business.

Scenario two: a China-based brand entering the US

A manufacturer with a strong product and almost no US footprint wants US buyers to find it through AI. The baseline is brutal: the engines do not know the brand exists, and answers about its category name only US and European incumbents. The work here is mostly entity-building, not page-tweaking. The brand needs a coherent English-language presence, consistent facts across every surface an engine reads, legitimate third-party coverage in US trade media, and reference-grade structured facts. It is slow, and it is the right slow. A brand that is genuinely unknown in a market cannot shortcut its way into being cited as an authority there; it has to become one. For the broader build-versus-marketplace question this raises, see building a brand versus selling on marketplaces, and for the export operating model, the China B2B export playbook.

Scenario three: a B2B software vendor with good rankings and no citations

A mid-market SaaS company ranks on page one for most of its category terms and still almost never appears when buyers ask the engines "best tool for X" or "X alternatives." The baseline shows the engines citing a review site, a competitor's comparison page, and a Reddit thread, none of which mention this vendor. Diagnosis: strong classic SEO, weak retrieval readiness and thin off-site corroboration. The fix layers cleanly onto the playbook. They reshape their comparison and pricing pages to answer the literal questions buyers ask, publish a first-party benchmark only they can produce, get the schema and authorship right, and run digital PR aimed at the exact review and trade surfaces the baseline showed winning. The AI-visibility work then feeds the top of a pipeline they run themselves with verified contact data, where the consultancy delivers the list and the client runs all outreach. That model, and the boundary around it, is laid out in the B2B growth engine.

Budgeting the work: effort, payoff, and patience

Because the two jobs run on different clocks, the smartest budgeting question is not "how much should I spend on GEO" but "how should I sequence effort against payoff." Front-load the cheap, fast retrieval-readiness work, because it is mostly one-time and it unblocks everything else. Treat off-site authority as a standing commitment that compounds, not a campaign that ends. The table below frames typical effort and payoff timing so you can set expectations honestly with whoever holds the budget. Treat the timing as illustrative ranges, not promises, since they shift by category, starting position, and how unknown your brand is.

Illustrative effort vs payoff by workstream (typical timing, not guarantees)
WorkstreamRelative effortTypical time to visible effectDurability
Crawl access and rendering fixesLow, mostly one-timeDays, once a page is re-fetchedHigh, until you break it again
Schema and authorship cleanupLow to medium, one-timeDays to weeksHigh, with light upkeep
Reshaping high-value pagesMedium, per pageWeeks on retrieval-heavy enginesMedium, needs freshness upkeep
Original data and researchMedium to highWeeks to months as it gets citedHigh, you become the source
Off-site authority and digital PRHigh, ongoingMonths, compoundingHighest, hardest to copy
Consistency and fact cleanupLow to medium, one-timeOften improves accuracy firstHigh, with periodic re-audit

Read down the durability column and the strategy writes itself. The cheap, fast items are table stakes you do once and maintain. The expensive, slow items are where the defensible advantage lives, precisely because competitors find them hard to copy. A brand that only ever does the fast work will keep pace with every other brand that read the same checklist. A brand that also does the slow work becomes the answer.

Common mistakes and pitfalls

Most GEO efforts do not fail because the brand did too little. They fail because the brand did the wrong things in the wrong order, or chased a shortcut that the engines are built to ignore. The recurring traps:

  • Blocking the very bots you need. A bot-management rule or a blanket robots.txt disallow can quietly exclude the live-retrieval agents. This is the most common single failure, and the most invisible.
  • Hiding content behind JavaScript. If the answer only appears after client-side rendering, a non-rendering crawler sees nothing. Server-render the substance.
  • Optimizing for keywords instead of questions. Engines answer questions. A page titled for a keyword string but written as marketing copy will not match how a customer phrases a query out loud.
  • Vague, adjective-heavy claims. "Industry-leading" and "best-in-class" are unquotable. They give the model nothing to lift and nothing to verify, so it reaches for a competitor who gave it a number.
  • Treating GEO as on-site only. No amount of page polish substitutes for off-site corroboration. A claim that lives only on your domain reads as marketing to every engine.
  • Chasing one engine. Tuning everything for ChatGPT while your customers ask Perplexity wastes the effort. Map your customers first.
  • Manipulating community signals. Fake Reddit threads, paid reviews, and astroturfed mentions get detected and discounted, and the cleanup costs more than doing it honestly would have.
  • Letting pages go stale. Freshness bias means your best page slowly loses citation share if you never update it. Treat top pages as living documents.
  • Believing the guarantee. No one controls what a model says. Anyone promising a guaranteed citation is selling certainty that does not exist.
  • Copy-pasting a US playbook into China, or vice versa. Domestic Chinese engines run on a different corpus and different authority signals. The two markets must be planned and measured separately, a point we return to in GEO vs SEO in 2026.

Metrics to watch and a working checklist

Classic SEO watches rankings. GEO watches citation share and how you are described, because being on page one means nothing if the answer above the links names someone else. Track a small set of metrics consistently rather than a large set occasionally.

Metrics worth tracking for AI visibility
MetricWhat it tells youHow to read it
Citation shareHow often you appear in answers to your question setThe headline number. Track per engine; they move independently.
Description accuracyWhether the model describes you correctlyAn early indicator. Accuracy usually improves before citations do.
Sentiment and framingWhether you are named as a leader, an option, or a cautionBeing mentioned negatively is its own problem to fix.
Competitor citationsWho gets cited instead of you, and from what sourceYour roadmap. Their sources are the gaps you need to close.
AI crawler hitsHow often GPTBot, PerplexityBot, and others fetch your pagesA leading indicator in your server logs. Access precedes citation.
AI-referred trafficVisits arriving from AI answer surfacesSmaller than search traffic, but high-intent. Watch the trend.

A pre-publish checklist for any page you want cited

  • The heading is phrased as a question a customer would actually ask.
  • The first 40 to 60 words answer that question on their own.
  • The brand and product are named explicitly, not hidden behind pronouns.
  • At least one real, contextualized number appears near the top.
  • Accurate Article and, where relevant, FAQPage schema is in place.
  • Author and publisher map to a real, named expert.
  • The core content is in server-rendered HTML and reachable by retrieval bots.
  • Dates are honest, and the page is on an update schedule.
  • At least one claim is corroborated by a credible third party off-site.

How GEO fits with the rest of your growth engine

Getting cited by AI is not a standalone channel bolted onto the side of marketing. It is the visibility layer that sits on top of everything else you do, and it draws strength from the same foundations. Strong organic SEO widens the eligibility pool the engines pull from. Digital PR builds the off-site corroboration. A clear product story and consistent facts make the entity legible. For B2B specifically, AI visibility feeds the top of a pipeline that you then run with verified prospect data, where our role is to deliver the list and yours is to run the outreach; we never contact your prospects. That model is laid out in the B2B growth engine and, for named-account work, in B2B ABM for named accounts. For consumer brands, the same visibility work supports the demand engine described in the B2C growth playbook.

The throughline is that none of this is a hack. It is the same discipline that wins classic search, applied to a surface that reads more carefully and trusts more selectively. Be verifiable, be structured, be genuinely the best answer to the question, and earn the outside corroboration that lets an engine repeat you with confidence. Do that consistently and the engines will pass you along. Try to shortcut it and they will route around you to a competitor who did the work.

Frequently asked questions

Can I pay to get my brand cited by ChatGPT or Perplexity?

No. There is no paid placement that makes a model name you in an organic answer, and anyone offering one is selling something that does not exist. You earn citations by being reachable, quotable, and corroborated by sources the engine trusts. Paid ads inside some AI products are a separate, clearly-labeled surface and are not the same as being cited in the answer itself.

How long does it take to start getting cited?

It varies by engine and by how unknown you start. Retrieval-heavy engines like Perplexity can pick up a fresh, well-structured page within days once crawl access is clean. Trained-memory effects and reference-grade authority build over months. A realistic expectation is foundation work in month one, early movement on live-retrieval engines in month two, and compounding authority gains across month three and beyond.

Is GEO just SEO with a new name?

They overlap but are not the same. Strong SEO widens the pool of pages an engine can consider and shares many fundamentals: crawlability, quality, structure. GEO adds its own demands: answer-first content, explicit entities, schema that hands the model question-and-answer units, and off-site corroboration weighted toward the specific surfaces each engine reads. Ranking on Google no longer guarantees a citation, so the two now need separate measurement. We unpack the boundary in GEO vs SEO in 2026.

Do I need a Wikipedia page to get cited?

It helps, especially for ChatGPT, because reference-grade sources feed trained memory across multiple models. But you cannot manufacture one. Wikipedia requires genuine notability and independent sourcing, and a promotional entry will be removed. The right path is to earn the independent coverage that makes a legitimate entry possible, and in the meantime keep your structured facts consistent on directories and knowledge-graph sources.

Should I block AI crawlers to protect my content?

Decide deliberately rather than by default. Blocking training crawlers is a legitimate choice if you object to your content training future models. Blocking live-retrieval agents, however, removes you from the answers those engines generate right now, which is usually the opposite of what a brand seeking visibility wants. Most brands allow the search and user agents while debating only the training crawlers.

Does llms.txt actually do anything yet?

As of early 2026, no major AI company has committed to reading it in production, so it is not a citation strategy on its own. It is cheap to publish and useful for some AI-assisted developer tools, so adding it is reasonable. Just do not let it crowd out the work that actually decides citations: crawlable, quotable pages and off-site corroboration.

What kind of content gets cited most often?

First-party data and original research lead, because there is nowhere else for the engine to source the figure. After that, bottom-funnel pages such as comparisons, pricing, and real case studies tend to outperform generic explainers, and clean definitions plus attributed expert quotes are highly quotable. The common thread is specificity the model can lift and verify.

How do I measure whether my GEO work is paying off?

Track citation share across your customer question set per engine, the accuracy of how you are described, which competitors and sources are cited instead of you, AI crawler hits in your server logs, and AI-referred traffic. Re-run the question set monthly and watch the trend. Description accuracy usually improves before citation share does, so treat it as an early signal.

Does the same playbook work for the Chinese market?

No, and copy-pasting it is a common mistake. Domestic Chinese engines run on a different corpus and reward different authority signals, so the two markets must be planned, executed, and measured separately. For brands going in either direction, that dual-market reality is part of why we treat US and China growth as two engines rather than one.

Can a small brand compete with established players for citations?

Yes, on the right questions. A small brand will struggle to outrank an incumbent for the broadest category terms, but it can win narrower, more specific customer questions where it genuinely has the best answer, especially with original data and clear, quotable structure. Pick the questions where you can credibly be the best source, and earn citations there first.