Getting cited by ChatGPT and Perplexity depends on three things: a clearly identifiable entity backed by consistent structured data, passages written to be extracted (self-contained, 40–80 words, answering one question completely), and claims that are verifiable against other sources. Letting AI crawlers access your site matters too, but identity, extractability and verifiability are what actually earn the citation.
Why classic SEO is not enough
A search engine returns a list and lets the user choose. An answer engine writes a response and picks a handful of sources to support it. Those are different selection problems. Ranking well helps, because most answer engines retrieve from search results, but it is not sufficient. Plenty of page-one sites are never quoted, and plenty of quoted sources are not first.
The difference comes down to three things: whether the model can confidently identify who you are, whether your page contains a passage that can be lifted as an answer, and whether your claims are verifiable against other sources.
1. Make your identity unambiguous
A model will not cite a source it cannot identify. If your site describes an organisation but nothing anchors that organisation to a verifiable identity, you are an unnamed string rather than a known entity.
In practice this means one consistent Organization record in your structured data, with a stable identifier reused on every page rather than re-declared anonymously each time, and a set of sameAs links pointing to profiles that genuinely exist and genuinely belong to you.
- Define the organisation once with a stable
@id, then reference that identifier from every other page instead of repeating the object. - List
sameAsURLs for every profile you actually control: LinkedIn, the company registry, Crunchbase, YouTube, your Google Business Profile. - Keep your name, address and phone number byte-identical everywhere, including in the schema, the footer and every external listing.
- Never list a
sameAsURL that 404s. An unverifiable claim is worse than no claim.
2. Write passages that can be lifted
Answer engines extract passages. A paragraph that only makes sense after reading the three paragraphs above it cannot be extracted, no matter how good it is. The unit that gets quoted is a self-contained block of roughly 40 to 80 words that answers one question completely.
What an extractable answer looks like
Lead with the direct answer in the first sentence. Follow with one or two sentences of specific support: a number, a timeframe, a named standard. Do not open with context, do not open with a question, and do not require a pronoun to resolve against an earlier paragraph.
A useful test: read the paragraph aloud with no other context. If a stranger would understand it and learn something specific, it is extractable. If they would ask what it refers to, it is not.
3. Be verifiable
Models weight sources that agree with other sources. A claim that appears only on your own site, with no corroboration anywhere, is treated cautiously. A claim that matches a company registry, a news mention or an independent profile is treated as established.
This is why off-site presence matters more for AI citation than for traditional ranking. You are not just building links, you are building corroboration. A business that appears consistently across a registry listing, a Google Business Profile, a LinkedIn page and a handful of independent mentions is straightforwardly verifiable. One that exists only on its own domain is not.
4. Publish the things people actually ask
Service pages rarely get cited, because they answer the question 'what do you sell' rather than 'how does this work'. Citations go to pages that answer a specific question better than the alternatives. That means genuinely useful writing on narrow topics, not broad overviews.
Narrower is better. A page on how a specific process works in a specific market will be cited far more often than a general guide to your industry, because there is far less competition for the specific question and far more precision in the answer.
5. Let the crawlers in
Finally, and obviously, check that you are not blocking the agents you want to be read by. Many sites unintentionally exclude AI crawlers through a restrictive robots.txt or an aggressive firewall rule. Our llms.txt guide covers the newer, complementary file worth adding alongside robots.txt.
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Google-Extended
Allow: /
Also add max-snippet:-1 and max-image-preview:large to your robots meta tag. The default snippet cap limits how much of your text an engine may display, which directly limits how much of it can appear in an answer.
The honest summary
None of this is a trick, and none of it produces results in a week. Being cited by an answer engine is the outcome of being genuinely identifiable, genuinely useful and genuinely corroborated. The technical work removes the obstacles. The content and the off-site presence are what actually earn the citation.