How to earn Claude AI citations: source selection and web search logic (2026)

Furkan Yaman
September 29, 2026
14 Mins
Articles

Claude cites sources inline in every web search answer, but eligibility is decided by a crawler most sites have never configured. This guide covers Claude citations end to end, from robots.txt to what actually gets quoted.

Key takeaways

  • Anthropic runs three separate crawlers. Blocking the training bot does nothing to the search bot. Blocking the search bot removes you from Claude citations entirely.
  • Claude has two knowledge sources, its training data and live web search. Only one of them can cite you, and it does not run on every question.
  • On Team and Enterprise accounts, web search is off until an owner enables it workspace-wide. Your B2B buyers may be asking Claude questions it answers without touching the web at all.
  • Claude quotes specific passages rather than linking a source tray, so extractable factual sentences matter more than page-level authority.
  • Cognizo tracks Claude alongside nine other platforms, reads the rendered answer through UI scraping, and reports six metrics rather than mention counts.

Every AI platform retrieves differently. The same question returns different brands in ChatGPT, Google AI Overviews and Claude, because the index, the trigger logic and the selection criteria are not shared. Treating them as one surface produces a strategy that fits none.

This article takes Claude on its own terms: how Claude decides to search, which crawler controls whether you can be cited, what Claude source selection rewards, and how to measure the result. It also covers what Anthropic does not publish about its retrieval stack, because a lot of advice on Claude citations treats guesses as documented fact.

Claude is not a search engine, and that changes the job

There is no Claude index you submit to, no Claude equivalent of Search Console, and no verification step. Claude answers from two places, and only one of them produces Claude citations.

Two knowledge sources, one citation path

The first source is the model's training data. Ask a stable factual question and Claude answers from what it already knows. No search runs, no sources appear, and nothing you publish this quarter affects the answer.

The second source is live web search. Anthropic's web search documentation describes it plainly. When a topic benefits from current information, Claude invokes a search tool to ground its response in live web content, and every response includes citations. Claude web search citations come from this path and no other.

That split matters for planning. Publishing more content does not change what Claude already believes about your category. It changes what Claude finds when it searches. Those are different levers with different timelines, and only the second one produces Claude citations this quarter.

Knowing which prompts fall on each side is the first thing worth measuring. Cognizo captures the rendered answer, so a prompt that returned no sources is visibly a training-data answer rather than a citation you lost.

When Claude decides to search

Anthropic's web search tool documentation publishes the trigger taxonomy, and it is more specific than most people assume.

Claude searches when a request depends on information that is current, changing, or outside its training data. The documented examples are recent events and announcements, current prices and statistics, explicit requests to look something up, and information about specific organizations, people or products that might have changed.

Read that last category again. Questions about specific organizations and products are named as a search trigger. That is most of your commercial prompt set, and it means brand and product queries are the ones where publishing actually moves Claude citations.

Claude answers directly, with no search and no sources, for established facts, maths and science fundamentals, coding concepts, creative writing, analysis of content already in the conversation, and ordinary conversational turns.

Users can also force a search by asking for one or suppress it by asking Claude not to. In the newer Claude experience there is no toggle, and Claude searches when it helps.

The practical read for a brand: your prompt universe splits into questions where you can compete now and questions where you cannot. Product comparisons, pricing, current capabilities and anything dated fall into the first group. Category definitions and stable concepts mostly fall into the second. Only the first group is winnable through publishing, and Cognizo's Prompt Volumes shows which prompts buyers actually use before you commit budget to either.

The 150-character citation window A ruler marked from zero to 150 characters with a cut line at 150. A claim spread across three sentences extends past the cut line and is truncated, so it never arrives intact in a citation. A claim written as one short sentence ends before the cut line and travels whole. 150 characters is roughly 25 words.
Each citation carries up to 150 characters of your text. Source: Anthropic web search tool documentation.

The three crawlers that decide whether Claude can cite you

This is the part that silently disqualifies sites from Claude citations, and it is the cheapest thing on this list to fix.

What each bot actually does

Anthropic's crawler documentation lists three robots, each with its own user-agent token and its own consequence when blocked.

ClaudeBot collects web content that may contribute to model training. Blocking it signals that your future material should be excluded from training datasets. It has no effect on whether Claude can cite you.

Claude-SearchBot navigates the web to improve search result quality, analysing content to make responses more relevant and accurate. Anthropic states that disabling it prevents the system from indexing your content, which may reduce your site's visibility and accuracy in search results. This is the bot that gates Claude citations.

Claude-User fetches pages when a user asks Claude a question. Anthropic states that disabling it prevents retrieval of your content in response to a user query, reducing visibility for user-directed web search.

Read those consequences again. One is a licensing decision. Two are visibility decisions.

Three bots, three consequences Three Anthropic crawlers compared. ClaudeBot collects content for model training and blocking it has no effect on citations. Claude-SearchBot indexes content for search quality and blocking it removes the site from Claude search results. Claude-User fetches pages on a user's request and blocking it makes the site unretrievable for user queries. A blanket user-agent wildcard rule catches all three at once.
One licensing decision, two visibility decisions. Source: Anthropic crawler documentation.

The mistake that removes you from Claude entirely

In 2024, many sites added a blanket AI crawler block. Search Engine Journal's coverage of Anthropic's documentation update noted a study finding that while 79% of top news sites block at least one AI training bot, 71% also block at least one retrieval or search bot, potentially removing themselves from AI search citations in the process.

That is the failure mode. A team decides not to feed model training, pastes in a block list, and disqualifies the domain from Claude citations without realising it. The symptom looks like an unexplained absence on one platform while others hold steady.

Blocking ClaudeBot does not block Claude-SearchBot. Each token needs its own directive, which is the whole point of the three-bot split.

Configuring robots.txt for Claude

If you want training excluded but Claude citations preserved, the configuration is explicit:

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

Four details from Anthropic's documentation that get missed:

  • Every subdomain needs its own robots.txt entry. The root domain does not cover the rest.
  • IP blocking is not a reliable opt-out. It stops Anthropic reading your robots.txt in the first place. Anthropic publishes its crawler IP ranges so you can verify traffic in your logs rather than guess.
  • Crawl-delay is supported, so rate limiting is available without blocking.
  • A blanket Disallow: / under User-agent: * catches every bot you did not explicitly allow, which is how most accidental blocks happen.

Then check your CDN and bot-protection rules. A security product blocking unfamiliar user agents leaves no trace in robots.txt. Our guide to optimizing your site for AI crawlers covers the full audit.

Cognizo's AI Traffic Analytics tracks ClaudeBot alongside GPTBot, OAI-SearchBot and other major bots, so you see crawl behaviour rather than intended configuration. Config files say what you meant. Logs say what happened.

The enterprise gate nobody checks

Two documented controls sit between your content and an enterprise audience, and neither has anything to do with your site.

The workspace switch

On Team and Enterprise accounts, web search is off until an Owner or Primary Owner enables it for the entire workspace in organization settings. Until that happens, every employee in that company gets answers from training data only.

Think about what that means if you sell to enterprises. Claude skews toward business and technical users, which is exactly why it matters for B2B pipeline. But some of those users sit inside workspaces where Claude cannot reach the web. For them your robots.txt is irrelevant and your newest content is invisible. What matters is what the model absorbed during training, which is a slower and far less direct game.

Domain allow-lists

Organizations using Claude through the API can also restrict which domains web search is allowed to touch, through allow-lists and block-lists set at the tool or console level. An allow-list is absolute. If a company has scoped Claude to a set of approved sources and your domain is not among them, nothing you publish will ever reach those users.

You cannot influence another company's allow-list directly. What you can do is be present on the domains that tend to make those lists, which are usually established industry publications, standards bodies and major review platforms rather than vendor blogs. That is one more argument for treating earned placement as primary rather than supplementary.

The conclusion is not to abandon either path. Claude citations for enterprise audiences run on three gates now: whether the workspace has search on, whether your domain clears any allow-list, and whether you win the retrieval. Our guide to SEO for Claude covers the on-page side for business buyers.

Three gates in front of an enterprise reader Three sequential gates stand between a brand and an enterprise Claude user. First, an administrator must have enabled web search for the workspace, or Claude answers from training data only. Second, the domain must clear any allow-list the organization has set, or the content is never reachable. Third, the page must win the retrieval. The first two gates are controlled by the reader's organization. Only the third is controlled by the publisher.
Two of the three gates belong to someone else.

How Claude cites sources

Claude's citation behaviour differs from the summary-plus-source-tray pattern used elsewhere, and the difference shapes what you should publish.

Inline attribution, with quotes

Anthropic documents that a web search response includes direct citations to sources, source links for further reading, and relevant quotes where appropriate. Claude web search citations therefore arrive attached to specific claims rather than pooled at the end.

The quoting is the part worth internalising. Claude does not just link you. It lifts a specific passage and attributes it, which means the unit competing for attention is the sentence rather than the page. A strong page with no quotable sentence loses to a weaker page with one.

Cognizo's citation share metric splits this into owned and earned, so you can see whether Claude quoted your page or somebody else's page talking about you. Those two outcomes call for completely different work.

One more thing the documentation settles

Image results in Claude are powered by Bing, and Anthropic says so directly. Text retrieval is a separate system, and Anthropic does not publish which backend serves it.

This matters because a lot of writing about Claude citations asserts a specific text search provider as established fact, usually citing vendor research. Anthropic has not documented it. Treat any claim about Claude's text retrieval backend as inference, and do not build a strategy that depends on it holding. Measure what Claude actually returns instead, which is what Cognizo's per-platform tracking is for.

The second cut you never see

Retrieval is not one decision any more. On newer versions of the web search tool, there is a filtering stage between being found and being read.

Anthropic documents it as dynamic filtering. With basic web search, every search result is loaded into Claude's context window, and much of that content is irrelevant to the request. With the February 2026 tool version and later, Claude instead writes and runs code that filters the results first, so only relevant content reaches the context window.

For a publisher this adds a stage that did not exist before. Your page can rank well enough to be retrieved, then be cut programmatically before the model ever reads it. Surviving that cut is a different test from surviving retrieval. It rewards pages where the relevant material is identifiable quickly and concentrated, and it penalises pages that bury one useful paragraph inside a long, loosely related article.

The practical response is the same discipline that helps everywhere else, applied harder. Put the answer near the top of the section. Make the section self-describing. Keep the page focused on the question it claims to answer, rather than covering six adjacent topics to chase more keywords.

Search volume per question also varies in a way worth knowing. Anthropic notes that simple factual queries typically use one to three searches, while comparative or multi-entity research can use ten or more. Comparison prompts therefore run more retrieval passes than definitional ones, which is a structural reason "best tools for X" questions are worth tracking separately. Cognizo's topic grouping handles that split, so comparison prompts are not averaged in with the easy ones.

Claude source selection: what earns the citation

With the mechanics settled, the question becomes what makes Claude choose your passage over someone else's.

Write sentences that can be quoted

Since Claude extracts passages, the highest-yield editing pass is making individual sentences self-contained. A sentence opening with "this approach" or "as mentioned above" cannot be lifted. Its subject lives somewhere the retrieval did not go.

Name the subject, state the claim, keep the qualifiers inside the sentence. Specific figures survive extraction better than ranges, since a number is hard to paraphrase away.

The 150-character constraint

Here is the most actionable detail in Anthropic's documentation, and almost nobody writing about Claude citations mentions it.

Each citation returned by the web search tool carries a cited_text field containing up to 150 characters of the cited content. One hundred and fifty characters is roughly 25 words. That is the size of the window your claim has to fit inside.

The implication is concrete. A claim spread across three sentences cannot be carried in a single citation. A claim compressed into one 20-word sentence can. When you have something worth being quoted on, a competitive figure, a definition, a benchmark, write it as one short self-contained sentence rather than as a well-argued paragraph.

This does not mean writing in fragments. It means that within a normal paragraph, the load-bearing claim should occupy one sentence that could stand alone at 25 words or fewer. The surrounding prose does the persuading. That one sentence does the travelling.

The second cut: filtering between retrieval and reading A four-stage pipeline. Search returns results, all of which are retrieved. A new filtering stage then runs code that removes some results before they reach the context window, so only the remainder are read by the model. Pages can be retrieved and still be cut before the model sees them.
Tile counts are illustrative. The filtering stage is documented. Source: Anthropic web search tool documentation.

Cognizo's UI scraping captures which passage Claude actually used, so you can compare the sentence you intended to be quoted against the one that got picked up.

Give it something only you have

Commodity passages are interchangeable. If forty domains carry the same paragraph, selection falls back on authority signals, and a mid-authority domain loses every time.

Original measurement, dated first-hand testing, named expert attribution and documented methodology all produce passages with no substitute in the candidate set. This is the biggest lever on Claude source selection, and the slowest to build.

Work the third-party surface

Most citations across AI platforms point at domains other than yours. Review sites, comparison articles, industry publications and community threads carry a large share of what models retrieve. Optimizing your own site therefore addresses only part of the problem.

Identify the domains Claude already cites on your topics, then work out which you can appear on. Cognizo's source mention rate metric names them rather than leaving it to guesswork. Our guide to third-party AI citations covers the outreach side.

Keep pages current and dated

Claude searches when a question benefits from current information, so freshness is not a ranking nicety here. It is part of the trigger condition.

There is also a mechanical reason, not just an editorial one. Every search result returned to Claude carries a page_age field recording when the site was last updated. Freshness is not inferred from your copy. It is handed to the model as a field alongside the content.

Date your claims explicitly in the text as well. A passage reading "as of March 2026, pricing starts at" survives a stale crawl better than "pricing currently starts at." The model can weigh an explicit date rather than assume one.

Your citation travels further than the Claude app

Claude web search runs inside the API as well as the consumer app, which means it powers products you will never see branded as Claude.

Anthropic's terms for the feature are direct: when displaying API outputs directly to end users, citations must be included to the original source. So a company that builds a research assistant, a support bot or an internal knowledge tool on Claude web search is obliged to surface your attribution when it uses your content.

This changes the size of the prize. A Claude citation is not only an impression inside one chat interface. It is a unit of attribution that propagates into every downstream product grounded on Claude's web search, most of which have their own audiences and none of which will send you a report.

It also explains a measurement gap. You can track what Claude itself says. You cannot track what a third-party product built on Claude says, and nobody can.

The practical answer is to measure the surface you can see, continuously rather than by sampling, and treat it as the leading indicator for everything downstream. That is what Cognizo's daily Claude tracking is for: your Claude citations in the app are the closest available proxy for your citations everywhere Claude is embedded.

How to track Claude citations

You cannot optimize what you cannot see, and Claude publishes nothing to site owners. No Search Console equivalent, no impression report, no query data, no notification when you gain or lose a citation.

That leaves two options. Ask Claude the same questions manually and record the answers, which collapses past a few dozen prompts. Or track Claude citations properly.

What Cognizo does here

Cognizo tracks Claude as one of ten platforms, alongside ChatGPT, Google AI Overviews, Google AI Mode, Gemini, Perplexity, Microsoft Copilot, Meta AI, Grok and DeepSeek. Retrieval differs per platform, so each is tracked separately rather than blended into one score.

It reads the rendered answer. Cognizo captures responses through UI scraping, meaning the answer as a user sees it on screen rather than an API sample. Claude attributes inline and quotes specific passages, so that distinction decides whether you can see which sentence was used and how prominently you were named.

It reports six metrics. Answer Engine Insights covers Visibility Score, share of voice, citation share, source mention rate, sentiment and positioning accuracy, broken down by model, topic, prompt and region. Citation share splits owned from earned, telling you whether your own pages or third-party properties are carrying you.

It runs daily on every tier. Retrieval changes day to day, so a monthly check cannot separate a decline from ordinary variance.

It closes the loop. Content Optimization turns a Claude citation gap into a prioritised brief and a draft. Prompt Volumes shows which prompts buyers actually use. MCP and API access ship on every paid tier, so the data reaches your own reporting or agent without a sales conversation.

Pricing is $499 a month on Growth and $999 on Pro. Enterprise runs on custom terms with all ten platforms and multiple workspaces. Our guide to Claude rank tracking covers the measurement mechanics in more depth.

Common mistakes

Pasting a blanket AI crawler block. It stops training collection and disqualifies you from Claude citations at once. Audit yours today.

Assuming Claude works like Google. No index to submit to, no verification, no reporting. Tactics built around a search console have nothing to attach to.

Optimizing only your own domain. Earned citations dominate, so much of the work sits on properties you do not own.

Treating asserted retrieval backends as fact. Anthropic documents Bing for image results and nothing for text retrieval. Strategies built on unverified plumbing break when the plumbing changes.

Measuring Claude inside a blended AI score. Averaging Claude with ChatGPT and AI Overviews hides the platform-specific gaps you are trying to fix. Our explainer on how AI search engines find, rank and cite covers why that divergence is structural.

The same reasoning applies surface by surface. What earns a citation in Claude is not what earns one elsewhere, which is why earning Gemini citations needs its own playbook rather than a copy of this one.

Frequently asked questions

What are the best tools to track Claude citations?

Claude provides no first-party reporting, so tracking Claude citations means running prompts against the platform and recording results systematically. Cognizo tracks Claude alongside nine other platforms, capturing responses through UI scraping so you see the rendered answer rather than an API sample. It reports Visibility Score, share of voice, citation share, source mention rate, sentiment and positioning accuracy per prompt, with daily tracking on every tier. Manual checking works for a few prompts and stops scaling after that.

How do I know if Claude can even reach my site?

Check your server logs rather than your robots.txt file, because what the server returns is the only thing that counts. Look for Claude-SearchBot and Claude-User over the last 30 days and confirm they receive 200 responses. Anthropic publishes its crawler IP ranges so you can verify the traffic is genuine. Then check CDN and bot-protection rules separately, since a security product can block crawlers your robots.txt permits. Cognizo's AI Traffic Analytics reports this continuously rather than leaving it as a one-off log review.

Does blocking ClaudeBot stop Claude from citing me?

No, and this is the most useful distinction in the whole topic. ClaudeBot collects training data only. Claude-SearchBot handles search indexing and Claude-User handles retrieval at a user's request, and both need their own robots.txt directives. You can block training while keeping citation eligibility by disallowing ClaudeBot and allowing the other two. The reverse mistake, blocking all three with a blanket rule, removes you from Claude answers entirely.

How is Claude source selection different from Google's?

Claude has no public index, no submission process and no reporting surface, so nothing transfers from a Search Console workflow. Claude also quotes passages inline with attribution rather than putting a summary above a link tray. That shifts competition from page-level authority toward individual extractable sentences. Retrieval only runs when Claude judges a question to need current information, so some questions about your category never touch the web.

Why does my brand appear in ChatGPT but not in Claude?

Different platforms use different indexes, trigger logic and selection criteria, so divergence is normal rather than a bug. Check crawler access first, since Claude-SearchBot needs its own permission and is often blocked by rules written for other bots. Then look at whether your ChatGPT presence comes from third-party domains that Claude weights differently. Tracking each platform separately is the only way to tell which explanation applies, which is why Cognizo keeps all ten surfaces on their own scorecards instead of one blended number.

Do Claude citations send traffic?

Some, and less than the citation count suggests. Claude attributes inline with clickable source links, so the click path exists, but many users read the answer and never follow it. The more reliable downstream signal is branded search volume rising while direct referral stays flat. Pair that with a "How did you hear about us?" field carrying an explicit AI option, since the path from a Claude citation to a conversion carries no UTM. Cognizo's AI Traffic Analytics connects crawler activity to the sessions and conversions that follow.

How long does it take to earn Claude AI citations?

Crawler access fixes can take effect quickly, since restoring permission restores eligibility as soon as recrawling happens. Content changes are slower. A new passage must be crawled, indexed, then win against an existing candidate set, which usually takes several weeks before it appears in answers. Third-party placement work is slower still. Sequence accordingly: fix access first, because everything else depends on it. Daily tracking is what tells you when each change actually lands.

Does Claude web search run for every question?

No. Claude evaluates whether a question benefits from current information and searches only when it does. Stable or definitional questions get answered from training data with no sources at all. Users can force a search by asking for one, or suppress it by asking Claude not to search. On Team and Enterprise accounts an owner must enable web search for the whole workspace first, so some business users have it switched off entirely.

‍