Coruve

Pillar guide

The complete guide to AI referral traffic

Assistants now send real people to real websites, and almost no analytics tool tells you which ones. Here is what the traffic actually looks like on the wire, which signals can be trusted, and where every method — including ours — goes blind.

7 min read

What AI referral traffic actually is

One sentence, and then the distinction the rest of this guide is built on.

AI referral traffic is a person who arrived at your website because an AI assistant pointed them at it. They asked ChatGPT or Perplexity or Gemini something, your page was named in the answer, and they clicked. They are a visitor like any other — they can read, buy, sign up and leave — and the only unusual thing about them is where the recommendation came from.

That is worth stating flatly because the phrase “AI traffic” is used for two completely different things, and mixing them up is the single most common error in this whole subject:

  • An AI referral is a human being. It is demand. It can convert.
  • An AI crawler is a program fetching your pages — to train a model, to build a search index, or because somebody just asked an assistant about your site. It cannot buy anything.

Both are interesting. They answer different questions, and adding them together produces a number that answers neither. We wrote a whole post on telling them apart — AI crawlers vs AI referrals — because the distinction is where most tools quietly go wrong.

Why it is growing and why your analytics probably cannot see it

Nothing is broken. The categories were drawn before assistants existed.

Every mainstream analytics tool sorts arrivals into a handful of channels: direct, organic search, social, referral, email, paid. That list is a good description of how people found websites for about twenty years. An AI assistant is none of those things, so for most of the last two years a visit from one had nowhere correct to go — and the tool did not tell you it had a problem. It just filed the visit somewhere.

This is changing, and it is worth being accurate about rather than selling against. Google Analytics 4 added an AI Assistant channel to its default channel group, announced on 13 May 2026. If you use GA4 and your traffic is recent, some of this is now categorised for you. The ChatGPT post goes through exactly what that channel does and does not cover, because the limits are specific and they matter.

But that channel classifies traffic from the point it arrived in your property onwards, not your history — and plenty of tools still have no AI channel at all. So for most sites, most of the AI arrivals to date are still sitting somewhere else. Where they landed depends on details you never chose:

  • If the browser sent a referrer, the visit usually becomes an ordinary referral, sitting in a long list next to a forum and a directory — technically true, completely uninformative.
  • If the referrer was stripped, or the link was opened from a phone app rather than a browser tab, the visit becomes direct, which is the honest word for “we have no idea”.
  • Worst case, it is added to organic search. Some assistant surfaces live on hostnames that a broad search-engine matching rule will happily swallow, and then AI arrivals are inflating a number you use to judge your SEO. We hit exactly this and had to fix it — the details are in the per-assistant post.

So the growth is invisible twice over: the traffic is new, and the reporting categories predate it. If you have noticed “direct” creeping up over the last year and have not been able to explain it, this is one of the plausible explanations, and it is a testable one.

The three signals that reveal it

There are exactly three, they are not equally trustworthy, and the difference matters more than the total.

1. The referrer — a header the browser sets

When someone clicks a link, their browser usually tells your server which page they came from. If that value is chatgpt.com or perplexity.ai or claude.ai, you have a strong, practical signal that an assistant sent this person.

Strong, but not proof. A referrer is a string the client chose to send, and any client can send anything. It is believed because browsers do in fact send it honestly, not because it was checked.

2. The user agent — which tells you it is not a person at all

The same header field that says “Chrome on a Mac” is where crawlers identify themselves: GPTBot, ClaudeBot, PerplexityBot and the rest. This signal does not find you referrals — it finds you the machines, so they can be counted separately instead of being mistaken for visitors.

3. The campaign tag — the weakest one, and still worth reading

ChatGPT stamps its outbound links with utm_source=chatgpt.com. That tag survives the two things that destroy a referrer: a privacy setting that strips it, and the hop from a native app into a browser where no referrer exists in the first place.

It is also the weakest signal on this page, because a campaign tag is simply text somebody typed into a link, and anybody can author a link. The right way to use it is narrow: read it only when every other signal has already come up empty, so it can turn “we know nothing” into a named guess, and can never overrule or rename a visit that a referrer already explained. That is the rule Coruve follows, and the ChatGPT post walks through it properly.

Why the three are never added together

A number built by summing a cryptographic proof, a believable header and a typed-in string inherits the trustworthiness of the weakest input. So Coruve keeps them in separate buckets — verified, inferred and unattributed — on every screen and in every export, and refuses to print a single blended “AI traffic” figure. It is less impressive and it is the number you can actually act on.

Four things can arrive from an AI company. Only one is a customer

Crawl, cite, click — plus the training fetch that started it all.

Most AI operators run more than one fetcher, and they do different jobs. Rolled into a single “bot hits” number they are mute; separated, they describe a funnel that runs from your text entering a model all the way to somebody knocking on your door.

What arrivedIs it a person?What it means for you
Training fetchNoYour writing is being read as material for a model. Nothing happens on your site today because of it.
Search-index fetchNoYour page is entering an index the assistant can draw on later. This is the groundwork for being cited.
On-demand fetchNo — but a person caused itSomebody is sitting in front of an assistant right now asking about your page. The closest a crawler ever gets to a visit.
A referral clickYesA visitor. Measure them like any other visitor: what they read, what they did, whether they came back.

The third row is the one people miss, and it is the interesting one: an on-demand fetch is a person’s curiosity, arriving as a program. It is still not a visit, and counting it as one inflates every visitor number you own — which is a mistake we made and had to fix, described in detail in the crawlers versus referrals post.

How to measure it

What to ask of any tool, including this one.

You do not need a particular vendor to see some of this. You need four things, and it is entirely reasonable to hold every tool to them:

  • Assistants get their own channel, not a shared bucket with forums and directories, and not folded into organic search.
  • Each assistant is named separately. “AI: 4%” is not an insight. “Perplexity sends a third as many people as ChatGPT but they read twice as long” is.
  • Humans and crawlers are counted apart, always, with no screen that quietly sums them.
  • The tool tells you how confident it is, and tells you what it cannot see.

Coruve does this natively rather than as a configuration exercise. It currently recognises 12 assistant products a person can arrive from and 17 AI crawler signatures, with the lists last reviewed by a human on 2026-08-10. Those two figures are computed from the matching tables themselves, which is why they are allowed on a marketing page at all — nobody can forget to update them.

Hosts are matched exactly, or as a true subdomain, and never as a fragment of a longer string. That sounds pedantic until you notice that a naive “contains claude.ai” rule also matches notclaude.ai and claude.ai.example.test, both of which anybody can register.

If you would like the answer for your own site before deciding anything, the free AI traffic checker reads seven days and shows you the split with its caveats attached. It needs no account.

The blind spot every tool has, including ours

If a vendor has not told you this, they have told you something incomplete.

Coruve measures with a small script in your page. So does almost every privacy-first analytics tool, and so does Google Analytics. That has real advantages, and it has one consequence nobody advertises: an AI that reads your site without running JavaScript leaves no trace at all. It fetches the HTML and leaves. The script never runs, so no event is ever sent, so there is nothing to count.

Most crawlers behave exactly that way. Which means any crawler figure produced by a script-based tool — ours included — is a floor, not a share. “We saw at least this many” is supportable. “AI crawlers are 3% of your traffic” is not, because the top and the bottom of that fraction were not measured by the same instrument.

This is why our own free checker prints a count for crawler activity rather than a percentage, on the page where a big confident number would be most persuasive. If you want a complete picture of crawling, the place to get it is your server or CDN logs, which see every request whether or not any script ran.

The referral half is different, and much better: those are real browsers running real scripts, so referral measurement is as complete as any other traffic measurement you have.

What to actually do about it

Five things, in the order they are worth doing.

  1. Measure before you strategise. For most small sites AI referrals are still a small share. Small and growing is worth watching; small and flat is worth knowing so you stop reading breathless articles about it.
  2. Look at what they do, not just how many. Assistant referrals often arrive further down the decision than a search visitor — they have already been given a summary and a reason. Compare their behaviour against your other channels before you conclude anything about their quality.
  3. Check which pages get cited. If assistants keep sending people to one explainer, that is a signal about what you are useful for, and it is usually not the page you spent the most money on.
  4. Decide your crawler policy deliberately. You can allow or block training crawlers and retrieval crawlers separately in robots.txt. They are genuinely different decisions: blocking retrieval makes it harder for an assistant to cite you at all.
  5. Do not over-fit. This is a fast-moving area, hostnames change, and anyone quoting precise industry-wide percentages at you is guessing. Your own numbers are the only ones about your site.

The rest of this guide

Four posts, each going properly into one part of the above.

See which assistants send you people

The free checker reads seven days of your traffic and shows the split, with the caveats attached. No account, and nothing to install to see the answer.

AI Referral Traffic: The Complete Guide | Coruve