- Most assistant influence is unattributable by construction. Someone reads an answer, then searches your brand — that visit lands as direct or branded search, and no configuration recovers it.
- Channel labels moved underneath everyone. In our own property the same assistants were classified as
referral, then no medium at all, thenai-assistant. Any comparison spanning that change measures the labelling, not the traffic. - Define the channel yourself against raw source values, so a vendor taxonomy change doesn't silently reshape your history.
- The only direct evidence is a manual citation log. Ten real customer questions, three assistants, once a month. Half an hour, and it tells you what you're cited for and who beats you.
- Blocking training and blocking citation are different decisions controlled by the same robots.txt. Blocking the agents that fetch to answer questions removes you from the answers.
- There is no separate AI optimisation discipline. Question-shaped headings, direct answers, checkable specifics and something only you can say — the same things that make a page useful to a human in a hurry.
Everyone wants to know whether AI assistants are sending them business. Most analytics setups genuinely cannot answer that question, and a fair number are quietly answering it wrong.
Before you can decide what to do about AI search, you need to be able to see it. This is how to set up the measurement, how to check whether you're being cited at all, what actually influences it, and how to think about the crawler question — which is a decision most sites make by accident.
Why you probably can't see it today
Three separate problems, and they compound.
Most of it arrives unattributed. Someone asks an assistant a question, gets an answer that mentions you, and then searches your name or types your URL. That's a real, assistant-caused visit, and it lands in your analytics as direct or branded search. The influence is invisible by construction, and no configuration fixes it.
What is attributable gets classified inconsistently. Referrals from assistants arrive with a referrer like chatgpt.com or perplexity.ai, and where those land in your channel groupings has changed over time.
Nobody set it up. Default channel groupings weren't built with this category in mind, so unless someone deliberately configured it, assistant traffic is scattered across referral, direct and other.
The classification shifted, and it broke the comparison
That's from our own analytics property. The same three hosts, plotted by how they were labelled over time.
Read the ChatGPT row. Until early June 2026 it was landing as a referral. There's a period where it arrived with no medium set at all. From late June it appears as ai-assistant. Same source, same kind of visit, three different labels — and no announcement that would appear in anyone's inbox.
The practical consequence: any year-over-year or half-over-half comparison spanning that change is measuring the labelling, not the traffic. If your dashboard shows AI traffic appearing from nothing in mid-2026, some of that is real growth and some is a reclassification, and you cannot separate them after the fact unless you kept the raw source data.
The lesson generalises past this one case. Platform-defined channel groupings are convenient and they are not stable. If a number matters to you, define it yourself against raw source and medium values, so that when a vendor changes their taxonomy your history doesn't silently change shape.
Setting up the measurement
The goal is one report that answers "how much traffic arrived from an AI assistant, and what did it do."
1. Build a custom channel group rather than relying on the default. Create a rule matching source against the assistant hosts, and keep the list in version control somewhere, because it grows. At time of writing the hosts worth matching include:
chatgpt.com, chat.openai.com, openai.com, perplexity.ai, claude.ai, gemini.google.com, bard.google.com, copilot.microsoft.com, bing.com/chat, you.com, poe.com, phind.com
2. Match on the raw source, not the platform's label. That's what protects you from the reclassification problem above.
3. Segment landing pages. Which pages assistants send people to is the single most actionable output. It tells you what you're being cited for, and it's usually not your homepage — it's a specific page that answered something well.
4. Look at behaviour, not just volume. Assistant referrals tend to arrive better-informed than search traffic, because they've already had their question partly answered. Compare engagement and conversion, not just sessions.
5. Keep a manual citation log. The most useful measurement isn't in analytics at all — see below.
Check whether you're actually being cited
Analytics only sees people who clicked. Citation without a click still shapes a decision, and you can only see it by looking.
Once a month, take the ten questions your customers actually ask — the ones your sales team answers weekly — and put them to the main assistants. Record whether you appear, which page is cited, and who is cited instead.
That log is more useful than any dashboard, for three reasons. It tells you whether you're in the answer set at all. It tells you which of your pages does the work. And it tells you who your real competitors are in this channel, which is frequently not who you think — often it's a trade publication, a directory, or a competitor's blog post from four years ago.
Keep it simple: a spreadsheet with the question, the date, the assistant, cited or not, and the URL. Ten questions across three assistants is half an hour a month, and it's the only direct evidence available.
What actually influences citation
Nobody outside these companies knows the ranking mechanics, and anyone selling you certainty is guessing. But the properties shared by pages that do get cited are consistent, and they're unsurprising:
- A question as a heading, answered immediately underneath. Assistants extract passages. A heading phrased the way a person asks, followed by two or three sentences that actually answer it, is dramatically more extractable than the same information buried mid-paragraph.
- Specifics that can be checked. Numbers, dates, steps, named constraints. Vague reassurance summarises to nothing and gets skipped.
- Structure a parser can follow. Real headings in order, short paragraphs, tables for comparisons, lists for sequences.
- Something only you can say. First-hand data, measurements, a method you actually ran. Rewritten consensus is already available from a hundred sources, and there's no reason to cite yours.
- Being crawlable. Content behind JavaScript, logins or interstitials may simply not be read.
- Conventional credibility. Being linked, being referenced, being an established source on the topic still matters.
For a worked version of this in one industry, see how it applies to legal practice pages. Notice that this is the same list that makes a page useful to a human in a hurry. There is no separate AI optimisation discipline here — there's writing clearly, structuring properly, and having something specific to say. Anyone selling "AEO" as a distinct service should be asked what they'd do that good content strategy wouldn't.
Structured data helps at the margin: FAQ and How-To markup where it genuinely matches the content, product and organisation schema, clean article markup. It's cheap, it's honest, and it makes machine extraction easier. It is not a shortcut past having something worth extracting.
The crawler decision
This is the part most sites decide by accident, usually by copying a robots.txt snippet from a blog post.
The important distinction: blocking training and blocking citation are different decisions, and the same file controls both.
Some agents crawl to build training corpora. Others fetch pages in order to answer a specific user's question right now, and those are the ones that produce citations and referral traffic. Block the second group and you have removed yourself from the answers — which is almost never what someone intends when they add a blanket block.
A defensible default for most businesses that want to be found: allow the agents that fetch to answer questions, and make your own decision about the training crawlers. If your content is your product — you sell research, courses, or a publication — that calculus differs, and blocking training while allowing citation is a coherent position.
Two practical cautions. Agent names change and new ones appear, so a robots.txt written a year ago is probably out of date; check the operators' current documentation rather than trusting a listicle. And robots.txt is a request, not an enforcement mechanism — well-behaved crawlers honour it, and it is not a security control.
A worked example of restructuring a page for extraction
Abstract advice is easy to nod at, so here's the concrete version.
Take a services page that currently opens like this: "With over fifteen years of combined experience, our team is passionate about delivering bespoke solutions tailored to the unique needs of each and every client."
That sentence contains no extractable fact. There is nothing in it a machine could quote in answer to a question, and nothing a human could use to decide anything. It is also, in some form, the opening of most service pages on the internet.
The restructured version asks the question the visitor arrived with, and answers it:
"How long does a kitchen installation take?" Most kitchen installations take eight to twelve working days once the units arrive. Removal and preparation is two days, first fix is three, and the worktop template adds a week between fitting and completion because the stone is cut off-site.
That's quotable, checkable, and specific. It also answers the actual question, which is why it works for both audiences. Nothing about it is a trick, and it took no longer to write — it just required knowing the answer and being willing to commit to it.
Three patterns that convert vague pages into extractable ones:
Turn claims into numbers. "Fast turnaround" becomes "most orders ship within two working days." If you can't put a number on it, you may not know it, and that's worth discovering.
Turn features into consequences. "Built with modern technology" becomes "your team can publish a page without a developer."
Turn generalities into conditions. "Suitable for all businesses" becomes "works well for teams of five to fifty; below five, the setup cost usually isn't worth it." Saying who something isn't for is among the most credible things a page can do, and it's rare enough that it stands out to readers and to models.
What not to do
The category is new enough to be full of confident bad advice. Things we'd avoid:
Don't write pages for machines. Text stuffed with question-shaped headings and thin answers reads badly to humans and isn't more extractable — it's less, because there's nothing worth extracting. The structure helps only when the substance is there.
Don't buy an "AEO audit" that's a renamed SEO audit. Ask specifically what would be done differently from good content and technical work. If the answer is schema markup and FAQ sections, that's a normal SEO engagement with a new label on the invoice.
Don't block every crawler in a panic. It's reversible, but while it's in place you're absent from a channel your competitors are in, and the effect is delayed enough that you may not notice what it cost.
Don't chase citation for its own sake. Being quoted in an answer about a topic you don't monetise is a vanity metric. Prioritise the questions that precede a purchase.
Don't trust a single dashboard number. Given the classification instability above, treat any AI-channel figure as provisional until you've checked how it's defined.
Don't rebuild your site for this. Nothing in this article requires a new platform, a new design, or a migration. It requires clearer writing and a measurement setup, and any proposal that turns it into a rebuild is using a trend as a sales opportunity.
How much should you actually care
Honestly: less than the volume of commentary suggests, and more than zero.
For most businesses today this channel is small compared to search — and it's growing, it's poorly measured, and the influenced-but-unattributed portion is invisible, which means the visible number understates it by an unknown amount. That combination makes it easy to over- and under-invest at the same time.
The proportionate response:
- Set up the measurement. An afternoon, and it stops you guessing.
- Start the citation log. Half an hour a month.
- Fix the crawler decision deliberately rather than by accident.
- Write the questions-and-answers content you should be writing anyway.
- Don't restructure your marketing around it yet, and be sceptical of anyone who tells you to.
The reason to do the first four now isn't that this channel is currently large. It's that all four are cheap, all four improve your site for humans regardless, and the measurement takes months to become meaningful — so starting late means being unable to answer the question later, when it matters more.
Which pages get cited, and why it's rarely the homepage
One consistent pattern worth planning around: when an assistant sends someone to a site, it is almost never to the front door.
Homepages are written to introduce a company to everyone, which means they answer no specific question completely. The pages that get cited are the ones that resolve one thing thoroughly — a comparison, a process, a definition, a set of numbers, a decision framework. That has two practical consequences.
First, your landing page report is your citation report. If assistant referrals concentrate on three pages, those three pages are what you're known for in this channel, whether or not they're what you sell. That's worth knowing, because the gap between "what we're cited for" and "what we'd like to be cited for" is a content plan.
Second, deep pages need to work as entry points. A page that assumes the reader arrived via your navigation will disorient someone who landed on it cold from an answer. Every substantive page needs to establish, near the top, what this is and who you are, and offer an obvious route onward — which is the same internal linking discipline that helps everything else.
There's a third-order effect worth mentioning. Because assistants summarise, a visitor who does click has often already absorbed your general position and arrives with a narrower, more advanced question. Pages that only restate the basics can feel redundant to that visitor. The pages that convert them tend to be the ones with something the summary couldn't carry: the specifics, the caveats, the "here's when this doesn't apply."
That's a slightly uncomfortable thought for anyone whose content strategy is built on covering the fundamentals of their category. If a machine can now summarise the fundamentals adequately, the value of publishing your own version of them falls, and the value of publishing what only you know rises. We'd treat that as the most important strategic implication of this whole subject — more important than any measurement setup — and it argues for fewer, more specific pages rather than more general ones.
What we'd do this quarter
If you want a concrete plan: build the custom channel group, tag the landing pages assistants actually send people to, run the ten-question citation log once, decide the crawler policy on purpose, and take your five best-performing pages and restructure them into question-headed sections with direct answers underneath.
That last item is the one that does double duty. It makes the pages easier to skim for the person who's in a hurry, easier to extract for the machine, and it's the same work whether or not the AI channel turns into anything.
If you want a second opinion on your setup, or on whether the AI traffic in your dashboard is real growth or a relabelled channel, send us a screenshot — the reclassification question takes about ten minutes to settle.
Keep reading
- The SEO strategy that ranks: topic clusters for 2026
- The 1,000-page site that barely links to itself
Work with Rough Works
Need a partner for your next site? We're a Vancouver-based digital agency building websites and AI-enhanced experiences for brands across Canada and the United States. Start a project →
Common questions
How do I track ChatGPT and other AI assistant traffic in GA4?
Build a custom channel group that matches the raw source value against assistant hostnames — chatgpt.com, chat.openai.com, perplexity.ai, claude.ai, gemini.google.com, copilot.microsoft.com and others — rather than relying on the default grouping. Match on raw source rather than the platform's own channel label, because those labels change. Then segment by landing page, which tells you what you are being cited for, and compare engagement and conversion rather than only sessions, since assistant referrals tend to arrive better informed than search traffic.
Why did my AI traffic suddenly appear in mid-2026?
Possibly because the classification changed rather than the traffic. In our own analytics property, the same assistant hosts were labelled as referral traffic until around June 2026, briefly arrived with no medium set, and then began appearing under an ai-assistant channel. Any period-over-period comparison spanning that change is partly measuring the relabelling. Check how your channel is defined and whether the underlying source values actually changed before treating a jump as growth.
Should I block AI crawlers in robots.txt?
Only after separating two different decisions that the same file controls. Some agents crawl to build training corpora; others fetch a page in order to answer a specific user's question, and those are the ones that generate citations and referral traffic. A blanket block removes you from the answers, which is rarely the intent. A defensible default for most businesses that want to be found is to allow the agents that fetch to answer questions and make a separate, deliberate decision about training crawlers — the calculus differs if your content is itself the product.
What is AEO and is it different from SEO?
Answer engine optimisation describes making content easy for AI assistants to extract and cite. In practice the work is not a separate discipline: question-shaped headings with direct answers underneath, checkable specifics rather than vague claims, clean structure, crawlability, and something only you can say. That is the same list that makes a page useful to a person in a hurry. If an agency proposes an AEO engagement, ask specifically what they would do differently from good content and technical SEO work — if the answer is schema markup and FAQ sections, it is a normal SEO engagement with a new label.
How can I tell if an AI assistant is citing my website?
Analytics only shows people who clicked, so keep a manual citation log. Once a month, take the ten questions your customers actually ask, put them to the main assistants, and record whether you appear, which page is cited, and who is cited instead. It takes about half an hour across three assistants and is the only direct evidence available. It also reveals who your real competitors are in this channel, which is often a trade publication or a directory rather than the businesses you expect.
Does structured data help with AI search?
It helps at the margin. FAQ, How-To, product, organisation and article markup make machine extraction easier and are cheap to add, provided the markup genuinely matches visible content on the page. It is not a shortcut past having something worth extracting — structure only helps when there is substance underneath it, and marked-up thin answers are less citable, not more.
How much should a small business invest in AI search right now?
Proportionately little, but not nothing. Set up the measurement, start a monthly citation log, make the crawler decision deliberately, and write the question-and-answer content you should be writing anyway. All four are cheap and all four improve the site for human readers regardless of what happens to the channel. What we would not do yet is restructure a marketing strategy around it, or commission a rebuild in its name — nothing about this work requires a new platform.

