Understand
Measuring AI referrals and references
Referrals come from stored page views. Crawler rows come from the User-Agent on the request. Citations stay empty until someone logs one.
AI referral traffic
A stored page view counts when its referrer host or utm_source is on the list below. 222 page views were scanned. 0 matched an assistant. A zero here is the stored count.
- ChatGPT
- Perplexity
- Gemini
- Copilot
- Claude
- You.com
- Poe
- Meta AI
- Phind
No stored page view has an AI referrer or an AI campaign source.
Days
No day has an AI referral in this scan.
Landing pages
No landing page has an AI referral in this scan.
GA4 channel group
In GA4, open Admin, Data display, Channel groups. Copy the default group and add a channel named AI referral above Referral. Set the source match to regex, case insensitive. Leave Organic Search alone: this pattern does not match google or bing, so ordinary search stays in Organic Search. A Bing chat URL and a Bing referral share the source bing in GA4, because GA4 does not keep the referrer path. This site’s classifier reads that path.
^(chatgpt\.com|chat\.openai\.com|chatgpt|openai|perplexity\.ai|perplexity|gemini\.google\.com|gemini|copilot\.microsoft\.com|copilot|claude\.ai|claude|you\.com|you|poe\.com|poe|meta\.ai|phind\.com|phind)$
Dark traffic
Many AI apps open a link with an empty referrer. The hit is stored as direct. A campaign parameter is the reliable mark, and only when the answer included your tagged URL.
Teams sometimes treat a rise in direct visits to pages people rarely type as a stand-in for those stripped clicks. That comparison needs a baseline this database does not have, so this page does not turn the direct count into an AI estimate.
28 stored page views are direct and landed somewhere other than the home page. That is the pool a stripped AI click would sit in.
Try a referral
Each button runs the same classifier with a referrer or a utm_source. Nothing is written to the visit log. When consent is on, a match sends ai_referral_detected with the assistant id only.
AI crawler visits
The request proxy, the file Next.js 16 uses in place of middleware, reads the User-Agent. A match on a document request is posted to the crawler log. Training crawlers build a corpus. Search crawlers fill an answer index. User-fetch crawlers run because a person asked the assistant to open the URL. Google-Extended and Applebot-Extended are robots.txt tokens, not User-Agents, so they never appear as rows. Postgres migration drizzle/0008_ai_crawler.sql creates ai_crawler_hit and ai_citation.
4 recent document hits are stored.
| Token | Kind | Hits | What it is |
|---|---|---|---|
| GPTBot | Training | 0 | OpenAI training crawl. |
| OAI-SearchBot | Search | 0 | OpenAI’s search index. |
| ChatGPT-User | User fetch | 0 | A person asked ChatGPT to open the page. |
| PerplexityBot | Search | 0 | Perplexity’s index. |
| Perplexity-User | User fetch | 0 | A person asked Perplexity to open the page. |
| ClaudeBot | Training | 4 | Anthropic training crawl. |
| Claude-User | User fetch | 0 | A person asked Claude to open the page. |
| Google-Extended | robots.txt token | — | Not a User-Agent. A robots.txt token. Googlebot fetches the page. Disallowing Google-Extended opts Gemini training out and leaves Google search crawling in place when Googlebot is still allowed. |
| Bingbot | Search | 0 | Microsoft’s search crawl. Copilot grounding often uses the same token, so a hit is not labeled Copilot from the User-Agent alone. |
| Applebot-Extended | robots.txt token | — | Not a separate User-Agent. A robots.txt token for Apple Intelligence training. The fetcher is Applebot. |
| CCBot | Training | 0 | Common Crawl, often reused in training sets. |
| Bytespider | Training | 0 | ByteDance crawl. |
| Amazonbot | Search | 0 | Amazon crawl. |
| meta-externalagent | Training | 0 | Meta AI training crawl. |
Recent fetches
- ClaudeBot · Training/lab/seo
- ClaudeBot · Training/robots.txt
- ClaudeBot · Training/sitemap.xml
- ClaudeBot · Training/lab/ask
robots.txt
The file allows / for every user agent and disallows account, admin, API, and the white-paper download and thanks paths. There is no GPTBot, ClaudeBot, or Google-Extended group. Training and search crawlers that obey robots.txt may fetch the public pages. This page does not change that policy.
User-Agent: * Allow: / Disallow: /account Disallow: /admin Disallow: /api Disallow: /white-paper/download Disallow: /white-paper/thanks Sitemap: https://www.blackboxpersonalization.com/sitemap.xml
Citation log
Brand mentions in an answer are not in the server log. The way to measure them is to ask a fixed prompt on a schedule and record whether the answer cited you. That record is manual. Add a row from Admin. The fields are the prompt, the assistant, the date, whether a URL was cited, and that URL.
No citation has been logged.
llms.txt
llmstxt.org is Jeremy Howard’s 2024 proposal for a Markdown file at /llms.txt. It gives a model a short map of the site: a title, a summary, and a few links.
It does not block a crawler. That is robots.txt. It is not the full URL list. That is the sitemap. Publishing the file is not a confirmed ranking signal.
/llms.txt names the site and points at the lab, the notes, and the white paper. What is llms.txt? walks the file line by line.
# Black Box Personalization > Opening the black box of personalization. A public site by Neil Black about how digital metrics are collected in the browser, on the network, and at the server, and how those signals can personalize a page. The public origin is https://www.blackboxpersonalization.com. Pages teach with the live mechanism. This file is a curated brief. It does not replace robots.txt or the sitemap. ## Start here - [Home](https://www.blackboxpersonalization.com/): Where a metric is read in the browser, on the network, and in the stored beacon. - [Lab](https://www.blackboxpersonalization.com/lab): Collection, understanding, and personalization, with the rule on the page. - [About](https://www.blackboxpersonalization.com/about): Neil Black. - [Notes](https://www.blackboxpersonalization.com/blog): Short essays. - [White paper](https://www.blackboxpersonalization.com/white-paper): How the site was built. ## Measurement - [AI referrals](https://www.blackboxpersonalization.com/lab/ai-referrals): How this site classifies AI referrers and crawler user agents. - [llms.txt](https://www.blackboxpersonalization.com/lab/llms-txt): What this file is, and how it differs from robots.txt and the sitemap. - [SEO](https://www.blackboxpersonalization.com/lab/seo): What a crawler reads in the HTML. - [Sitemap](https://www.blackboxpersonalization.com/sitemap.xml): The full public URL list.
Recommended next
Why these picksReading path transitions…