Original research · Karbon Agency · published Sep 16, 2026 · updated Oct 2, 2026
AI Crawlers and AI Referrals: What Our Own Logs Show, 2026
Second edition, data through Oct 1, 2026. Two first-party datasets, both ours. 111,049 tracked sessions across 14 websites (May 24 – Oct 1, 2026), and 17,225 server-side crawler fetches from our own domain (Sep 14 – Oct 1, 2026). The pixel tells us who arrives; the server log tells us who reads. They disagree by two orders of magnitude, and that gap is the story.
Are AI assistants sending real traffic to local-business websites yet?
Very little — but they are reading more every week. AI-assistant crawlers fetched 7,811 pages vs 2,071 for Google, Bing and Apple combined (Sep 14 – Oct 1, 2026, n=17,225), and OAI-SearchBot alone rose 57% week over week. Yet only 125 of 21,043 human sessions (0.59%) came from an assistant (Jun 1 – Oct 1, 2026).
Every figure on this page is printed with the sample size and date window it was computed over. Where a cohort is too small to support a rate, we say so instead of publishing one.
What changed since the first edition
The first edition (published Sep 16, 2026) was built on the crawler log’s first days, Sep 14 – 16, 2026. This edition covers Sep 14 – Oct 1, 2026, long enough to compare the log week against week: Sep 18 – Sep 24, 2026 against Sep 25 – Oct 1, 2026, two blocks of 7 complete days each.
- OAI-SearchBot is now the busiest crawler on the site. It went from 1,960 fetches to 3,069 (+57%) and, in the latest block, out-fetched Amazonbot (1,859). OAI-SearchBot builds the index ChatGPT search answers from, so this is the crawler to watch.
- Meta’s meta-externalagent jumped from 122 to 2,519, but as one burst. 2,501 of its 2,519 latest-block fetches came on a single day (Oct 1), the last day in the data, and Meta’s link-preview fetcher spiked the same day (66 of its 78). One day cannot tell a new crawl campaign from a one-off; the next refresh will.
- GPTBot’s burst has passed. Its 523 fetches in the earlier block were almost all one day (522 on Sep 24); the latest block had 172 (−67%). Training crawls arrive in campaigns, not as a steady rate. It was the only crawler in the table below to fall.
- Most other crawlers rose. PerplexityBot went from 222 to 391 (+76%), ChatGPT-User from 254 to 369 (+45%), Bingbot from 598 to 705 (+18%) and Googlebot from 56 to 70. ClaudeBot, silent through the earlier block, made 73 fetches in the latest one.
- Live, question-driven fetches grew. ChatGPT-User went from 62 fetches in the first edition’s window to 715 now, and Claude-User from 3 to 55. Section 2 shows which pages they asked for.
- AI referrals rose, almost entirely on our own site. In the days since the first edition’s window closed (Sep 16 – Oct 1, 2026), 56 of 3,657 human sessions (1.53%) came from an AI assistant, against 0.40% over Jun 1 – Sep 15, 2026. But 49 of those 56 landed on karbonagency.com, which is also the site doing the most AI-search work. Client sites received 7. That is a result for one site, not a market trend.
| Crawler | Sep 18 – Sep 24 | Sep 25 – Oct 1 | Change |
|---|---|---|---|
| OAI-SearchBot | 1,960 | 3,069 | +57% |
| meta-externalagent | 122 | 2,519 | +1965% |
| Amazonbot | 1,561 | 1,859 | +19% |
| Bingbot | 598 | 705 | +18% |
| PerplexityBot | 222 | 391 | +76% |
| ChatGPT-User | 254 | 369 | +45% |
| GoogleOther | 2 | 239 | +11850% |
| GPTBot | 523 | 172 | −67% |
| facebookexternalhit | 9 | 78 | +767% |
| ClaudeBot | 0 | 73 | new |
| Googlebot | 56 | 70 | +25% |
| Applebot | 34 | 41 | +21% |
| Claude-User | 16 | 34 | +113% |
Corrections to the first edition
- AI referral counts. The first edition reported 58 AI-referred sessions (50 from ChatGPT) and that 68.0% of ChatGPT visitors scrolled past halfway. The query behind those figures was not saved and we could not reproduce them. Re-running the documented query over the same window (Jun 1 – Sep 15, 2026) gives 69 AI-referred sessions (60 from ChatGPT) and 50.7% scrolling past halfway. This edition publishes only reproducible figures, and every query now ships with the report.
- Search pages per session. The first edition counted accounts.google.com sign-in redirects as search referrals, which pushed search to 1.38 pages per session. Without them the same window gives 1.28.
- Core Web Vitals. The first edition said the pixel captured no CLS or INP. It does, on the exit beacon; we had looked in the wrong event. CLS is now in section 6.
- Local development traffic (sessions on localhost) is now excluded from every session count.
1. AI crawlers now out-fetch search crawlers 3.8 to 1 on our own site
Our JavaScript pixel cannot see these visitors at all — crawlers do not execute JS. So we read the server log directly. Over Sep 14 – Oct 1, 2026 it recorded 17,225 crawler fetches against a single domain (karbonagency.com, our own site — see the methodology note on why this dataset is one site and not 14).
| Crawler | Operator | Category | Fetches | Unique paths | Fetches / path |
|---|---|---|---|---|---|
| OAI-SearchBot | OpenAI | AI assistant | 5,132 | 571 | 8.99 |
| Amazonbot | Amazon | AI training / platform | 4,350 | 1,559 | 2.79 |
| meta-externalagent | Meta | AI training / platform | 2,899 | 228 | 12.71 |
| Bingbot | Microsoft | Traditional search | 1,533 | 409 | 3.75 |
| PerplexityBot | Perplexity | AI assistant | 984 | 381 | 2.58 |
| GPTBot | OpenAI | AI assistant | 811 | 340 | 2.39 |
| ChatGPT-User | OpenAI | AI assistant | 715 | 103 | 6.94 |
| GoogleOther | Traditional search | 242 | 60 | 4.03 | |
| Googlebot | Traditional search | 208 | 147 | 1.41 | |
| ClaudeBot | Anthropic | AI assistant | 114 | 80 | 1.43 |
| facebookexternalhit | Meta | Link preview | 89 | 8 | 11.13 |
| Applebot | Apple | Traditional search | 88 | 88 | 1.00 |
| Claude-User | Anthropic | AI assistant | 55 | 27 | 2.04 |
| DuckAssistBot | DuckDuckGo | AI training / platform | 5 | 5 | 1.00 |
| Total | 17,225 | — | — | ||
Grouped by who operates them: crawlers belonging to companies whose main product is an AI assistant (OpenAI, Perplexity, Anthropic) made 7,811 fetches. The traditional search indexers — Googlebot, Bingbot, Applebot and GoogleOther — made 2,071. That is 3.8 AI fetches for every search fetch, up from 1.5 in the first edition.
The largest single crawler was an AI-assistant one: OAI-SearchBot, at 5,132 fetches, ahead of Amazonbot (4,350). Amazonbot, Meta’s meta-externalagent and DuckAssistBot (7,254 fetches between them) feed AI systems but belong to companies with other main products, so we count them in their own bucket rather than inflating the AI-assistant number with them. Either way the direction is the same: on this domain, in this window, Googlebot was the 9th-busiest crawler.
Sweeping versus re-reading
The fetches-per-path column separates two behaviours. Single-pass sweepers — Applebot (1.00), ClaudeBot (1.43), Googlebot (1.41) — fetch each URL about once in 18 days. OAI-SearchBot re-read its 571 paths 8.99 times each: it is keeping a working set of pages fresh, not discovering new ones. Only two crawlers ran a higher ratio, and neither is steady re-reading — meta-externalagent (12.71) is one day’s burst piled onto 228 paths, and facebookexternalhit (11.13) renders link previews for 8 shared URLs. ChatGPT-User (6.94 per path) looks similar for a different reason — live fetches pile onto a few pages, above all the homepage.
2. What the assistants fetch when someone asks a question
ChatGPT-User and Claude-User are not crawlers in the indexing sense. They fetch a page because a person asked the assistant something and it went to read the source. Each fetch is evidence of a real question, which makes them the closest thing in the log to demand. Across Sep 14 – Oct 1, 2026 they made 770 live fetches, against 65 in the first edition’s window.
| Page type | Live fetches | Distinct pages |
|---|---|---|
| Homepage | 490 | 1 |
| Best-of list | 83 | 24 |
| Industry page | 51 | 33 |
| Original research report | 49 | 2 |
| Benchmark page | 32 | 13 |
| Cost guide | 30 | 20 |
| Other page | 28 | 18 |
| Service page | 7 | 7 |
The homepage took 490 of the 770 — most plausibly assistants checking who we are before answering. Beyond it, 280 live fetches spread across 117 distinct pages, led by best-of lists (83), industry pages (51), original research reports (49) and benchmark pages (32) — mostly pages that answer one narrow question with a list or a number. The most-fetched individual pages were / (490), /data/ai-crawler-and-referral-report-2026 (48), /benchmarks/industry/insurance-agent (12), /best/dermatologists (11) and /best/gyms (9). This report itself became one of the most-fetched pages within days of publication.
What this cannot tell you: the assistants do not say who asked, so we cannot count distinct people. One person repeating a question, or a scheduled prompt, produces many fetches, and some may be our own team checking how assistants describe us. Read the page-type mix as the signal, not the absolute count.
3. Which assistants actually send visitors? Essentially one
Across 21,043 human sessions on 13 sites (Jun 1 – Oct 1, 2026), we classified every referrer. AI assistants accounted for 125 sessions in total — 0.59%.
| Assistant | Sessions | Share of human sessions |
|---|---|---|
| ChatGPT (chatgpt.com) | 116 | 0.55% |
| Perplexity | 3 | insufficient sample |
| Gemini | 3 | insufficient sample |
| Claude | 2 | insufficient sample |
| Microsoft Copilot | 1 | insufficient sample |
ChatGPT is the entire channel (116 of 125). Perplexity (3), Gemini (3), Claude (2), Microsoft Copilot (1) are each far below our 30-session publication floor, so we report their raw counts and refuse to compute a rate from them. Anyone publishing a tidy percentage split of “AI search market share” from referral data at this volume is reporting noise.
The cohort is also concentrated: it spans 4 sites, and one of them — karbonagency.com, our own — accounts for 69.6% (87 sessions). One caveat cuts the other way: assistants increasingly answer in place without a click, and a visitor who reads about a business in ChatGPT and then types the domain lands in our Direct bucket. The 125 sessions are a floor, not a measurement of AI’s influence.
4. AI-referred visitors scroll further, but do not stay longer
With 125 sessions the AI cohort is now above our floor for every rate in the table, so we can compare it with the other channels. Remember that 69.6% of it is one site.
| Channel | Sessions (n) | Avg pages | Multi-page | Scrolled past 50% | Median engaged (n) |
|---|---|---|---|---|---|
| AI assistant | 125 | 1.50 | 26.4% | 48.0% | 36.5s (100) |
| Search engine | 3,402 | 1.40 | 16.6% | 16.4% | 52.0s (973) |
| Social | 1,963 | 1.43 | 24.9% | 12.3% | 18.0s (1,142) |
| Direct / no referrer | 13,904 | 1.40 | 15.8% | 10.5% | 17.0s (1,963) |
48.0% of AI-referred sessions scrolled past the halfway point, against 16.4% of search sessions, and 26.4% went multi-page against 16.6%. But the median engaged time ran the other way: 36.5 seconds (n=100) against 52.0 seconds for search (n=973).
A reading that fits both: the assistant has already answered the easy questions, so the visitor who clicks through skims further down the page for the one thing they still need — and leaves once they find it. The first edition’s stronger claim (AI visitors engage far more) does not survive the larger, reproducible sample.
5. Bot share is not a stable number, and single-month benchmarks mislead
Across the full 111,049 sessions (May 24 – Oct 1, 2026), 80.8% were bots. But the monthly figure is volatile enough that the overall average describes no individual month well.
| Month | Sessions | Bot sessions | Bot share |
|---|---|---|---|
| May 2026 (partial) | 2,131 | 1,840 | 86.3% |
| Jun 2026 | 22,042 | 20,157 | 91.4% |
| Jul 2026 | 17,349 | 12,994 | 74.9% |
| Aug 2026 | 6,504 | 3,887 | 59.8% |
| Sep 2026 | 59,266 | 47,441 | 80.0% |
| Oct 2026 (partial) | 3,757 | 3,396 | 90.4% |
Bot share ranged from 59.8% in Aug to 91.4% in Jun — a 31.7-point swing on a largely unchanged set of sites. Automated traffic arrives in campaigns, and a benchmark quoted from one month’s data (including ours) can be off by thirty points. If you are comparing your bot rate to a published figure, compare distributions over months, not point estimates.
6. The pages the assistants are reading are not fast
From 17,050 performance beacons carrying an LCP value (human sessions, Jun 1 – Oct 1, 2026), the 75th-percentile Largest Contentful Paint was 2,776 ms — above the 2,500 ms “good” threshold, so the tracked set fails that bar at p75. 72.8% of individual page loads were inside it; the slow quarter drags the percentile out. TTFB was 882 ms at p75.
Layout stability is worse. The 75th-percentile Cumulative Layout Shift was 0.249 (n=10,471 page exits), above the 0.1 "good" line: a quarter of the measured page views scored 0.249 or worse.
We include this mainly to head off a wrong inference. Page speed is weak as a citation lever, and nothing in our crawler log suggests these crawlers are deterred by a slow page — they fetched the pages anyway. Fix LCP and CLS because slow, jumpy pages lose humans, not because it will win you an AI citation.
INP is shown only as an upper bound (104 ms at p75, n=5,377) — see “What the data did not support” for why.
What the data did not support
We tested more angles than we published. These are the ones we dropped, and why — listed because a study that only reports its successes is not telling you how it was made.
| Angle tested | Why it is not in this report |
|---|---|
| Which form field people abandon most | 140 abandoned-field events were recorded in the window and the largest single field cohort was 31. At most one field clears our 30-session publication floor, and one field is not a ranking. |
| Device split of AI-referred visitors | The AI-referred cohort is 125 sessions (desktop 99, mobile 25, tablet 1). Only 1 of those cells clears the 30-session floor, so there is no split to compare. |
| Whether AI assistants send visitors to deeper pages than search | 25.6% of AI-referred sessions landed on a homepage vs 27.0% of search sessions. At n=125 that gap sits inside sampling noise. We are not calling it a finding. |
| HTTP status codes returned to crawlers | The status column in our crawler log is NULL for 100% of rows in this window. The field is not being populated. |
| INP as a Core Web Vital | The pixel only records interactions of 40 ms or longer, so page views whose interactions were all fast report nothing. The p75 over what is left (104 ms, n=5,377) overstates real INP, so we give it as an upper bound only, not a pass/fail figure. |
| Time-of-day and day-of-week demand | Already covered in our State of Local Business Web Traffic 2026 report; re-running it here produced nothing new. |
Methodology
Where the numbers come from
Two independent first-party instruments, both operated by Karbon Agency.
Session data — our own tracking pixel (k.js) on 14 websites: karbonagency.com plus 13 client sites, mostly US local and small businesses. 111,049 sessions, May 24 – Oct 1, 2026. Referral, engagement and performance analyses use the non-bot subset from Jun 1 onward: 21,043 human sessions, Jun 1 – Oct 1, 2026. Sessions on localhost (our own development) are excluded.
Crawler data — server-side request logging on marketing pages, recording user agent and path for known crawler families. 17,225 fetches, Sep 14 – Oct 1, 2026. The log started on Sep 14, 2026 and runs on one domain only: karbonagency.com, our own site. It is not the 14-site session sample, and 18 days is enough to compare one week with the next, not to call a trend. Treat the crawler sections as an early read that we will keep extending.
How channels are assigned
A session’s channel comes from its referrer’s host. AI assistant: chatgpt.com, perplexity.ai, gemini.google.com, claude.ai or Copilot — or, when the referrer is empty, a utm_source naming one of them, because ChatGPT tags the links it shows with utm_source=chatgpt.com even when no referrer arrives. Search: Google, Bing, DuckDuckGo, Yahoo, Kagi and similar search hosts (sign-in and mail subdomains excluded). Social: Facebook, Instagram, Threads, TikTok, LinkedIn, X, YouTube, Pinterest, Reddit. Direct: no referrer. Sessions referred from one tracked site to another are excluded so internal navigation cannot inflate engagement.
What the pixel does not capture
k.js only sees clients that execute JavaScript. Crawlers overwhelmingly do not, which is exactly why the crawler counts come from server logs and cannot be cross-referenced against the session data — GPTBot and its peers are absent from site_sessions by construction. The two datasets are reported side by side, never combined.
Bot classification in the session data is heuristic — user-agent and behavioural signals, not a verified reverse-DNS check. It will both miss sophisticated bots and occasionally flag an unusual human. The month-to-month volatility in section 5 is partly real traffic and partly the heuristic’s own variance; we cannot fully separate the two.
Engaged time and CLS come from an exit beacon that does not always arrive — a fast bounce or a killed tab returns nothing. We therefore report them only over sessions that returned a beacon, and state that sub-sample beside every such figure rather than treating a missing beacon as zero.
Sample bias
These 14 sites are Karbon’s own and its clients’, so the sample skews to the verticals we serve — local service businesses, entertainment venues and small professional practices in the US. It is not a random sample of the web. Traffic is also unevenly distributed across the sites, so the aggregate leans toward the busier ones, and the AI-referral cohort leans heavily toward our own site.
Privacy
Every figure is a cross-client aggregate. We publish no client name, business name, individual session, visitor identifier, email, IP address, or city-level figure tied to any single site. Cohorts below 30 observations are not published as rates — where a cohort falls under that floor we print the raw count and label it “insufficient sample” rather than deriving a percentage from it. The one domain named anywhere in this report is our own.
Dates and reproducibility
First published Sep 16, 2026. This edition’s data runs through Oct 1, 2026 (UTC), and the copy was re-read against it on Oct 2, 2026. Every figure is regenerated by one committed script running the same read-only queries, so each edition re-runs identical logic. Figures are baked in at that point rather than queried live, so this page shows the numbers a human actually checked.
Want to know which crawlers are reading your site?
The same instrumentation that produced this report — pixel plus server-side crawler logging — will tell you which AI bots fetch your pages, how often, and which pages they keep coming back to. Compare against our published benchmarks.
Book a free strategy callPublished by Karbon Agency. Citing this report? Link to this page and credit “Karbon Agency, AI Crawlers and AI Referrals 2026.” Journalists and researchers can request the underlying aggregate tables via the contact page. See also our State of Local Business Web Traffic 2026 and Bot Traffic Report 2026.