Original research · Karbon Agency · published Sep 16, 2026 · updated Oct 2, 2026

AI Crawlers and AI Referrals: What Our Own Logs Show, 2026

Second edition, data through Oct 1, 2026. Two first-party datasets, both ours. 111,049 tracked sessions across 14 websites (May 24 – Oct 1, 2026), and 17,225 server-side crawler fetches from our own domain (Sep 14 – Oct 1, 2026). The pixel tells us who arrives; the server log tells us who reads. They disagree by two orders of magnitude, and that gap is the story.

Are AI assistants sending real traffic to local-business websites yet?

Very little — but they are reading more every week. AI-assistant crawlers fetched 7,811 pages vs 2,071 for Google, Bing and Apple combined (Sep 14 – Oct 1, 2026, n=17,225), and OAI-SearchBot alone rose 57% week over week. Yet only 125 of 21,043 human sessions (0.59%) came from an assistant (Jun 1 – Oct 1, 2026).

Every figure on this page is printed with the sample size and date window it was computed over. Where a cohort is too small to support a rate, we say so instead of publishing one.

3.8x
AI-assistant crawlers fetched 7,811 pages vs 2,071 for Google, Bing and Apple combined
n=9,882 · Sep 14 – Oct 1, 2026
+57%
OAI-SearchBot fetches, Sep 18 – Sep 24, 2026 vs Sep 25 – Oct 1, 2026 (1,960 → 3,069)
n=5,029 · Sep 14 – Oct 1, 2026
0.59%
of human sessions arrived from an AI assistant — 125 of 21,043
n=21,043 · Jun 1 – Oct 1, 2026
80.8%
of all tracked sessions were bots, but the monthly figure swung 59.8%–91.4%
n=111,049 · May 24 – Oct 1, 2026

What changed since the first edition

The first edition (published Sep 16, 2026) was built on the crawler log’s first days, Sep 14 – 16, 2026. This edition covers Sep 14 – Oct 1, 2026, long enough to compare the log week against week: Sep 18 – Sep 24, 2026 against Sep 25 – Oct 1, 2026, two blocks of 7 complete days each.

Crawler fetches on karbonagency.com, Sep 18 – Sep 24, 2026 vs Sep 25 – Oct 1, 2026 (complete UTC days). Crawlers with fewer than 30 fetches across both blocks are left out.
CrawlerSep 18 – Sep 24Sep 25 – Oct 1Change
OAI-SearchBot1,9603,069+57%
meta-externalagent1222,519+1965%
Amazonbot1,5611,859+19%
Bingbot598705+18%
PerplexityBot222391+76%
ChatGPT-User254369+45%
GoogleOther2239+11850%
GPTBot523172−67%
facebookexternalhit978+767%
ClaudeBot073new
Googlebot5670+25%
Applebot3441+21%
Claude-User1634+113%

Corrections to the first edition

1. AI crawlers now out-fetch search crawlers 3.8 to 1 on our own site

Our JavaScript pixel cannot see these visitors at all — crawlers do not execute JS. So we read the server log directly. Over Sep 14 – Oct 1, 2026 it recorded 17,225 crawler fetches against a single domain (karbonagency.com, our own site — see the methodology note on why this dataset is one site and not 14).

Crawler fetches by user agent. n=17,225 fetches, one domain, Sep 14 – Oct 1, 2026.
CrawlerOperatorCategoryFetchesUnique pathsFetches / path
OAI-SearchBotOpenAIAI assistant5,1325718.99
AmazonbotAmazonAI training / platform4,3501,5592.79
meta-externalagentMetaAI training / platform2,89922812.71
BingbotMicrosoftTraditional search1,5334093.75
PerplexityBotPerplexityAI assistant9843812.58
GPTBotOpenAIAI assistant8113402.39
ChatGPT-UserOpenAIAI assistant7151036.94
GoogleOtherGoogleTraditional search242604.03
GooglebotGoogleTraditional search2081471.41
ClaudeBotAnthropicAI assistant114801.43
facebookexternalhitMetaLink preview89811.13
ApplebotAppleTraditional search88881.00
Claude-UserAnthropicAI assistant55272.04
DuckAssistBotDuckDuckGoAI training / platform551.00
Total17,225——

Grouped by who operates them: crawlers belonging to companies whose main product is an AI assistant (OpenAI, Perplexity, Anthropic) made 7,811 fetches. The traditional search indexers — Googlebot, Bingbot, Applebot and GoogleOther — made 2,071. That is 3.8 AI fetches for every search fetch, up from 1.5 in the first edition.

The largest single crawler was an AI-assistant one: OAI-SearchBot, at 5,132 fetches, ahead of Amazonbot (4,350). Amazonbot, Meta’s meta-externalagent and DuckAssistBot (7,254 fetches between them) feed AI systems but belong to companies with other main products, so we count them in their own bucket rather than inflating the AI-assistant number with them. Either way the direction is the same: on this domain, in this window, Googlebot was the 9th-busiest crawler.

Sweeping versus re-reading

The fetches-per-path column separates two behaviours. Single-pass sweepers — Applebot (1.00), ClaudeBot (1.43), Googlebot (1.41) — fetch each URL about once in 18 days. OAI-SearchBot re-read its 571 paths 8.99 times each: it is keeping a working set of pages fresh, not discovering new ones. Only two crawlers ran a higher ratio, and neither is steady re-reading — meta-externalagent (12.71) is one day’s burst piled onto 228 paths, and facebookexternalhit (11.13) renders link previews for 8 shared URLs. ChatGPT-User (6.94 per path) looks similar for a different reason — live fetches pile onto a few pages, above all the homepage.

2. What the assistants fetch when someone asks a question

ChatGPT-User and Claude-User are not crawlers in the indexing sense. They fetch a page because a person asked the assistant something and it went to read the source. Each fetch is evidence of a real question, which makes them the closest thing in the log to demand. Across Sep 14 – Oct 1, 2026 they made 770 live fetches, against 65 in the first edition’s window.

Live assistant fetches (ChatGPT-User + Claude-User) by page type. n=770 fetches, karbonagency.com, Sep 14 – Oct 1, 2026.
Page typeLive fetchesDistinct pages
Homepage4901
Best-of list8324
Industry page5133
Original research report492
Benchmark page3213
Cost guide3020
Other page2818
Service page77

The homepage took 490 of the 770 — most plausibly assistants checking who we are before answering. Beyond it, 280 live fetches spread across 117 distinct pages, led by best-of lists (83), industry pages (51), original research reports (49) and benchmark pages (32) — mostly pages that answer one narrow question with a list or a number. The most-fetched individual pages were / (490), /data/ai-crawler-and-referral-report-2026 (48), /benchmarks/industry/insurance-agent (12), /best/dermatologists (11) and /best/gyms (9). This report itself became one of the most-fetched pages within days of publication.

What this cannot tell you: the assistants do not say who asked, so we cannot count distinct people. One person repeating a question, or a scheduled prompt, produces many fetches, and some may be our own team checking how assistants describe us. Read the page-type mix as the signal, not the absolute count.

3. Which assistants actually send visitors? Essentially one

Across 21,043 human sessions on 13 sites (Jun 1 – Oct 1, 2026), we classified every referrer. AI assistants accounted for 125 sessions in total — 0.59%.

Referring AI assistant, human sessions only. n=21,043 sessions, 13 sites, Jun 1 – Oct 1, 2026.
AssistantSessionsShare of human sessions
ChatGPT (chatgpt.com)1160.55%
Perplexity3insufficient sample
Gemini3insufficient sample
Claude2insufficient sample
Microsoft Copilot1insufficient sample

ChatGPT is the entire channel (116 of 125). Perplexity (3), Gemini (3), Claude (2), Microsoft Copilot (1) are each far below our 30-session publication floor, so we report their raw counts and refuse to compute a rate from them. Anyone publishing a tidy percentage split of “AI search market share” from referral data at this volume is reporting noise.

The cohort is also concentrated: it spans 4 sites, and one of them — karbonagency.com, our own — accounts for 69.6% (87 sessions). One caveat cuts the other way: assistants increasingly answer in place without a click, and a visitor who reads about a business in ChatGPT and then types the domain lands in our Direct bucket. The 125 sessions are a floor, not a measurement of AI’s influence.

4. AI-referred visitors scroll further, but do not stay longer

With 125 sessions the AI cohort is now above our floor for every rate in the table, so we can compare it with the other channels. Remember that 69.6% of it is one site.

Engagement by acquisition channel. Human sessions only, Jun 1 – Oct 1, 2026. Per-row sample sizes in the Sessions column; median engaged time is computed only over sessions that returned an engagement beacon (that sub-sample is shown in brackets).
ChannelSessions (n)Avg pagesMulti-pageScrolled past 50%Median engaged (n)
AI assistant1251.5026.4%48.0%36.5s (100)
Search engine3,4021.4016.6%16.4%52.0s (973)
Social1,9631.4324.9%12.3%18.0s (1,142)
Direct / no referrer13,9041.4015.8%10.5%17.0s (1,963)

48.0% of AI-referred sessions scrolled past the halfway point, against 16.4% of search sessions, and 26.4% went multi-page against 16.6%. But the median engaged time ran the other way: 36.5 seconds (n=100) against 52.0 seconds for search (n=973).

A reading that fits both: the assistant has already answered the easy questions, so the visitor who clicks through skims further down the page for the one thing they still need — and leaves once they find it. The first edition’s stronger claim (AI visitors engage far more) does not survive the larger, reproducible sample.

5. Bot share is not a stable number, and single-month benchmarks mislead

Across the full 111,049 sessions (May 24 – Oct 1, 2026), 80.8% were bots. But the monthly figure is volatile enough that the overall average describes no individual month well.

Bot share by month. n=111,049 sessions across 14 sites, May 24 – Oct 1, 2026. Partial months are marked.
MonthSessionsBot sessionsBot share
May 2026 (partial)2,1311,84086.3%
Jun 202622,04220,15791.4%
Jul 202617,34912,99474.9%
Aug 20266,5043,88759.8%
Sep 202659,26647,44180.0%
Oct 2026 (partial)3,7573,39690.4%

Bot share ranged from 59.8% in Aug to 91.4% in Jun — a 31.7-point swing on a largely unchanged set of sites. Automated traffic arrives in campaigns, and a benchmark quoted from one month’s data (including ours) can be off by thirty points. If you are comparing your bot rate to a published figure, compare distributions over months, not point estimates.

6. The pages the assistants are reading are not fast

From 17,050 performance beacons carrying an LCP value (human sessions, Jun 1 – Oct 1, 2026), the 75th-percentile Largest Contentful Paint was 2,776 ms — above the 2,500 ms “good” threshold, so the tracked set fails that bar at p75. 72.8% of individual page loads were inside it; the slow quarter drags the percentile out. TTFB was 882 ms at p75.

Layout stability is worse. The 75th-percentile Cumulative Layout Shift was 0.249 (n=10,471 page exits), above the 0.1 "good" line: a quarter of the measured page views scored 0.249 or worse.

We include this mainly to head off a wrong inference. Page speed is weak as a citation lever, and nothing in our crawler log suggests these crawlers are deterred by a slow page — they fetched the pages anyway. Fix LCP and CLS because slow, jumpy pages lose humans, not because it will win you an AI citation.

INP is shown only as an upper bound (104 ms at p75, n=5,377) — see “What the data did not support” for why.

What the data did not support

We tested more angles than we published. These are the ones we dropped, and why — listed because a study that only reports its successes is not telling you how it was made.

Angle testedWhy it is not in this report
Which form field people abandon most140 abandoned-field events were recorded in the window and the largest single field cohort was 31. At most one field clears our 30-session publication floor, and one field is not a ranking.
Device split of AI-referred visitorsThe AI-referred cohort is 125 sessions (desktop 99, mobile 25, tablet 1). Only 1 of those cells clears the 30-session floor, so there is no split to compare.
Whether AI assistants send visitors to deeper pages than search25.6% of AI-referred sessions landed on a homepage vs 27.0% of search sessions. At n=125 that gap sits inside sampling noise. We are not calling it a finding.
HTTP status codes returned to crawlersThe status column in our crawler log is NULL for 100% of rows in this window. The field is not being populated.
INP as a Core Web VitalThe pixel only records interactions of 40 ms or longer, so page views whose interactions were all fast report nothing. The p75 over what is left (104 ms, n=5,377) overstates real INP, so we give it as an upper bound only, not a pass/fail figure.
Time-of-day and day-of-week demandAlready covered in our State of Local Business Web Traffic 2026 report; re-running it here produced nothing new.

Methodology

Where the numbers come from

Two independent first-party instruments, both operated by Karbon Agency.

Session data — our own tracking pixel (k.js) on 14 websites: karbonagency.com plus 13 client sites, mostly US local and small businesses. 111,049 sessions, May 24 – Oct 1, 2026. Referral, engagement and performance analyses use the non-bot subset from Jun 1 onward: 21,043 human sessions, Jun 1 – Oct 1, 2026. Sessions on localhost (our own development) are excluded.

Crawler data — server-side request logging on marketing pages, recording user agent and path for known crawler families. 17,225 fetches, Sep 14 – Oct 1, 2026. The log started on Sep 14, 2026 and runs on one domain only: karbonagency.com, our own site. It is not the 14-site session sample, and 18 days is enough to compare one week with the next, not to call a trend. Treat the crawler sections as an early read that we will keep extending.

How channels are assigned

A session’s channel comes from its referrer’s host. AI assistant: chatgpt.com, perplexity.ai, gemini.google.com, claude.ai or Copilot — or, when the referrer is empty, a utm_source naming one of them, because ChatGPT tags the links it shows with utm_source=chatgpt.com even when no referrer arrives. Search: Google, Bing, DuckDuckGo, Yahoo, Kagi and similar search hosts (sign-in and mail subdomains excluded). Social: Facebook, Instagram, Threads, TikTok, LinkedIn, X, YouTube, Pinterest, Reddit. Direct: no referrer. Sessions referred from one tracked site to another are excluded so internal navigation cannot inflate engagement.

What the pixel does not capture

k.js only sees clients that execute JavaScript. Crawlers overwhelmingly do not, which is exactly why the crawler counts come from server logs and cannot be cross-referenced against the session data — GPTBot and its peers are absent from site_sessions by construction. The two datasets are reported side by side, never combined.

Bot classification in the session data is heuristic — user-agent and behavioural signals, not a verified reverse-DNS check. It will both miss sophisticated bots and occasionally flag an unusual human. The month-to-month volatility in section 5 is partly real traffic and partly the heuristic’s own variance; we cannot fully separate the two.

Engaged time and CLS come from an exit beacon that does not always arrive — a fast bounce or a killed tab returns nothing. We therefore report them only over sessions that returned a beacon, and state that sub-sample beside every such figure rather than treating a missing beacon as zero.

Sample bias

These 14 sites are Karbon’s own and its clients’, so the sample skews to the verticals we serve — local service businesses, entertainment venues and small professional practices in the US. It is not a random sample of the web. Traffic is also unevenly distributed across the sites, so the aggregate leans toward the busier ones, and the AI-referral cohort leans heavily toward our own site.

Privacy

Every figure is a cross-client aggregate. We publish no client name, business name, individual session, visitor identifier, email, IP address, or city-level figure tied to any single site. Cohorts below 30 observations are not published as rates — where a cohort falls under that floor we print the raw count and label it “insufficient sample” rather than deriving a percentage from it. The one domain named anywhere in this report is our own.

Dates and reproducibility

First published Sep 16, 2026. This edition’s data runs through Oct 1, 2026 (UTC), and the copy was re-read against it on Oct 2, 2026. Every figure is regenerated by one committed script running the same read-only queries, so each edition re-runs identical logic. Figures are baked in at that point rather than queried live, so this page shows the numbers a human actually checked.

Want to know which crawlers are reading your site?

The same instrumentation that produced this report — pixel plus server-side crawler logging — will tell you which AI bots fetch your pages, how often, and which pages they keep coming back to. Compare against our published benchmarks.

Book a free strategy call

Published by Karbon Agency. Citing this report? Link to this page and credit “Karbon Agency, AI Crawlers and AI Referrals 2026.” Journalists and researchers can request the underlying aggregate tables via the contact page. See also our State of Local Business Web Traffic 2026 and Bot Traffic Report 2026.