Most of the argument about AI crawlers is about whether to let them in. Far less of it is about what they actually do once they arrive: how often they come, who sends the most, and what a site gets back for the pages it hands over.
The numbers below come from server-side logs of 924 websites running the LovedByAI WordPress plugin, which records every page request that reaches the site from a recognised bot. Nearly three in four are small and mid-sized businesses: shops, clinics, law and accounting firms, trades, restaurants, hotels and agencies. The median site has about 100 pages. Together they logged 38.6 million identified bot requests in the 30 days from 6 September to 5 October 2026, 24.3 million of them from AI crawlers. Every figure is tested for the obvious ways it could be wrong: a few huge sites, part-month data and fake bots. Every outside figure carries its source and date. We update this page every quarter.
AI crawler statistics 2026: the key numbers
- On the average small business website, AI crawlers made 40.7% of identified bot page requests, level with every search engine combined at 40.8% (6 Sep to 5 Oct 2026).
- The median site received 3.1 AI crawler requests for every Googlebot request (6 Sep to 5 Oct 2026; 95% confidence interval 2.8 to 3.4).
- 79% of sites got more requests from AI crawlers than from Googlebot (6 Sep to 5 Oct 2026).
- The median site received about 1,050 AI crawler page requests between 6 September and 5 October 2026, roughly 35 a day, from 11 different AI crawlers.
- Meta-ExternalAgent is the single most active AI crawler, at 21.6% of the average site's AI crawler requests (6 Sep to 5 Oct 2026).
- Counting each company's crawlers together, OpenAI took the largest share on the average site from 6 September to 5 October 2026 (27.2%), ahead of Meta (24.6%) and Anthropic (21.5%).
- On the average site, 69.9% of AI crawler requests were for model training, 18.8% for AI search and 11.3% were triggered by a person asking an assistant a question (6 Sep to 5 Oct 2026).
- The busiest 1% of sites received 75% to 87% of the requests from GPTBot, ClaudeBot, Meta-ExternalAgent and Amazonbot. For Googlebot the figure is 42% (6 Sep to 5 Oct 2026).
- On the same sites every month from June to September 2026, AI crawling moved with Googlebot: the ratio between them held between 2.2 and 2.4 to 1.
- OpenAI's crawlers made 33.7 page requests for every visitor ChatGPT sent, on the median site (6 Sep to 5 Oct 2026).
- 85% of restaurant, hotel and travel sites got at least one visitor from ChatGPT between 6 September and 5 October 2026, against 53% of marketing agencies.
- In our blocking study of 27 to 30 August 2026, 43% of websites refused at least one AI crawler, but only 7% said so in robots.txt.
- AI bots requested an llms.txt file 31 times between December 2025 and May 2026, across 82 million crawler requests.
How much small business website traffic comes from AI crawlers?
On the average small business website, AI crawlers made 40.7% of page requests from identified bots between 6 September and 5 October 2026, level with every search engine combined at 40.8%. Training crawlers alone made 30.8%, and Googlebot alone 16.8%.

You will see bigger numbers quoted, and our own data can produce one. Pooled across every request in the same window, AI crawlers made 63% of identified bot traffic. That figure is carried by a handful of very large sites that AI crawlers hammer: remove the busiest 1% of sites and the pooled share falls to 44%. The average-site figure weights every site equally, so it describes what a typical owner actually sees.
From 6 September to 5 October 2026, the median site's AI share was 39.6%, and on 32.8% of sites AI crawlers made the majority of all identified bot requests.
These are bot requests, not all traffic. For the share of all web traffic that is automated, the standard external figure is the Imperva 2025 Bad Bot Report, which put bots at 51% of all web traffic in 2024, the first year in a decade that automated traffic outnumbered humans.
AI crawlers vs Googlebot on small business websites: how many AI bot requests for every Google crawl?
The median small business website received 3.1 page requests from AI crawlers for every request from Googlebot between 6 September and 5 October 2026, with a 95% confidence interval of 2.8 to 3.4. 79% of sites got more requests from AI crawlers than from Googlebot, and against every traditional search engine combined, AI crawlers were level on the median site.

| Measure, 6 Sep to 5 Oct 2026 | Median small business website |
|---|---|
| AI crawler page requests | about 1,050 (35 a day) |
| Googlebot page requests | 345 (11 a day) |
| AI crawler requests per Googlebot request | 3.1 |
| AI crawler requests per request from all search engines | 1.0 |
| Sites with more AI crawler than Googlebot requests | 79% |
| Sites with more AI crawler requests than all search engines | 50% |
Why AI crawlers are 3 to 1 against Googlebot but level with search engines. Googlebot is only part of search engine crawling. On the average small business website between 6 September and 5 October 2026, Googlebot made 16.8% of bot requests, Bingbot 13.0% and Applebot, YandexBot and the other search engines 11.0%. AI crawlers, at 40.7%, match all of them together and far outnumber Googlebot alone. One thing these logs cannot show: Google has no separate AI crawler. AI Overviews and Gemini draw on the pages Googlebot already fetches, and Google-Extended is only a robots.txt switch, so Google's AI use is counted here as Googlebot. Divide those averages and AI crawlers come out at 2.4 times Googlebot; the 3.1 is the ratio on the middle site. Both put AI crawlers at two to three times Googlebot.
The lead is not a quirk of one kind of site. It holds on small, medium and large sites and on every domain group. Online stores are where Google crawls hardest, and AI crawlers still lead there, at 2.3 to 1. The full breakdown is in how we tested these numbers.
The comparison to watch is against older data. In December 2024, Vercel measured GPTBot and Claude together at about 20% of Googlebot's request volume across its network. The populations differ, so the two figures are not a like-for-like trend, but the direction is not subtle.
Which AI crawlers visit websites the most?
Meta-ExternalAgent is the single most active AI crawler on small business websites, at 21.6% of the average site's AI crawler requests from 6 September to 5 October 2026, followed by ClaudeBot at 18.5% and GPTBot at 12.2%. By reach, ChatGPT-User and OAI-SearchBot visited the most sites. "Sites reached" is the share of sites a crawler visited at least once in those 30 days.
| Crawler, 6 Sep to 5 Oct 2026 | Company | Purpose | Share of the average site's AI requests | Sites reached |
|---|---|---|---|---|
| Meta-ExternalAgent | Meta | Training | 21.6% | 77.1% |
| ClaudeBot | Anthropic | Training | 18.5% | 82.6% |
| GPTBot | OpenAI | Training | 12.2% | 83.5% |
| Amazonbot | Amazon | Training | 10.9% | 71.3% |
| OAI-SearchBot | OpenAI | Search | 7.8% | 90.6% |
| ChatGPT-User | OpenAI | User-triggered | 7.2% | 91.7% |
| PerplexityBot | Perplexity | Search | 4.9% | 73.9% |
| Bytespider | ByteDance | Training | 3.1% | 46.9% |
| Meta-WebIndexer | Meta | Search | 3.0% | 54.3% |
| CCBot | Common Crawl | Training | 2.3% | 68.2% |
| YouBot | You.com | Search | 1.9% | 51.4% |
| Claude-User | Anthropic | User-triggered | 1.8% | 57.9% |
| Claude-SearchBot | Anthropic | Search | 1.2% | 26.4% |
| DuckAssistBot | DuckDuckGo | User-triggered | 0.7% | 48.2% |
| DeepSeekBot | DeepSeek | Training | 0.6% | 38.6% |
| cohere-ai | Cohere | Training | 0.6% | 51.2% |
| MistralAI-User | Mistral | User-triggered | 0.4% | 33.4% |
| Googlebot, for reference | Search engine | n/a | 100% |
Grouped by company, the picture for the same 30 days is close to a three-way split. OpenAI's three crawlers take 27.2% of the average site's AI crawler requests, Meta's two take 24.6% and Anthropic's three take 21.5%.

The two columns disagree at the top, and both are right. OpenAI crawls almost every site steadily, so it has the biggest average share. Meta crawls some sites far harder than anyone else, so it is the single biggest AI crawler on more sites: 32.2% of them, against 28.2% for OpenAI.
Cloudflare's network data agrees on the order. In the seven days to 6 October 2026, Cloudflare Radar ranked Meta-ExternalAgent the most active AI crawler at 17.8% of AI bot and crawler traffic, then ClaudeBot at 11.4%, behind only Googlebot at 23.7%.
How often do AI crawlers visit a small business website?
The median small business website received about 1,050 AI crawler page requests between 6 September and 5 October 2026, roughly 35 a day, from 11 different AI crawlers. Googlebot made 345 requests to the same median site, and 99.6% of sites were visited by at least one AI crawler.

| AI crawler requests per site, 6 Sep to 5 Oct 2026 | Page requests |
|---|---|
| Quietest quarter of sites (25th percentile) | 354 |
| Median site | about 1,050 |
| Busiest quarter of sites (75th percentile) | 3,827 |
| Busiest tenth of sites (90th percentile) | 14,650 |
Individual crawlers, on the sites they visited at all:
| Crawler | Median page requests per site, 6 Sep to 5 Oct 2026 |
|---|---|
| Googlebot (reference) | 345 |
| Meta-ExternalAgent | 213 |
| ClaudeBot | 181 |
| GPTBot | 91 |
| Amazonbot | 87 |
| OAI-SearchBot | 58 |
| PerplexityBot | 57 |
| ChatGPT-User | 45 |
ChatGPT-User is the one to read closely. It only fetches a page when a person asks ChatGPT something that needs it, and it reached 91.7% of sites between 6 September and 5 October 2026. Its timing gives it away as human demand: on Wednesday 30 September 2026 its busiest hour carried ten times the traffic of its quietest, while GPTBot, ClaudeBot and Googlebot ran within a factor of two around the clock.
Sites are found fast. For sites that installed the plugin between March and June 2026, the median time to the first logged visit from an AI search or answer crawler was zero days, and 95% were visited within a week. Being crawled is not being cited: this says nothing about how soon a site starts appearing in answers.
Why AI bots crawl: training vs search vs user requests
On the average small business website, 69.9% of AI crawler requests between 6 September and 5 October 2026 were for model training, 18.8% were for AI search indexes and 11.3% were fetches triggered by a person asking an assistant a question.
Cloudflare's August 2025 analysis found 80% of AI crawling from July 2024 to July 2025 was for training, 18% for search and 2% for user actions. The training and search shares line up closely. The user-triggered share is far higher on small sites, because a ChatGPT-User fetch is a much bigger fraction of a small site's AI traffic than of a large one's.
Purpose follows each operator's own documentation. Meta-WebIndexer is counted as search because Meta describes it as improving Meta AI search results; Meta-ExternalAgent is counted as training.
Meta-ExternalAgent: is Meta's AI crawler the most active bot?
Meta-ExternalAgent is the most active AI crawler on small business websites, at 21.6% of the average site's AI crawler requests and 43% of all AI crawler requests pooled, from 6 September to 5 October 2026. Meta documents it as crawling "for use cases such as training foundation AI models".
Most sites never see the extreme. On the median site Meta-ExternalAgent visited, it made 213 page requests in those 30 days. The pooled volume comes from a small number of sites it crawls relentlessly: the single most-crawled site in our logs received 3.66 million requests from Meta-ExternalAgent in the same 30 days, about 85 a minute, around the clock.
That pattern holds for every major training crawler.

From 6 September to 5 October 2026, the busiest 1% of sites received between 75% and 87% of the requests from each of the four biggest training crawlers. Googlebot is concentrated too, but at 42% it spreads its crawling far more evenly. For anyone reporting on AI crawler load, this is the caveat that matters: pooled totals describe the sites being hammered, and medians describe everyone else.
AI crawlers by industry: which businesses does ChatGPT send visitors to?
Restaurant, hotel, travel and events sites were the most likely to get a visitor from ChatGPT: 85% received at least one in the 30 days to 5 October 2026, against 53% of marketing and web agencies. AI crawlers out-requested Googlebot in every industry we could measure, by between 2.2 and 4.2 to 1.

The split fits the kind of question that ends in a click: where to eat, where to stay, what to buy, which clinic to call. Agencies and professional firms are crawled as hard as anyone, but a ChatGPT answer is far less often the thing that sends someone to them.
The gap between the top and the bottom of the chart is solid. It holds with Israeli sites excluded (82% against 50%), with only sites of 30 to 300 pages compared (84% against 54%), and with only the sites we classified confidently (89% against 56%). The industries in between overlap, so treat their order as approximate.
| Industry | AI crawler requests per Googlebot request, median site, 6 Sep to 5 Oct 2026 |
|---|---|
| Restaurants, hotels, travel and events | 2.2 |
| Online stores | 2.5 |
| Health, medical and wellness | 2.3 |
| Home services and trades | 3.1 |
| Education and training | 2.9 |
| Software and IT | 3.8 |
| Professional services | 3.6 |
| Publishers and blogs | 4.2 |
| Marketing and web agencies | 4.1 |
These ratios are not precise enough to rank. Most of their confidence intervals overlap, and the order moves when Israeli sites are excluded. What every row shows is the same thing as the headline: AI crawlers lead Googlebot in every industry.
Industry comes from a language model reading each site's name, title and meta description and assigning one of 12 groups. Where a site had already told us its industry, the confident classifications agreed with it 78% of the time, and several of the disagreements were wrong self-reports. Each industry shown has between 47 and 116 sites. Industries with fewer than 40 sites, sites we could not classify, and a cluster of gambling affiliate sites are left out.
Is AI crawler traffic growing? The AI bot traffic trend, June to September 2026
On the same 477 websites every month, AI crawling rose and fell with Googlebot between June and September 2026. The ratio of AI crawler to Googlebot requests held between 2.2 and 2.4 to 1 in every month.
| Index, June 2026 = 100 | Jun | Jul | Aug | Sep |
|---|---|---|---|---|
| AI crawlers | 100 | 102 | 156 | 136 |
| Googlebot | 100 | 108 | 137 | 129 |
| Bingbot | 100 | 115 | 124 | 119 |
Four months is too short, and the series too jumpy, to call a growth rate: August's jump was shared by every crawler, AI and search alike, and September gave part of it back. What the series does show is that AI crawling is not pulling away from search crawling on established sites. It has run at more than twice Googlebot's volume, at a steady ratio, for as long as we can measure it on a like-for-like basis.
This series starts in June on purpose. From 25 May 2026 our logs stopped recording robots.txt, sitemap and feed requests, which AI crawlers request heavily, so earlier months are not comparable. It also counts only the crawlers we detected for the whole period, so bots added to our detection list later cannot show up as growth.
ChatGPT crawl-to-referral ratio: how many pages does OpenAI crawl per visitor?
On the median small business website that received at least one visitor from ChatGPT between 6 September and 5 October 2026, OpenAI's crawlers made 33.7 page requests for every visitor ChatGPT sent. 632 of the 924 sites, about two in three, received at least one ChatGPT visit in those 30 days, 26,288 visits in total.

| OpenAI crawler, 6 Sep to 5 Oct 2026 | Page requests per ChatGPT visitor, median site |
|---|---|
| All three OpenAI crawlers | 33.7 (middle half of sites: 12 to 101) |
| OAI-SearchBot and ChatGPT-User (search and user fetches) | 17.1 |
| GPTBot (training) | 8.5 |
| ChatGPT-User alone | 6.8 |
Each row is its own median across sites, so the rows do not add up.
Network-wide ratios run far higher. In the seven days to 6 October 2026, Cloudflare Radar put OpenAI at 392 pages crawled for every referral, Anthropic at 473, Perplexity at 3,700, Microsoft at 41 and Google at 4.9. Pooled across all of our sites for 6 September to 5 October 2026, OpenAI's ratio was 153 to 1, the same order of magnitude as Cloudflare's. The median small site simply sits far below the pooled figure, for the same reason as everything else on this page: a few heavily crawled sites drive the totals.
We count a ChatGPT visit when it arrives tagged utm_source=chatgpt.com or with a chatgpt.com referrer. Claude and Perplexity visits arrive with an attributable referrer too rarely for a reliable 30-day ratio, so we do not publish one here. Our 108,000-page study compares all four engines, measured as visits per million crawls, over January to July 2026.
How many websites block AI crawlers?
In our study of 314 websites between 27 and 30 August 2026, 43% of websites refused at least one AI crawler, but only 7% had written that rule in robots.txt. The rest refused at the server, usually with a 403, while robots.txt allowed everything.
These figures come from our AI crawler blocking study, which fetched each site as 38 different AI crawlers.

| Crawler | Sites refusing it, robots.txt or server, 27 to 30 Aug 2026 |
|---|---|
| Bytespider | 32.2% |
| Amazonbot | 29.6% |
| ClaudeBot | 18.8% |
| GPTBot | 14.0% |
| CCBot | 11.1% |
| Meta-ExternalAgent | 10.2% |
| PerplexityBot | 8.0% |
| ChatGPT-User | 7.6% |
| OAI-SearchBot | 7.0% |
Of the 23 sites with an AI crawler rule in robots.txt, 18 had it written by Cloudflare's managed block, not by the owner. The blocking that most site owners have is blocking they did not choose and cannot see.
Do AI crawlers read llms.txt?
AI crawlers almost never request llms.txt. Across 82 million crawler requests to WordPress sites between December 2025 and 24 May 2026, AI bots fetched an llms.txt file 31 times. In May 2026 alone, the same AI crawlers made 177,353 requests for sitemap XML.

Of all 192 llms.txt fetches in that period, 161 came from search engines and SEO tools, and the single biggest requester was Googlebot, from the company that says it does not use the file. The full breakdown is in our llms.txt study.
AI bot traffic statistics from Cloudflare, Imperva and Vercel
The most cited outside figures, each checked against its original source.
| Source | Finding | Period |
|---|---|---|
| Cloudflare Radar | Most active crawlers by share of AI bot and crawler traffic: Googlebot 23.7%, Meta-ExternalAgent 17.8%, ClaudeBot 11.4%, Applebot 8.3%, Bingbot 6.7% | 7 days to 6 October 2026 |
| Cloudflare Radar | Pages crawled per referral: Perplexity 3,700, Anthropic 473, OpenAI 392, Microsoft 41, Google 4.9 | 7 days to 6 October 2026 |
| Cloudflare | 80% of AI crawling was for training, 18% for search, 2% for user actions | July 2024 to July 2025 |
| Imperva 2025 Bad Bot Report | Bots were 51% of all web traffic; malicious bots alone were 37% | 2024 |
| Vercel | GPTBot made 569 million requests in a month and Claude 370 million, together about 20% of Googlebot's 4.5 billion | The month before 17 December 2024, as published |
Is this AI crawler data reliable? How we tested it
The headline numbers for 6 September to 5 October 2026 held up under every test we ran. We looked for the usual ways a dataset like this misleads: a few giant sites, part-month data, fake bots and plain sampling error. Here is what each test showed.

It is not a few huge sites. Pooled totals are dominated by the busiest 1% of sites, so every headline figure here is a median or an average that weights each site equally. Where we quote a pooled figure, we say so.
It is not part-month data. Sites that joined during the window had only a few days of logs and looked artificially quiet. Every snapshot figure uses only sites logged on at least 25 of the 30 days.
They are real bots, not crude fakes. In every request logged on Sunday 4 October 2026, more than 99% of requests from GPTBot, ClaudeBot, Meta-ExternalAgent, Amazonbot, OAI-SearchBot and ChatGPT-User carried the exact user agent format each operator publishes. User-triggered fetches also behave like people, with a strong daily rhythm that training crawlers do not have. We cannot rule out a scraper copying a user agent string exactly, because the logs do not include IP addresses.
It is not sampling error. Resampling the sites 400 times puts the median ratio between 2.8 and 3.4, and the share of sites where AI crawlers lead between 76% and 82%.
It matches outside data. In the week to 6 October 2026, Cloudflare's network ranked the same two crawlers, Meta-ExternalAgent and ClaudeBot, as the most active AI crawlers, and our pooled OpenAI crawl-to-referral ratio sits in the same order of magnitude as Cloudflare's.
How we measured AI crawler traffic (methodology)
Data source. Server-side request logs from WordPress sites running the LovedByAI plugin. The plugin records each request that reaches WordPress from a recognised bot user agent. Nothing here comes from a third-party panel or estimator.
Window and population. 6 September to 5 October 2026 for every snapshot figure, counting only the 924 sites logged on at least 25 of the 30 days. The trend uses a fixed set of 477 sites with at least 25 logged days in each month from June to September 2026.
Average site, median site and pooled. "Average site" figures weight every site equally: each site's own shares are averaged, so they still add up to 100%. "Median site" figures are the middle site. "Pooled" figures add every request together, so the largest sites dominate them; we label them wherever they appear.
What counts as a request. Page requests only. Since 25 May 2026, requests for robots.txt, llms.txt, sitemaps, feeds, .xml, .txt, .json and .php files, and anything under /wp- (including images and the REST API) are dropped before they are stored.
Counts are floors. Requests answered by a page cache or CDN without reaching WordPress are never seen, so every count here undercounts its crawler. The reverse also applies: a request that reaches WordPress and is then refused by a security plugin is still logged, so a request here is a request that arrived, not a page that was served.
Identification. By user agent string, not verified by IP address. Tokens that are robots.txt switches with no crawler behind them (Google-Extended, Applebot-Extended), retired tokens (anthropic-ai, Claude-Web) and GoogleOther are never counted as AI. We could not match GrokBot to a crawler xAI documents, so it is left out of the ranking table.
Referral visits are not crawls. Our logs record a synthetic ChatGPT-User, Claude-User or Perplexity-User row for each AI referral visit. Those are people, not bots, so each site's referral visits are subtracted from the matching crawler before anything is counted.
Purpose. Training, search or user-triggered, following each operator's own documentation.
Who these sites are. Websites running the LovedByAI WordPress plugin. Classified by industry, 670 of the 924 (73%) are small and mid-sized business sites; the rest are publishers and blogs, non-profits, gambling affiliate sites and sites we could not classify. The median site has about 100 pages and 78% of business sites have fewer than 300. On the 670 business sites alone the headline figures do not move: 3.1 AI crawler requests per Googlebot request, AI ahead on 80% of sites, and an AI share of 40.9% against 40.1% for search engines. The owners installed a plugin for AI search, so the sample leans towards sites that care about AI visibility. About half are on .com domains and about three in ten on Israeli .il domains. Israeli sites have the lowest AI-to-Googlebot ratio of any domain group, so if anything they pull the headline figure down. This describes these WordPress sites, not the whole web.
What we do not publish. No site is identified, and no figure is tied to a named customer. Headline figures are medians and per-site shares rather than totals, because totals across a growing set of monitored sites mostly measure how many sites we monitor.
Reproducibility. Every figure comes from one set of SQL queries against the production log tables. If you want the query behind a specific number, ask us and we will send it.
How to cite these AI crawler statistics
Every first-party figure on this page is free to quote, chart or embed, with a link to this page as the source. Please keep the window with the number, because these figures move: "LovedByAI, AI Crawler Statistics 2026, data for 6 September to 5 October 2026" is enough. Journalists who need a cut we have not published can contact us.
To see which of these crawlers can actually reach your own site, run it through the free AI search checker. It fetches your homepage as each AI crawler and names any that get refused.

