Are AI Crawlers Eating Your Hosting Bandwidth? What the Data Really Shows
Bots are now the majority of traffic on the web, and you are quietly paying for them in bandwidth and server load. Whether that costs you anything real depends almost entirely on how many URLs your site has.

Here is a fact that has not sunk in yet for most business owners: for two years running, more than half of everything hitting websites has been automated. Imperva measured automated traffic at 51% of all web traffic in 2024, the first time in a decade that bots passed people, then at more than 53% in 2025.
Every one of those requests costs somebody something. Bandwidth, CPU time, database queries, an entry process on a shared hosting plan. You are not billed for it as a line item, but you are paying for it, and the bots are not shopping. Most coverage of this is either panic or vendor marketing, so below is what has actually been measured, who has actually been hurt, and what is worth doing about it.
How Much of That Bot Traffic Is Actually AI?
Less than the headlines suggest, and that matters.
Cloudflare, which handles roughly 20% of the web, published its 2025 numbers: across the year, AI bots other than Googlebot accounted for 4.2% of HTML requests on its network, from a low of 2.4% in early April to a peak of 6.4% in late June. Googlebot on its own was 4.5%, slightly more than every AI bot combined. Akamai measured AI bots at nearly 1% of total bot traffic on its platform as of November 2025, while DataDome tracked LLM crawler traffic climbing from 2.6% of verified bot traffic in January 2025 to over 10.1% by August 2025.
Those numbers disagree because they measure different networks, but they agree on the shape: AI crawling is growing fast and is still a minority of automated traffic. Within Cloudflare's verified bot traffic in 2025, Googlebot was over 28%, OpenAI's GPTBot about 7.5% and Microsoft's Bingbot about 6%. Over the year GPTBot's share of AI crawling more than doubled, from 4.7% to 11.7%, and ClaudeBot went from about 6% to about 10%.
So the honest framing is not that AI has taken over your server. A fast-growing new category of traffic has arrived, on top of everything that was already there, and nobody asked you.
The Part That Should Actually Annoy You: The Crawl-to-Refer Gap
Search crawlers have always been a fair trade. Googlebot takes your bandwidth and sends you customers. Cloudflare quantified that trade with a metric called crawl-to-refer: how many pages a company crawls for every one pageview it sends back to you.
In July 2025, Google's ratio was 5.4 to 1 and Microsoft's was 40.7 to 1. Those are recognizable business relationships.
Then there is the AI side. In January 2025 Anthropic's ratio was 286,930 to 1. By July it had improved to 38,065 to 1, an 86.7% drop, though across 2025 it peaked near 500,000 to 1. OpenAI ran 1,217 to 1 in January and 1,091 to 1 in July. Perplexity went the other way, from 54 to 1 up to 194 to 1, a 256.7% increase in crawling per referral. Mistral was the rare exception at 0.1 to 1 during a sample week in June 2025, sending ten times more referrals than it made crawl requests.
The reason for the gap is not mysterious. Cloudflare's data for the twelve months to mid-2025 found that 80% of AI crawling was for training, 18% for search and 2% for user actions. Year over year, training rose from 72% to 79% while search fell from 26% to 17%.
Four fifths of the AI crawling hitting your site is not looking for something to show a user. It is collecting material. You are supplying the inventory and paying the shipping.
Who Has Actually Been Hurt (And What They Have in Common)
The horror stories are real and well documented, and the pattern in them tells you whether you are exposed.
Wikimedia. Bandwidth for multimedia downloads grew 50% between January 2024 and April 2025, driven by scrapers rather than people. The Wikimedia Foundation reported that bots are about 35% of pageviews but at least 65% of its most expensive traffic, because bots bulk-read obscure pages that miss the edge cache and have to be served from the core datacenter.
Read the Docs. In May 2024 a single unnamed crawler downloaded 73 TB of zipped HTML, almost 10 TB of it in one day, costing over $5,000 in bandwidth. A separate Meta content downloader pulled 10 TB in June 2024. After blocking AI crawlers, file-download bandwidth fell 75%, from about 800 GB a day to about 200 GB, saving roughly $1,500 a month.
SourceHut. The git.sr.ht service logged 168 hours and 30 minutes of degradation between March 17 and March 24, 2025, attributed to LLM crawlers. The fix involved proof-of-work challenges and blocking several cloud providers outright, including Google Cloud and Azure.
Triplegangers. In January 2025 a seven-person company with more than 65,000 product pages of 3D human scans was knocked offline during US business hours by OpenAI's crawler using 600 IP addresses and tens of thousands of requests. The CEO described it to TechCrunch as basically a denial of service attack. The company had a terms-of-service ban on bots but no correctly configured robots.txt at the time.
Now look at what those four have in common. Every one has an enormous number of URLs that are cheap for a bot to request and expensive for the server to generate. A media archive. Build artifacts. Git blame and log pages. Sixty-five thousand product pages.
A brochure site for a Buffalo roofer has about forty URLs. If every AI crawler on earth pulled all forty every day for a month, that is a rounding error on any hosting plan sold this decade.

What This Means for a Normal Small Business Site
Blunt version: if you run a typical small business website, twenty to a couple hundred pages, mostly static HTML or PHP includes, images sized sensibly, no ads, on decent shared hosting, this is not your problem. Look at your own logs, confirm it, and move on.
Any vendor pointing at the "53% of traffic is bots" headline to sell you a bot-protection add-on is selling you something. That number includes every uptime monitor, every SEO tool, every scraper that has existed since 2010, and Googlebot itself.
When It Genuinely Does Affect You
Crawler load becomes a live issue if any of these is true:
- You have a product catalog with filter, sort or faceted URLs. Three filters with ten values each turns twenty products into a thousand crawlable URLs. Crawlers find all of them.
- You have an events calendar with next-month pagination. That is an infinite URL space, and it is the single most common self-inflicted crawl trap on small sites.
- Your internal search results page is crawlable. Same problem, unlimited variations.
- You run an image-heavy gallery or portfolio where every thumbnail click triggers a separate full-size fetch. Sizing and serving those correctly is worth doing anyway, and we covered the mechanics in our guide to image SEO and web design.
- You are on cheap shared hosting with CPU or entry-process limits. On those plans bandwidth is rarely the binding constraint. Concurrent PHP processes are. A burst of a few hundred simultaneous crawler requests against uncached PHP can throttle your site while your bandwidth graph looks perfectly calm.
That last one is the realistic failure mode for a small site, and it deserves naming because it does not look like a bill. It looks like the site being slow on Tuesday. Nobody connects that to crawlers, so nobody fixes it. This is the point where the quality of your hosting stops being an invisible line item, which is the argument we made at length in why website hosting matters for SEO.
Not Sure What Is Hitting Your Site?
Run a free scan of your site's technical health, then let us look at your actual server logs. Data first, decisions second. We would rather tell you that you have no problem than sell you a fix you do not need.
Free SEO Score & Grade Talk to AldoMediaThe Cloudflare Change Everyone Is About to Get Wrong
This is the time-sensitive part, and it is how a small business is most likely to get hurt by the AI crawler story in 2026. Not through a bandwidth bill. Through a checkbox.
On July 1, 2026, Cloudflare split AI traffic into three categories: Search, Agent and Training, with the controls available on all plans including Free. From September 15, 2026, Cloudflare has announced that Training and Agent bots will be blocked by default on ad-displaying pages, while Search bots stay allowed. It applies to newly onboarding domains, to new sites added by existing customers, and to free-tier customers. Owners can opt out in Security settings before that date.
Here is the side effect almost nobody is reporting. Multi-purpose crawlers are judged, in Cloudflare's own words, according to all of their behaviors. Googlebot, Applebot and Bingbot all do more than one job. So if you block Training, you block them too, even with Search set to allowed. Blocking AI training on Cloudflare can block Googlebot. That is a self-inflicted deindexing risk sitting behind a scary-sounding checkbox that a worried business owner is very likely to tick.
Most small business brochure sites do not display ads, so the new default probably does not touch them. But Cloudflare has not published how it decides that a page is ad-displaying, so this is worth checking rather than assuming. If you run display ads or affiliate units anywhere, set a reminder before September 15 and go look at your Security settings.
Cloudflare CEO Matthew Prince framed the change around the traffic mix rather than fairness, telling TechCrunch that now that the majority of traffic on the internet is non-human, the company must go further and act faster. Either way, the operational lesson is the same: check that box carefully, or not at all.
Blocking Is Not Free
The instinct when you read all this is to block everything. Resist it, because there is measurement on the other side of that decision.
Researchers at Rutgers Business School and Wharton studied publishers who blocked AI crawlers via robots.txt. Their December 2025 paper found blocking reduced large publishers' total traffic by roughly 23% and human-only traffic by roughly 14%, without reliably reducing AI citation rates. A later revision reports a smaller figure, closer to 7%, on a weekly measurement window, so treat the magnitude as unstable. The direction is what matters: blocking cost traffic and did not buy the protection it was supposed to buy.
For a local business that would quite like to be the answer ChatGPT gives when somebody asks for a roofer in Buffalo, blanket AI blocking works against you. We wrote a whole post on that tension, how to block ChatGPT in robots.txt and why you probably should not, and the 2026 data has only strengthened the case.
Robots.txt is also a request, not a fence. A peer-reviewed study presented at the ACM Internet Measurement Conference in 2025 tracked 130 self-declared bots over 40 days and concluded that compliance is selective and unreliable, with certain categories of bots, including AI search crawlers, rarely checking robots.txt at all.
Some crawlers also work around blocks. On August 4, 2025 Cloudflare accused Perplexity of stealth crawling, reporting an undeclared crawler doing 3 to 6 million requests a day across tens of thousands of domains while spoofing a macOS Chrome user agent, and de-listed it as a verified bot. Perplexity disputed the finding.
What Actually Works: Make the Waste Stop
Here is the statistic that should reframe the whole problem. Cloudflare reports that more than 50% of crawl traffic from good bots goes to re-fetching pages that have not changed. More than half. Not malicious, not AI-specific, just waste, because the server never told the crawler that nothing was new.
That is the hinge. The fix that pays best is not blocking. It is caching correctly, so a bot asking for an unchanged page gets a tiny 304 response instead of a full page render. Google's crawling infrastructure supports both ETag with If-None-Match and Last-Modified with If-Modified-Since, prefers ETag when both are present, and Google explicitly recommends serving a 304 to conserve crawl budget and server bandwidth. It is free, standards-based, officially supported, and it makes the site faster for humans at the same time. Google's own crawl budget documentation update is worth reading alongside this.
The Practical Checklist, Cheapest First
- Look at real logs first. Raw access logs or analytics, filtered by user agent. Most sites will find this is a non-issue and can stop right here. Everything below costs time or money, so earn it with data.
- Fix your crawl traps. Disallow in robots.txt plus noindex on filter and facet URLs, internal search results, and calendar pagination. Ordinary SEO hygiene that happens to cut crawler load hard, and it is free.
- Set correct cache headers. Cache-Control, ETag and Last-Modified, so repeat fetches get a 304 instead of rebuilding the page. This is where the waste lives, and it doubles as a speed win. Our guide to website speed and how to optimize it covers the same headers from the human-visitor side.
- Put Cloudflare's free tier in front of the site and open AI Crawl Control. It is on all plans, needs no configuration, and shows which AI services are crawling you, with per-crawler allow and block controls.
- Use rate limiting as a backstop, not a first move. Set it loose enough that Googlebot never trips it. A rate limit that catches your own search crawler is worse than none.
- Only then consider blocking, and be surgical. Block a specific misbehaving crawler you can see in your own logs. Do not flip a category-wide switch you do not fully understand.
Hosting Quality Is Now an SEO Question and a Cost Question at Once
The same choices now pay off twice. Correct caching cuts crawler waste and speeds up the site for customers. A host with real CPU headroom, instead of a heavily oversold shared box, absorbs a crawler burst without throttling, so Googlebot never sees a slow or failing response from you. Clean URL structure with no infinite spaces means crawl capacity goes to pages that matter instead of the 4,000th variation of a filter query.
Ten years ago you could argue hosting was a commodity and any cheap plan was fine. In a web where most requests are automated, that argument is finished. Our hosting plans and Buffalo web hosting pages spell out what we provision and why, and our Buffalo SEO work starts from the technical layer rather than bolting it on later.
None of this requires panic. It requires looking at your logs once, fixing the two or three things that are genuinely broken, and not paying for a solution to a problem you do not have.
Want Hosting That Absorbs This Without You Noticing?
AldoMedia hosts and maintains sites for businesses across Buffalo and Western New York, with caching, sane resource limits and crawl hygiene handled as part of the job rather than as an upsell.
See Hosting Plans Buffalo Web HostingSources
- Imperva: 2025 Bad Bot Report, How AI is Supercharging the Bot Threat
- Imperva: Bad Bot Report 2026, Bots in the Agentic Age
- Cloudflare: The 2025 Radar Year in Review
- Cloudflare: The crawl-to-click gap, data on AI bots, training and referrals
- Cloudflare: The crawl before the fall of referrals
- Wikimedia Diff: How crawlers impact the operations of the Wikimedia projects
- Read the Docs: AI crawlers need to be more respectful
- SourceHut status: LLM crawlers continue to DDoS SourceHut
- TechCrunch: How OpenAI's bot crushed this seven-person company's website like a DDoS attack
- Fastly: AI Bots in Q2 2025, Threat Insights Report
- Akamai: AI Bots Threaten the Foundation of Web-Based Business Models
- DataDome 2025 Global Bot Security Report
- Cloudflare: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
- TechRadar: Perplexity hits back after Cloudflare slams its online scraping tools
- Scrapers selectively respect robots.txt directives (ACM IMC 2025)
- Cloudflare: Your site, your rules, new AI traffic options for all customers
- Cloudflare: Making AI search smarter
- TechCrunch: Cloudflare's new policy pushes AI companies to pay for publishers' content
- Help Net Security: Cloudflare changes AI crawler access rules
- Cloudflare docs: AI Crawl Control
- Zhao and Berman: Strategic Response of News Publishers to Generative AI
- Google Search Central: Google crawlers and fetchers overview
- Search Engine Journal: Google recommends using 304 to conserve crawl budget
Frequently Asked Questions
Are AI crawlers really costing my small business money?
Almost certainly not in an amount you would notice. Every documented cost disaster involved a site with tens of thousands to millions of URLs. A normal brochure site of twenty to two hundred pages, fully re-crawled by every AI bot on earth every single day, still adds up to a few hundred megabytes a month. Look at your own server logs before you spend a dollar on a fix.
What percentage of web traffic is bots?
Imperva measured automated traffic at 51% of all web traffic in 2024 and more than 53% in 2025, the second straight year that bots outnumbered people. AI crawlers are only a slice of that total. Across 2025, Cloudflare measured AI bots other than Googlebot at 4.2% of HTML requests, and Googlebot alone at 4.5%.
Should I block AI crawlers in robots.txt?
Usually not, and almost never all of them. Researchers at Rutgers Business School and Wharton found that publishers who blocked AI crawlers lost traffic without reliably reducing how often AI tools cited them. If you want your business to be the answer an AI assistant gives when somebody asks for a contractor in Buffalo, blanket blocking works against you.
What is Cloudflare changing on September 15, 2026?
Cloudflare has announced that from September 15, 2026 it will block Training and Agent bots by default on ad-displaying pages, while Search bots stay allowed. It applies to newly onboarding domains, new sites added by existing customers, and free-tier customers. Owners can opt out in their Security settings before that date.
Can blocking AI training crawlers hurt my Google rankings?
Yes, and this is the real risk for a small business. Cloudflare judges multi-purpose crawlers on all of their behaviors, so choosing to block Training also blocks Googlebot, Applebot and Bingbot even when Search is set to allowed. Ticking the wrong box threatens your search visibility far more than any crawler ever threatened your hosting bill.
What is the cheapest way to cut crawler load on my site?
Fix crawl traps and cache properly. Blocking crawlable filter URLs, internal search results and endless calendar pagination costs nothing and cuts crawler load hard. Then set correct Cache-Control, ETag and Last-Modified headers so an unchanged page returns a 304 instead of being rebuilt. Cloudflare says more than half of good-bot crawl traffic is re-fetching pages that have not changed.
