● early access
Edge security

What Bot Traffic Really Costs Your Origin (and How to Estimate It)

Bots cost more than bandwidth: compute, database load, cache misses and skewed analytics. A simple model to estimate what bot traffic costs your origin.

BannedRobots Teampublished 4 min read

Industry reports, such as Imperva’s annual Bad Bot Report, have estimated for several years that automated traffic is roughly half of all web traffic, and a large share of it is unwanted. Whatever the exact number for your site, bots are probably a significant line in your infrastructure bill, even though nobody lists them there.

This article breaks down where bot traffic costs money, gives a simple model to estimate it for your own site, and explains why blocking at the edge saves more than blocking at the origin.

Where the cost comes from

1. Egress bandwidth

Every response you serve is data leaving your infrastructure. Major cloud providers charge for outbound data transfer. At list prices, the first tier on the big clouds has long been in the range of several cents per gigabyte (check your provider’s current pricing). Scrapers fetch your heaviest pages repeatedly, and some download images and media in bulk.

2. Compute

Dynamic pages cost CPU to render. A scraper walking your product catalog forces your application to build every page, often pages real users rarely visit, which are also the ones least likely to be cached.

3. Database load

Behind each dynamic page are queries. Crawlers that enumerate search results, filters and pagination produce query patterns real users don’t: deep offsets, unusual filter combinations, every page of every listing. These are expensive queries, and they compete with real users for connections.

4. Cache misses

Caches are sized for real traffic patterns. Bots break them in two ways:

  • Long-tail access. Crawlers request everything, evicting popular content to make room for pages nobody else will ask for.
  • Cache busting. Adding random query strings (?_=1727712345) or varying headers makes every request a miss, sending all of it to origin.

A drop in cache hit ratio multiplies every other cost on this list.

5. Things that don’t show up on the bill

  • Skewed analytics: conversion rates, bounce rates and A/B test results computed over traffic that includes bots.
  • Capacity planning: you scale for peak load, and part of the peak isn’t customers.
  • Security incidents: vulnerability scanners are bots too. Each probe that reaches your application is a chance for one to succeed.
  • Content and data theft: scraped prices, listings and articles have business value that doesn’t appear in infrastructure metrics.

A simple estimation model

You can estimate the direct cost from numbers you already have. For a given month:

bot_requests      = total_requests × bot_share
egress_cost       = bot_requests × avg_response_size_GB × egress_price_per_GB
compute_cost      = bot_requests × origin_miss_rate × cost_per_origin_request
total_bot_cost    ≈ egress_cost + compute_cost

Where:

  • bot_share comes from your own analysis. A rough start is the share of requests from self-declared tools, unverified “search engines” and hosting networks. See signs your site is being scraped.
  • origin_miss_rate is the share of bot requests that aren’t served from cache. For bots it’s typically much higher than for users.
  • cost_per_origin_request is your monthly compute and database cost divided by origin requests. It’s crude, but it gives you a scale.

A worked example

These numbers are hypothetical. Substitute your own.

Input Value
Total requests / month 200 M
Bot share 30 %
Average response size 150 KB
Egress price $0.08 / GB
Origin miss rate (bots) 60 %
Cost per origin request $0.000005
bot_requests   = 200M × 0.30                  = 60M
egress         = 60M × 0.00015 GB × $0.08     ≈ $720
compute        = 60M × 0.60 × $0.000005       ≈ $180
direct total                                   ≈ $900 / month

For many sites the direct number is modest. The bigger costs are usually indirect: extra capacity provisioned for bot-inflated peaks, database instances sized for crawler query patterns, engineering time on incidents, and the business impact of scraped data. The model gives you a floor, not a ceiling.

Why where you block matters

Blocking a bot at the origin, in your application or its web server, still costs you:

  • the request reaching your infrastructure (and, depending on your setup, the network transfer to get there);
  • a worker or connection to handle it;
  • whatever your application does before the block decision is made.

Blocking at the edge, before the request is forwarded, removes almost all of that. The CDN absorbs the request and your origin never sees it. The cost per blocked request drops from “a small slice of your server” to “a small slice of an edge function invocation”.

The ideal is also precise. A block that also catches real customers, as broad IP or ASN bans tend to (see why IP blocking fails), exchanges an infrastructure cost for a revenue cost. That’s usually a bad deal.

Measuring the savings

After deploying edge blocking, track:

  • Origin request count and egress volume. The direct effect.
  • Cache hit ratio. Often the most visible improvement.
  • Database load during crawl-heavy periods.
  • p95 and p99 latency for real users. Less contention helps everyone else.
  • Block rate and appeal rate. Blocks nobody complains about are a good sign. A rising appeal rate means the rules are too broad.

Key takeaways

  • Bot traffic costs egress, compute, database capacity and cache efficiency, plus analytics accuracy and security exposure.
  • Estimate direct cost with bot share × response size × egress price, plus origin misses × cost per request.
  • Direct costs are a floor. Capacity headroom, incidents and scraped data usually cost more.
  • Block at the edge, and block precisely, so savings don’t come at the expense of real users.
bot trafficinfrastructure costegresscachingfinops