Why IP Blocking Fails Against Modern Bots (and What to Use Instead)
IP bans hit real users behind CGNAT and proxies while scrapers rotate addresses for pennies. Why IP-based bot blocking fails, and what to key on instead.
Blocking an IP address is usually the first thing anyone does about a scraper. You see thousands of requests from one address in the logs, add a deny rule, and the traffic stops. For a day.
IP blocking made sense when one address meant one machine and getting new addresses was expensive. Neither is true anymore. This article covers why IP-based blocking fails in both directions, blocking real users and missing bots, and what identifier to use instead.
An IP address is not a client
An IP address identifies a network endpoint at a moment in time. It says nothing reliable about who or what is sending the request. Several very common setups put many unrelated clients behind one address:
- Carrier-grade NAT (CGNAT). Mobile carriers and many residential ISPs put thousands of subscribers behind a small pool of public IPv4 addresses. There’s even a dedicated address block for it,
100.64.0.0/10from RFC 6598, used between the subscriber and the carrier’s NAT. - Corporate proxies and VPNs. An entire office, sometimes an entire company, reaches the internet through a handful of egress IPs.
- Privacy relays. Services such as iCloud Private Relay route users through shared egress addresses on purpose.
- Universities, libraries, hotels, cafés. Shared Wi-Fi means shared addresses.
If a single scraper is running on a phone network, blocking its IP also blocks every other subscriber who happens to share that exit address. You’ll rarely hear about it. Those users just see an error page and leave.
Bots don’t care about their IP
Scraper operators have adapted to IP blocking completely:
| Technique | What it costs the operator | What it does to your IP blocklist |
|---|---|---|
| Cloud VMs | Cents per hour, new IP per instance | Blocklist grows forever, bot barely slows |
| Datacenter proxy pools | Cheap, thousands of IPs | Rotation per request defeats per-IP rate limits |
| Residential proxy networks | More expensive, millions of IPs | Traffic comes from real home connections you can’t ban |
| IPv6 | A single /64 has 2⁶⁴ addresses | Per-address blocking is meaningless |
Residential proxies are the hardest case. The requests really do come from consumer ISPs, often from devices whose owners installed an app that resells their bandwidth. Banning those addresses means banning ordinary households.
IPv6 makes per-address blocking pointless on its own terms. A single customer allocation is commonly a /64 or larger. You can block by prefix instead, but then you’re back to the shared-address problem at a different scale.
Blocking by ASN or country
When IP blocking stops working, teams often move up a level and block entire autonomous systems (ASNs) or countries. This does stop the bot, but it’s the same trade-off with a bigger blast radius:
- Blocking a cloud provider’s ASN also blocks legitimate integrations, monitoring services, and users on corporate VPNs hosted there.
- Blocking a country blocks every real customer there, and scrapers just pick an exit node elsewhere.
ASN is still a useful signal. Traffic from a hosting provider is more likely to be automated than traffic from a consumer ISP. The mistake is using a signal like that as the identifier you block on.
User-agent rules have the opposite problem
The other traditional approach is matching the User-Agent header. Blocking python-requests or curl does catch lazy scripts, and it’s worth doing. But the user agent is a string the client writes itself. Any scraper that cares will send a current Chrome user agent, and a user-agent rule can’t tell it apart from real Chrome.
So IP rules block too many people, and user-agent rules block too few bots.
What to key on instead: the client’s fingerprint
What you want is an identifier that:
- Is shared by all requests from the same client software, so blocking it actually stops the bot even when it changes IPs.
- Is not shared by unrelated users on the same network, so blocking it doesn’t cause collateral damage.
- Is hard for the client to fake, unlike a header it writes freely.
Request fingerprints fit all three. Every HTTPS connection starts with a TLS ClientHello, and the exact cipher suites, extensions and parameters in it are decided by the client’s TLS library, not by a string the client sets. Python’s ssl module, Go’s crypto/tls, curl built against OpenSSL and Chrome’s BoringSSL all produce recognisably different handshakes. We cover this in detail in TLS fingerprinting explained.
Combine that with how the client speaks HTTP and the network it comes from, and you get a composite fingerprint that is:
- stable across IP rotation, because the bot’s software doesn’t change when its proxy does;
- different from the real Chrome user next to it on the same CGNAT address, because their TLS stacks differ even when the user-agent strings match.
A practical migration path
You don’t have to drop IP rules overnight. A sensible progression:
- Log fingerprints alongside IPs. Start recording TLS and HTTP signals for every request, so you can see how many distinct clients sit behind your worst IPs. It’s often surprising.
- Keep IP rules only for clear-cut cases. A single-tenant server hammering you with exploit probes can still be blocked by IP, ideally with a short expiry.
- Move sustained blocking to fingerprints. For recurring scrapers, block the fingerprint. Blocks then follow the bot across IP changes and leave the network’s other users alone.
- Exempt verified crawlers explicitly. Search engines and monitoring services should be checked properly, not by user-agent string. See how to verify Googlebot.
Key takeaways
- An IP address identifies a network endpoint, not a client. CGNAT, proxies and relays put many users behind one address.
- Bots rotate IPs cheaply through cloud VMs, proxy pools and IPv6, so IP blocklists grow without stopping anything.
- ASN and country blocks stop bots but have a huge blast radius. Use them as signals, not as identifiers to block on.
- Fingerprints built from the TLS handshake and HTTP behavior survive IP rotation and separate a bot from a real user on the same IP.