Spoofed Googlebot: How to Verify Search Engine Crawlers
Many requests claiming to be Googlebot aren't. Verify crawlers with reverse and forward DNS or published IP ranges, and learn why user-agent checks aren't enough.
Every site wants to be crawled by Google, and every scraper knows it. Pretending to be Googlebot is one of the oldest bot tricks. Many sites exempt anything with Googlebot in its user agent from rate limits and blocks, so claiming it is a free pass.
The fix is simple: never trust a crawler’s user agent. Verify it. This article covers how verification works, with code, and how to treat clients that fail it.
Why the user agent isn’t proof
The Googlebot user agent looks something like this:
Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
Any HTTP client can send that string. When you check your logs for requests claiming to be Googlebot, a meaningful share usually comes from hosting providers and residential networks that have nothing to do with Google.
Exempting by user agent has a second cost. Blocking legitimate crawlers by mistake, for example by rate-limiting “Googlebot” traffic that is really Googlebot, hurts your search visibility. You need to tell the real one from the fakes, and the only reliable way is to check where the request comes from.
Method 1: reverse and forward DNS
Google documents a two-step DNS check, and Bing and other major search engines support the same approach:
- Reverse DNS: look up the hostname for the requesting IP. For Googlebot it must end in
googlebot.com,google.comorgoogleusercontent.com. - Forward DNS: resolve that hostname back to IP addresses. The original IP must be among them.
Step 2 matters. Anyone who controls an IP block can set its reverse DNS to crawl-1-2-3-4.googlebot.com. They can’t make Google’s DNS resolve that name back to their IP.
import socket
from functools import lru_cache
CRAWLER_DOMAINS = {
"googlebot": (".googlebot.com", ".google.com", ".googleusercontent.com"),
"bingbot": (".search.msn.com",),
}
@lru_cache(maxsize=50_000)
def verify_crawler(ip: str, crawler: str) -> bool:
suffixes = CRAWLER_DOMAINS[crawler]
try:
host, _, _ = socket.gethostbyaddr(ip) # reverse lookup
except (socket.herror, socket.gaierror):
return False
if not host.endswith(suffixes):
return False
try:
_, _, addresses = socket.gethostbyname_ex(host) # forward lookup
except (socket.herror, socket.gaierror):
return False
return ip in addresses
Note that gethostbyname_ex only returns IPv4 addresses. For IPv6 clients, use socket.getaddrinfo(host, None) and compare normalized addresses.
Performance: DNS lookups are slow compared to request handling. Don’t run them inline on every request. Cache results aggressively, since crawler IPs are stable, and verify asynchronously: let the first request through, verify in the background, and apply the result to later requests.
Method 2: published IP ranges
Google also publishes its crawler IP ranges as JSON files, split by crawler type: common crawlers like Googlebot, special-case crawlers, and user-triggered fetchers. Other large operators publish similar lists.
Checking an IP against a list of CIDR ranges is fast and needs no DNS:
import ipaddress, json, urllib.request
GOOGLEBOT_RANGES_URL = "https://developers.google.com/static/search/apis/ipranges/googlebot.json"
def load_ranges(url: str):
with urllib.request.urlopen(url, timeout=10) as r:
data = json.load(r)
return [
ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix"))
for p in data["prefixes"]
]
GOOGLEBOT = load_ranges(GOOGLEBOT_RANGES_URL)
def in_ranges(ip: str, ranges) -> bool:
addr = ipaddress.ip_address(ip)
return any(addr in net for net in ranges if net.version == addr.version)
Refresh the lists regularly (daily is plenty) and keep the last good copy if a download fails. For large lists, a prefix trie or sorted-interval search beats a linear scan.
Method 3: let your edge provider do it
If you’re behind a CDN with bot management, verification may already be done for you. Cloudflare maintains a list of verified bots and exposes the result to firewall rules and, on zones with Bot Management, to Workers through request.cf.botManagement.verifiedBot, along with a verified-bot category.
Using the provider’s verification has two advantages: zero lookup latency, and coverage of many more bots (SEO tools, uptime monitors, social link previewers) than you’d maintain yourself.
What to do with each outcome
Verification gives you three groups, and each needs a different policy:
| Claims to be a crawler? | Verified? | Treatment |
|---|---|---|
| Yes | Yes | Exempt from bot rules. Apply only generous rate limits. |
| Yes | No | Spoofed. Treat this as a strong bot signal, not a neutral one. |
| No | n/a | Evaluate normally. See signs your site is being scraped. |
The middle row deserves emphasis. A client that claims to be Googlebot and isn’t has lied about its identity on purpose. That’s much more suspicious than a client with a generic user agent, and it’s safe to block. In a rule-based pipeline, “claims a search-engine user agent and fails verification” works well as a ban rule of its own.
Put the exemption first in your rule order. A verified crawler that happens to request /wp-login.php while following a broken link should not be banned as an exploit scanner.
Don’t forget the non-search crawlers
Search engines aren’t the only automated traffic you want. Consider allowing (after verification):
- Social media link previewers, so shared links show cards.
- Uptime and performance monitors you or your customers use.
- Feed readers for your RSS feed.
- AI crawlers, which is a policy decision. Some sites welcome them for visibility, others block them. Decide deliberately rather than by accident.
Key takeaways
- A crawler’s user agent is a claim, not an identity. Scrapers use
Googlebotprecisely because sites trust it. - Verify with reverse DNS plus forward DNS confirmation, or against the operator’s published IP ranges.
- Cache results and verify off the request path. CDN-provided verified-bot flags avoid lookups entirely.
- Exempt verified crawlers before any other rule, and treat failed verification as a strong bot signal.