Roughly half of all web traffic is automated. Your analytics dashboard probably says it's close to zero. Both numbers are real — and the gap between them is the whole problem this guide solves.
Bot traffic skews every number you make decisions with: conversion rates, bounce rates, A/B tests, ad attribution, "how many people read this post." It also includes the traffic that actually costs you money — scrapers stealing content, credential stuffers hammering your login page, and exploit scanners probing for a way in. You can't make a single good decision about any of it until you can see it.
Here's how detection actually works, from the weakest signal to the strongest, and how to check your own traffic today.
Why your analytics can't see bots
Almost every analytics tool works the same way: a JavaScript snippet runs in the visitor's browser and phones home. That design fails against bots in both directions at once.
Simple bots don't run JavaScript. curl, Python scripts, Go scrapers, most crawlers — they fetch your HTML and never execute the snippet. To your analytics they don't exist. Your server is doing the work; your dashboard shows nothing.
Sophisticated bots run a real browser. Headless Chrome via Puppeteer or Playwright executes your snippet perfectly, fires every event, and gets counted as a human — often an engaged human, since scrapers visit lots of pages.
So JavaScript-based analytics undercounts the dumb bots and miscounts the smart ones as people. The only place every request is visible — bot or not, JavaScript or not — is the server side: your logs, or your edge. Keep that in mind for everything below, because it determines where detection has to happen.
Seven signs you have bot traffic right now
Before any fancy fingerprinting, these patterns are visible in whatever analytics you already have:
- Zero-duration sessions in bulk. A spike of visits with one pageview, no scroll, no clicks, gone in under a second.
- Traffic spikes at 3 a.m. Humans follow time zones. Cron jobs don't.
- Geography you don't sell to — especially datacenter cities. Ashburn (Virginia), Frankfurt, Singapore, and Council Bluffs appear in everyone's analytics for the same reason: that's where the servers live, not the readers.
- Superhuman browsing speed. Forty pages a minute is not a person skimming. It's a crawler walking your sitemap.
- Ancient browser versions. "Chrome 101" showing up in 2026 means someone pinned a user-agent string years ago and never updated the lie.
- Hits on pages humans never visit.
/wp-login.phpon a site that doesn't run WordPress,/.env,/phpmyadmin— that's exploit scanning, and it never comes from a person. - Referrer spam. Traffic "referred" by domains you've never heard of, hoping you'll click the referrer out of curiosity.
Any one of these is suggestive. The clusters are damning. But spotting symptoms isn't the same as detection — for that you need signals a bot can't easily fake.
The five layers of bot detection
Think of these as a ladder: each layer is harder to spoof than the one below it.
Layer 1: The user-agent string
Every request announces an identity: Mozilla/5.0 (compatible; MJ12bot/v1.4.8; ...) or python-requests/2.31 or a full Chrome string. This is self-declared — the bot equivalent of a name tag — and it catches two big groups: honest crawlers that identify themselves (the named AI crawlers, search engines, SEO tools) and lazy tooling that doesn't bother lying (curl, wget, HTTP libraries).
A useful rule: a bot that tells you its name usually takes no for an answer. The ones that lie about their name are the ones worth defending against — and the user-agent is useless against exactly those. One flag in a script turns "python-requests" into "Chrome." That's why this is layer 1 of 5, not the whole system, and why any tool whose detection stops at user-agent matching is theater.
Layer 2: The network
Every IP belongs to a network (an ASN) with a known operator. A request from Comcast or Vodafone plausibly carries a person. A request from AWS, Hetzner, or a bulletproof VPS host does not — people don't read blogs from inside a datacenter. Datacenter ASN lists are public, and this single check unmasks an enormous amount of "Chrome" traffic that's actually scrapers on cloud boxes.
Two honest caveats. First, VPN exit nodes live in datacenters too, and VPN users are real humans — a good classifier carves VPNs out rather than punishing them. Second, serious scrapers rent residential proxies, routing through real home connections precisely to beat this layer. The network signal is strong evidence, not proof — which is why the ladder keeps going.
Layer 3: Protocol and TLS fingerprint
This is the lie detector. A real Chrome 125 doesn't just say it's Chrome — it speaks like Chrome: HTTP/2 or HTTP/3, TLS 1.3, a specific cipher list in a specific order, a specific shape of TLS handshake. Those properties come from the browser's compiled networking stack, and an HTTP library can't fake them without reimplementing the browser.
So when a request claims Chrome 101 on Android but arrives speaking HTTP/1.1 over TLS 1.2 from an Amazon datacenter IP — a real example from our own feed — every layer disagrees with the name tag. The user-agent says phone; the handshake says script; the network says server rack. That contradiction is the detection. No single signal had to be perfect; the mismatch did the work.
This is also the layer that's simply unavailable to JavaScript analytics: the TLS handshake happens before any page exists. Only the server (or the edge terminating the connection) ever sees it.
Layer 4: Behavior
Identity signals say what a visitor is; behavior says what it's doing. The strong behavioral tells:
- Request velocity per identity — sustained superhuman pacing from one fingerprint.
- Attack paths — a request for
/.envor/wp-adminis a verdict all by itself, whatever the user-agent claims. - No subresources. Real browsers fetch CSS, images, and fonts. A "browser" that requests only HTML, page after page, is parsing you, not rendering you.
- Coordinated pools — many fresh identities surfacing from one network in a burst is one operator with a proxy list, not a sudden fan club.
Behavior is also your defense against the residential-proxy problem from layer 2: the IP looks like a home, but homes don't browse 400 product pages in arithmetic progression.
Layer 5: Verification
The strongest signal answers the opposite question: not "is this a bot?" but "is this bot who it claims to be?" Major crawlers are verifiable — published IP ranges, reverse-DNS checks, edge-level verified-bot attestation. Real Googlebot passes. The fake "Googlebot" scraping you from a VPS fails.
That gives you the single most valuable rule in bot policy: a claimed good bot that fails verification is a bad bot, automatically. No judgment call needed. It told you a checkable lie; the check failed. This is how you block AI crawlers without ever touching your search rankings — verified search crawlers are provably themselves, so nothing you do to the liars can hurt them.
The cheat sheet
| Signal | Catches | Fooled by |
|---|---|---|
| User-agent | Honest crawlers, lazy scripts | One line of code |
| Network / ASN | Datacenter scrapers, VPS fleets | Residential proxies |
| TLS + protocol fingerprint | Headless browsers, HTTP libraries lying about identity | Real browser stacks (expensive) |
| Behavior | Velocity, attack probes, proxy pools | Slow, careful, human-paced bots |
| Verification | Impostors of known bots | Nothing — it's cryptographic-grade attribution |
No layer is sufficient. The stack is the detector: every request scored on all five, with the contradictions doing the heavy lifting.
Check your own site today
If you have server or proxy logs, three greps will tell you more than a month of dashboard-staring:
- Count requests whose user-agent doesn't start with
Mozilla. That's your floor — the bots that aren't even pretending. - Look up the ASN of your top 20 IPs (any whois tool). Every datacenter network in that list is automation, whatever its user-agent claims.
- Find "Chrome" speaking HTTP/1.1. Modern Chrome has preferred HTTP/2+ for nearly a decade. Old protocol + new browser string = script with a name tag.
If you can't do this — because your host doesn't expose logs and everything you know comes from a JavaScript dashboard — then you've found the real issue: you have no visibility into the majority of your traffic at all. Detection has to live where every request passes, which in 2026 means the edge: classify the request on arrival, from the handshake up, before deciding what to serve it.
Classify, don't just "detect"
A binary bot-or-human switch fails immediately in practice, because the right response differs wildly by which bot:
- Verified search crawlers — welcome, always. That's your traffic.
- AI crawlers — a policy decision you should make deliberately, per bot. Either answer is defensible; not choosing isn't.
- Neutral bots — SEO tools, archivers, monitors. Harmless to most sites, blockable on your call.
- Bad bots — impostors, scrapers, scanners, credential stuffers. Block them — or better, make them suffer.
- Humans — everyone above is in service of not annoying these.
And trust signals in the right order: verification beats network evidence, network evidence beats behavior, and the self-declared name tag comes last. Our crawler directory profiles the named bots — what each collects, whether it honors robots.txt, and how to verify it — and if you just want a policy file, the free robots.txt generator builds one in a minute. Evaluating products instead? We compared the serious bot detection tools honestly, including the ones that aren't us.
The short version
Your analytics can't see bot traffic because it measures JavaScript, and bots either skip JavaScript or fake it perfectly. Real detection is server-side, layered, and adversarial: take the request's claimed identity, then check it against the network it came from, the way it speaks TLS, how it behaves, and — for known crawlers — whether it can prove it's really them. Honest bots identify themselves. Liars contradict themselves. The contradictions are the detection.
TrafficDATA runs all five layers on every request at the edge — TLS fingerprint, network reputation, behavioral scoring, and verified-bot checks — and classifies each visitor as human, good bot, AI crawler, neutral, or bad bot in the live feed, in real time. The free tier covers 100K page views, which is more than enough to find out what your traffic is actually made of.