Surendra Tamang

How antibot systems fingerprint your scraper

· Surendra Tamang

When a scraper gets blocked, most people change the User-Agent and add a proxy. Sometimes that works for a day. Then the blocks come back, because the User-Agent was never the main thing being checked.

Cloudflare, DataDome, Akamai, PerimeterX (HUMAN) and Kasada all score many signals together. This post walks through them in the order a request actually meets them, and ends with what holds up in production.

What a protection vendor sees before your first request finishes

Before your HTTP request is even read, the server already knows several things:

  1. Your IP and its network. Is it a datacenter range (AWS, Hetzner, OVH), a residential ISP or a mobile carrier? Has this IP been seen on other protected sites recently?
  2. Your TLS handshake. Which ciphers, extensions and curves your client offered, in what order.
  3. Your HTTP/2 connection. The SETTINGS values, window sizes and pseudo-header order your client sends.

Only then come headers, cookies, JavaScript challenges and behaviour. Each layer adds to a risk score. Crossing a threshold gets you a challenge page, a block or, worse, quietly fake data.

TLS fingerprinting: you’re identified at the handshake

Every TLS client says hello differently. Chrome, Firefox, Python requests, Go’s net/http and Node all send a distinct ClientHello. Hashing its fields gives a fingerprint. JA3 was the original; JA4 is the newer, more robust version.

This is why this fails:

import requests
requests.get(url, headers={"User-Agent": "Mozilla/5.0 ... Chrome/124.0 ..."})

The header claims Chrome, but the TLS handshake says Python urllib3. That mismatch is one of the cheapest signals a vendor can check, and one of the most reliable.

The fix is a client that reproduces a real browser handshake, like curl_cffi:

from curl_cffi import requests
r = requests.get(url, impersonate="chrome")

Now the TLS and HTTP/2 layers look like Chrome. Headers alone never could.

HTTP/2 frames and header order

HTTP/2 adds another fingerprint. Browsers send specific SETTINGS values (header table size, initial window size, max concurrent streams), a WINDOW_UPDATE of a specific size, and pseudo-headers (:method, :authority, :scheme, :path) in a fixed order. Each browser family does this slightly differently, and HTTP libraries differ from all of them.

Regular header order matters too. Chrome sends sec-ch-ua, sec-ch-ua-mobile and sec-ch-ua-platform near the top, then user-agent, accept and the sec-fetch-* headers in a consistent order. A request with the right headers in the wrong order, or with sec-ch-ua claiming Chrome 124 while the User-Agent says 120, stands out.

The rule: every layer must tell the same story. IP location, TLS, HTTP/2, headers, and later the JavaScript environment all need to agree on who you are.

Browser-side signals: canvas, fonts, and API probing

If the site serves a JavaScript challenge, it runs code in your browser and reports back. Typical probes:

  • Automation flags: navigator.webdriver, and traces left by the Chrome DevTools Protocol that Playwright and Puppeteer use.
  • Environment consistency: does navigator.platform match the User-Agent? Do the screen size, timezone and language match the IP’s country?
  • Rendering: canvas and WebGL output, the GPU renderer string, and installed fonts. Headless browsers on Linux servers produce recognisable results.
  • API behaviour: how native functions stringify, whether properties were patched, and timing quirks that differ in headless mode.
  • Behaviour: mouse movement, scrolling, and time between page load and interaction.

Challenge scripts are obfuscated and change often, so the exact checks shift. The categories stay the same.

What actually works in production

No single trick. It’s a set of habits:

  • Use the lightest client that passes. Try a TLS-impersonating HTTP client first. It’s fast and cheap. Go to a real browser only for sites whose JavaScript challenge requires one.
  • When you need a browser, use one built for it. Patched or stealth-focused browsers (for example Camoufox, or Chromium driven without the usual automation leaks) pass far more often than stock headless Chrome.
  • Match the proxy to the target. Datacenter IPs are fine for unprotected sites. Heavily protected ones need residential or mobile IPs, geolocated to match your browser’s timezone and language.
  • Keep sessions sticky. Reuse one IP, one fingerprint and one cookie jar per session. Rotating the IP every request while keeping the same cookies is a red flag, not camouflage.
  • Slow down. Most blocks I debug come from request rate, not fingerprints. Realistic concurrency per IP beats clever evasion.
  • Measure block rate. Log challenge pages and 403s per source. A rising block rate warns you before your data drops. That’s the core of monitoring a scraper properly.

And stay on the right side of the line. Collect publicly available data, respect rate limits, avoid personal data you don’t need, and read the site’s terms before you start.

FAQ

Why do I get blocked even with residential proxies? Because the IP is only one signal. If your TLS fingerprint says Python while your User-Agent says Chrome, a clean residential IP won’t save you.

Is headless Chrome enough? For lightly protected sites, often yes. For DataDome, Kasada or strict Cloudflare settings, stock headless Chrome is easy to detect. Use a stealth-focused browser, or find the site’s underlying API.

What’s JA3 and JA4? Ways to hash a TLS ClientHello into a fingerprint. JA4 is the newer format, and it’s harder to fool by shuffling extension order.

Can I bypass these systems permanently? No. Vendors update their checks continuously. What lasts is a setup that detects blocks early and adapts: monitoring, a fallback client, and time budgeted for maintenance.

#antibot#reverse-engineering#web-scraping