Agent Analytics

Crawler Verification

Crawler verification process

A service name in a User-Agent does not establish that the request came from that service. Verification checks DNS, published IP ranges, ASN information, and behavioral patterns against the claimed crawler identity.

Why verification matters

  • Distinguish impersonation
  • Improve analytics accuracy
  • Understand resource use
  • Evaluate access against security requirements
  • Improve delivery to legitimate crawlers

Primary verification techniques

MethodWhat it checks
Reverse DNSHostname and organization domain associated with the IP
IP rangesMembership in published crawler address ranges
ASNSource network ownership
HeuristicsCharacteristic signatures and access patterns

A shared-cloud ASN, User-Agent, or behavioral pattern alone cannot establish service authenticity. The list below describes verification methods; it does not imply equal verification confidence for every crawler.

Platform-specific verification

Platforms verifiable through published information

Partially verified platforms

Some platforms use common cloud provider IPs or don't publish verification methods, making complete verification challenging:

  • You.com
  • Bytedance
  • Yahoo
  • OpenClaw
  • DeepSeek
  • Baidu
  • BaiduSpider: Reverse DNS verification
  • Huawei
  • Yandex
  • Gemini

Stay updated

Review verification methods as platforms emerge, crawler infrastructure changes, new methods become available, and security requirements evolve. Do not treat a fixed list as sufficient for future traffic.

Important note about data updates

Changes to identification methods can change historical and current classifications. When traffic shifts substantially, distinguish actual volume changes from verification or classification updates.