Web crawler lookup

What is Diffbot?

Diffbot is a generic crawler. Use this crawler profile to identify user-agent tokens, operator signals, platform hints, and recommended handling.

Diffbot

Data collection crawler

Version
1.0
First seen
Jul 21, 2026
Confidence
Inferred user-agent token

Crawler tags

Data collection crawlerAI trainingObeys robots.txtSystem-triggeredBrowser-like UA

Directory facts

AI model training
Listed as training
Acts on behalf of user
No, system-triggered
Obeys directives
Yes, listed as obeying robots.txt

What is Diffbot?

Diffbot is a generic crawler. Diffbot is classified as a generic crawler that may fetch pages for search, SEO, monitoring, or data collection workflows.

Diffbot was inferred from a bot-like product token in the user-agent string.

How to identify Diffbot in logs

Search server logs for Diffbot. Matching those tokens is useful for discovery, but IP verification is still recommended before trusting the identity.

Diffbot

Mozilla/5.0 (compatible; Diffbot/1.0; +https://diffbot.com)

Use this as a web-crawler lookup reference for identifying how this user-agent presents itself in server logs. After you identify crawler traffic, run the AI Agent Readiness Scanner to confirm whether AI crawlers and agents can understand your site.

User-agent signals

Product tokens
Mozilla/5.0, Diffbot/1.0
Documentation
https://diffbot.com
Contact
None found
Platform
Unknown · Unknown
Browser profile
Unknown · Unknown
Browser-like UA
Yes
HTTP library
Unknown
Spoof risk
High

Questions answered by this crawler profile

What is Diffbot?

Diffbot is classified as a generic crawler that may fetch pages for search, SEO, monitoring, or data collection workflows.

Who operates Diffbot?

The operator for Diffbot is not known from the user-agent alone.

How do I identify Diffbot in logs?

Search server logs for Diffbot. Matching those tokens is useful for discovery, but IP verification is still recommended before trusting the identity.

Should I allow Diffbot?

Verify IP ownership or behavior before making security decisions because user-agent strings can be spoofed. Monitor crawl rate and paths, then allow normal traffic or rate-limit/block if behavior becomes abusive.

Does Diffbot respect robots.txt?

Robots.txt compliance cannot be proven from a user-agent string alone. Check the crawler operator documentation and your own logs before assuming behavior.