resulthack

All AI crawlers /

Diffbot

Diffbot · AI search · robots.txt: Not always

General, proactive web crawling for Diffbot’s search engine, so pages can be found and cited from the Diffbot Knowledge Graph and Diffbot’s web search services.

Block Diffbot

User-agent: Diffbot
Disallow: /

Blocking Diffbot refuses the crawl that lets pages be discovered and cited from the Diffbot Knowledge Graph and web search services. Diffbot says that crawl is not used for AI training.

Allow Diffbot

User-agent: Diffbot
Allow: /

An allow group only changes anything if a broader rule would block Diffbot. Under the robots.txt standard (RFC 9309, section 2.2.1) a crawler follows the group that names it and uses the User-agent: * group only when no group does — so this group lets it in even if your * group says Disallow: /. Token matching is case-insensitive.

Put these lines in /robots.txt at the root of each host (each subdomain has its own file). robots.txt is a request to well-behaved crawlers, not access control.

Facts

Operator
Diffbot
Purpose
AI search index

Diffbot says this crawl builds a general search engine, so sites can be discovered and cited from the Diffbot Knowledge Graph and its web search services, and that it is not used for AI training.

Obeys robots.txt
Not always. Diffbot says its web crawls follow robots.txt by default, including Disallow and Crawl-delay, but that in specific cases, typically a partnership or agreement with the site being crawled, the robots.txt instruction can be ignored or overridden. Crawls that Diffbot customers run on its Extract and Crawlbot APIs are set up by those customers, who are advised to send their own user-agent.
In your server logs
Unconfirmed by the operator Diffbot's documentation we checked does not give its user-agent string.
How to verify it
Unconfirmed by the operator Diffbot's documentation we checked gives no IP list or DNS check for this crawler, so a request claiming to be Diffbot cannot be verified against the operator.

What Diffbot says

Quoted word for word from the operator's documentation.

“General, proactive web crawling for building a general search engine. This allows websites to be discovered and cited from the Diffbot Knowledge Graph and web search services in response to keyword queries. It is not used for AI training.”

“By default Diffbot's web crawls adhere to a site’s robots.txt instructions, including the disallow and crawl-delay directives. In specific cases — typically because of a partnership or agreement you have with the site to be crawled — the robots.txt instruction can be ignored/overridden.”

“When users use Extract or Crawlbot APIs, they are defining their own crawl parameters using hosted software. The best practices would be to set a User-Agent for your crawl representing your organization and to enable robots.txt adherence feature in Crawlbot, which is toggled on by default.”

Sources

Other crawlers from Diffbot

Same purpose, other operators

← All 39 AI crawlers