AIWebIndex
Lyrenth · AI search · robots.txt: Yes
Crawls the public web to build Lyrenth's AI-readable web index, which serves each page to AI agents as one clean, attributed document.
Block AIWebIndex
User-agent: AIWebIndex
Disallow: /
Unconfirmed by the operator Lyrenth's documentation we checked does not say what blocking AIWebIndex changes beyond the crawler itself.
Allow AIWebIndex
User-agent: AIWebIndex
Allow: /
An allow group only changes anything if a broader rule would block AIWebIndex. Under the robots.txt standard (RFC 9309, section 2.2.1) a crawler follows the group that names it and uses the User-agent: * group only when no group does — so this group lets it in even if your * group says Disallow: /. Token matching is case-insensitive.
Put these lines in /robots.txt at the root of each host (each subdomain has its own file). robots.txt is a request to well-behaved crawlers, not access control.
Already have a robots.txt? Paste it into the robots.txt checker to see whether it blocks AIWebIndex on a given path, and which line decides.
Who blocks it
AIWebIndex is not in the 2026-10-09 survey: it was added to this directory after we last read the most-visited sites' robots.txt files, so we have no count for it yet.
Facts
- Operator
- Lyrenth
- Purpose
- AI search index
Lyrenth says AIWebIndex builds a search index of the public web that it serves to AI agents as one clean document per page, and that it does not train foundation models on what it crawls. Fetches a customer asks for by URL come from AIWebIndex-Agent, which Lyrenth says honours robots.txt the same way.
- Obeys robots.txt
- Yes. Lyrenth says the crawler implements RFC 9309 and honours Disallow, Crawl-delay and Sitemap, backs off on HTTP 429/503, fetches at most once every 2 seconds per domain, and fails closed when robots.txt cannot be read.
- In your server logs
- The user-agent string starts "AIWebIndex/2.0"; on-demand fetches a customer asked for start "AIWebIndex-Agent/2.0", and ownership checks read "AIWebIndex/2.0 verification".
- How to verify it
- Lyrenth gives three checks: the source IP is on its published list, the request carries a Web Bot Auth signature (RFC 9421), and the IP has forward-confirmed reverse DNS under lyrenth.com.
Published list: www.lyrenth.com/bot/ip-ranges.json — 22 IPv4 ranges when we fetched it, list dated 16 Aug 2026. Always match against the live list; operators update them. To count real and fake AIWebIndex requests in your own server log against this list, drop the log into the log reader.
What Lyrenth says
Quoted word for word from the operator's documentation.
“The simplest opt-out is your robots.txt: User-agent: AIWebIndex Disallow: /”
“AIWebIndex/2.0 is the crawler that builds Lyrenth, the AI-readable web index: a search index of the public web that serves every page as one clean, attributed AIDocument.”
“It is an indexer, not a model trainer: we do not train foundation models on the content we crawl.”
“Lyrenth, operated by Aleksma Ai, Inc., is the reference commercial implementation.”
“It identifies itself honestly and is verifiable three ways on this page, honors RFC 9309 robots.txt including Crawl-delay, keeps per-domain rate limits, and fails closed when robots.txt cannot be read.”
“Our crawler implements RFC 9309 and honors Disallow rules, Crawl-delay, HTTP 429/503 backoff, and Sitemap directives.”
“On-demand index fetches (a customer asked the index for that exact URL) identify separately so your logs can tell the two apart; both honor robots.txt the same way:”
“Verification probes use a slightly different UA so you can distinguish crawl fetches from ownership checks:”
“Confirm a request is genuinely ours three ways: 1 Source IP is on our published list at /bot/ip-ranges.json.”
“2 Requests carry a Web Bot Auth signature (RFC 9421); our public key directory is at api.lyrenth.com/.well-known/http-message-signatures-directory.”
“Rate: at most one fetch every 2 seconds per domain across our whole fleet, or your Crawl-delay when it is longer.”
“3 Every IP has forward-confirmed reverse DNS under lyrenth.com.”
Sources
- Lyrenth: Bot identification (AIWebIndex/2.0) — fetched and checked 2026-10-09
- Lyrenth: AIWebIndex IP ranges (ip-ranges.json) — fetched and checked 2026-10-09