← AI crawler directory
AI training
Omgilibot
Webz.ioHonours robots.txt
Collects web data resold as training and market-intelligence datasets. Data resale rather than a consumer surface. Common on training-opt-out lists.
User-agent string
Omgilibot/0.4 +http://omgili.com
How to verify it
Webz.io published ranges.
What blocking costs
Data resale rather than a consumer surface. Common on training-opt-out lists.
robots.txt rule
# Block Omgilibot
User-agent: Omgilibot
Disallow: /
# Allow everything else
User-agent: *
Allow: /Robots.txt is enough
Webz.io documents Omgilibot as honouring robots.txt. Verify with your logs after deploying — a spoofed agent will keep coming.
Generate a full robots.txt
Pick which AI crawlers to allow or block across all 42 agents and copy the file out.
Open the generator →