Control which AI crawlers
access your applications.
Scalix Shield identifies the AI crawlers and bots reaching your deployed applications and lets you allow, block, or rate-limit each one — per crawler, per category, or with custom rules. Enforced at the gateway, with no code changes required.
49 known crawlers
Built-in database of 43 AI crawlers and 6 search engines across 6 categories, maintained from operators’ published documentation.
Allow / block / rate-limit
Set a policy per crawler or per category. Rate limits are enforced per minute at the gateway.
IP verification
Crawler identities are validated against operator-published IP ranges, refreshed automatically — spoofed user-agents are detected.
Generated robots.txt
Your policies compile to a robots.txt you can serve, so compliant crawlers see your rules before requesting a page.
See who's reading your site.
Shield identifies every crawler that touches your applications and judges it against your policy in real time. Identities are verified against operator-published IP ranges — a spoofed user-agent doesn't get Googlebot's privileges.
Set a policy. Ship a robots.txt.
Choose what each crawler may do. Shield enforces the decision at the gateway in real time and compiles your policy into a robots.txt that compliant crawlers read before requesting a page.
Interactive demo — your real policies are managed in the console, per project.
Six crawler categories
Search engines
Googlebot, Bingbot, Applebot and others — typically allowed to preserve search visibility.
AI assistants
Crawlers that fetch pages in response to a user’s question, such as ChatGPT-User and PerplexityBot.
AI coding agents
Agents that read documentation and repositories while writing code.
AI training scrapers
Crawlers that collect content for model training. Allow, block, or rate-limit each one.
AI data providers
Companies that resell crawled content as datasets.
AI agents
Autonomous agents acting on a user’s behalf.
Custom rules and full visibility
Define your own user-agent rules for crawlers not yet in the database — they take precedence over built-in policies. Every decision is logged: which crawlers reached you, how often, and the action Shield took — in the console or via the API.