Control which AI crawlers
access your applications.
Scalix Shield identifies the AI crawlers and bots reaching your deployed applications and lets you allow, block, or rate-limit each one. Per crawler, per category, or with custom rules. Enforced at the gateway, with no code changes required.
49 known crawlers
Built-in database of 43 AI crawlers and 6 search engines across 6 categories, maintained from operators’ published documentation.
Allow / block / rate-limit
Set a policy per crawler or per category. Rate limits are enforced per minute at the gateway.
IP verification
Crawler identities are validated against operator-published IP ranges, refreshed automatically. Spoofed user-agents are detected.
Generated robots.txt
Your policies compile to a robots.txt you can serve, so compliant crawlers see your rules before requesting a page.
See who's reading your site.
Shield identifies every crawler that touches your applications and judges it against your policy in real time. Identities are verified against operator-published IP ranges, a spoofed user-agent doesn't get Googlebot's privileges.
Set a policy. Ship a robots.txt.
Choose what each crawler may do. Shield enforces the decision at the gateway in real time and compiles your policy into a robots.txt that compliant crawlers read before requesting a page.
Interactive demo, your real policies are managed in the console, per project.
Six crawler categories
Search engines
Googlebot, Bingbot, Applebot and others. Typically allowed to preserve search visibility.
AI assistants
Crawlers that fetch pages in response to a user’s question, such as ChatGPT-User and PerplexityBot.
AI coding agents
Agents that read documentation and repositories while writing code.
AI training scrapers
Crawlers that collect content for model training. Allow, block, or rate-limit each one.
AI data providers
Companies that resell crawled content as datasets.
AI agents
Autonomous agents acting on a user’s behalf.
Custom rules and full visibility
Define your own user-agent rules for crawlers not yet in the database. They take precedence over built-in policies. Every decision is logged: which crawlers reached you, how often, and the action Shield took, in the console or via the API.