Effective-rule audit
Merge matching groups, honor case-sensitive paths and specificity, and show sampled access rather than static labels.
Audit the policy your SaaS serves today, see how documented search and AI crawlers reach real pages, and make explicit changes without guessing or erasing unrelated rules.
It allows public product, pricing, documentation, evidence, and render assets needed for discovery; blocks only paths the owner intentionally chooses; declares verified sitemaps; and never pretends crawl policy protects private customer data.
Merge matching groups, honor case-sensitive paths and specificity, and show sampled access rather than static labels.
Keep search, AI citation, training, extended-use, and user-triggered tokens as distinct owner decisions.
Preserve by default; allow only verified crawler tokens and validated literal path prefixes.
Require acknowledgment before files that can suppress discovery, rendering, or whole-site crawling are copied or downloaded.
Fetch only the root policy within strict network, UTF-8, byte, redirect, and parser boundaries.
Apply effective rules to representative SaaS and discovery URLs for each source-verified crawler token.
Preserve existing behavior or make explicit crawler and literal-path decisions with affected-URL previews.
Review critical warnings, download the exact output or scoped patch, then test it on the public host.
The matcher follows RFC 9309 and current Google robots.txt guidance. Crawler identifiers link to official OpenAI, Anthropic, Google, and Bing documentation with a last-verified date.
It audits the root policy, applies RFC 9309 matching, shows sampled access for source-verified search, AI-search, training, and user-triggered crawler tokens, then preserves the current file unless you explicitly choose a crawler or literal-path change.
Not reliably. robots.txt controls crawling, not authorization or guaranteed deindexing. A blocked URL can still be known from links. Use access control for private data and an indexable response carrying noindex when search removal is the goal.
That is a product and content-licensing decision, not one generic switch. The tool keeps AI citation/search agents, model-training agents, Google Search, Google-Extended, and user-triggered agents separate, using only tokens linked to current official provider documentation.
Replacing robots.txt can silently change host-wide crawling. The safe default retains comments, Allow and Disallow rules, Sitemap declarations, and unknown extension records. Explicit changes get a before-publish warning and representative URL review.
Yes, with case-sensitive literal path prefixes after analysis. The tool refuses full URLs, raw wildcards, query strings, fragments, encoded delimiters, control characters, and duplicates, and previews matching sampled URLs. A folder-scoped crawl returns only a review patch.
No. It reads authorized public responses and returns a private file, evidence, and repository-aware implementation prompt. You or your coding agent must review, test, and publish the policy.
Blogged turns product context into reviewed, interlinked SaaS content while keeping technical discovery implementation-ready.
Start free