Answer: The Robots.txt Generator produces your output instantly from the input you provide — everything runs in your browser, free, with no signup required.
Generate crawler directives for Googlebot, Bingbot, GPTBot & more
This robots.txt generator builds a valid robots.txt file from your rules — which crawlers may or may not fetch which paths — with live preview and one-click copy. It runs entirely in your browser.
robots.txt is the oldest and simplest crawler control on the web: a plain-text file at the site root that well-behaved crawlers read before fetching pages. Simple, but with sharp edges worth knowing.
The file groups rules by user agent (crawler). Each group contains Allow and/or Disallow path prefixes, and crawlers pick the most specific matching group for their name. A bare Disallow: (empty value) permits everything; Disallow: / blocks the entire site for that agent.
Google's documented precedence: the most specific (longest) matching path wins, so 'Allow: /folder/page' beats 'Disallow: /folder'. A robots.txt is advisory for compliant crawlers — it blocks fetching, not indexing (a page linked elsewhere can still appear in results without content; to keep pages out of indexes entirely, use noindex meta tags or authentication).
The table shows the effect of canonical directives exactly as crawlers interpret them. Note the Crawl-delay line: Google ignores it, Bing honors it as milliseconds historically (seconds in some eras) — treat it as a politeness hint, not a throttle you can rely on.
Also common: blocking crawl traps (faceted search URLs like /search? or calendar pages that generate infinitely), staging folders, and API endpoints with crawl budget implications. Sitemap declarations in robots.txt help discovery: a 'Sitemap:' line is read by all major engines.
| Directive(s) | Effect |
|---|---|
| User-agent: * Disallow: | Allows all crawlers everywhere |
| User-agent: * Disallow: / | Blocks entire site for all crawlers |
| Disallow: /admin/ | Blocks the /admin/ directory |
| Disallow: /*?s= | Blocks search query URLs (wildcard) |
| Allow: /folder/page$ Disallow: /folder/ | Permits one page, blocks rest of folder |
| Sitemap: https://example.com/sitemap.xml | Declares sitemap location |
On large sites, robots.txt is a crawl-budget tool: Googlebot allocates limited fetch capacity per site, and wasting it on parameter permutations or duplicate paths delays indexing of pages you care about. Blocking low-value URL spaces (filters, sorts, internal search) redirects that budget to meaningful pages.
Measure before and after: Google Search Console's Crawl Stats and Pages reports show fetch volume and discovered-vs-indexed counts. Change robots.txt gradually — an over-broad Disallow silently deindexes a whole section, and Google caches robots.txt for up to 24 hours, so mistakes linger.
The file must be served at https://example.com/robots.txt (site root, lowercase name) with a 200 status and text/plain type. Subdomains each need their own. Comments start with #. Rule of thumb: keep it small — every major engine reads only the first few hundred kilobytes.
Validate before deploying: Google Search Console has a robots.txt tester that checks rule syntax and lets you test URLs against the file, and the robots exclusion protocol is documented at robots-txt.com and by each engine. This generator emits only syntactically valid directives in standard order.
["Wildcards (*) and end-anchors ($) exist in the robots exclusion extensions Google supports: 'Disallow: /*.pdf$' blocks all PDFs, 'Disallow: /*?' blocks every parameterized URL. Use them surgically — an anchored rule that's too broad is invisible until traffic data shows a section gone.", "AI crawlers are the new neighbors: GPTBot, ClaudeBot, CCBot, and Google-Extended each honor robots.txt under their own user-agent names, so opting in or out of AI training crawls is a per-agent decision made in this same file. Blocking an agent that doesn't exist yet is harmless; the file is read fresh on every major crawl.", "Finally, robots.txt is public — anyone can read your Disallow list, so never use it to hide sensitive paths (that's a directory for attackers). Sensitive areas belong behind authentication. And after any deploy, fetch /robots.txt directly in a browser to confirm it's served as expected; a broken rewrite rule serving HTML where the text file should be quietly drops all rules for compliant crawlers."]
Where do I put robots.txt?
At the root of your domain: https://example.com/robots.txt — served with HTTP 200 and text/plain content type. Subdomains each require their own file; paths like /blog/robots.txt are ignored.
Does robots.txt block a page from Google?
It blocks crawling, not indexing. A URL blocked in robots.txt can still appear in results (without content) if other pages link to it. To fully exclude a page, allow crawling and use a noindex meta tag or X-Robots-Tag header.
How do I block all bots except Google?
Write a permissive group for Googlebot (User-agent: Googlebot with Disallow: of only what's needed), then a restrictive group for everything else: User-agent: * with Disallow: /. Specific agents pick their own group; * applies to the rest.
What does Disallow: / do?
It blocks the entire site for the user agents in that group — the nuclear option. Used intentionally for staging servers and internal tools; accidentally it deindexes a live site's visibility in short order.
Do I need a robots.txt if I want everything crawled?
No — an absent file means 'no restrictions' (served as 404). Many sites still ship one to declare their sitemap URL and manage crawl budget on large URL spaces.