Robots.txt Guide for Beginners: What It Is and How to Use It Safely

Robots.txt Guide for Beginners: What It Is and How to Use It Safely

The robots.txt file is a small text file with a big responsibility: it tells search engine crawlers which parts of your site they may or may not visit. Used correctly, it keeps crawlers focused. Used incorrectly, it can hide your whole site from Google.

What Is Robots.txt?

Robots.txt is a plain text file placed at the root of your domain (yourdomain.com/robots.txt) that gives instructions to web crawlers. It follows the Robots Exclusion Protocol, which major search engines respect.

Basic Syntax

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://yourdomain.com/sitemap_index.xml
  • User-agent – which crawler the rules apply to. * means all crawlers.
  • Disallow – paths crawlers should not visit.
  • Allow – exceptions inside a disallowed folder.
  • Sitemap – the location of your XML sitemap.
Robots.txt Rules to Remember
Robots.txt Rules to Remember

Robots.txt Controls Crawling, Not Indexing

This is the most misunderstood point. Blocking a URL in robots.txt stops Google from crawling it, but Google may still index the URL (without its content) if other pages link to it. To keep a page out of search results, use a noindex meta tag instead — and do not block that page in robots.txt, or Google will never see the noindex tag.

WordPress and Robots.txt

WordPress creates a virtual robots.txt automatically, which is fine for most blogs. SEO plugins like Yoast SEO let you edit it through Yoast SEO → Tools → File editor. A simple, safe version for most WordPress blogs is the example shown above.

What You Should Not Block

  • CSS and JavaScript files – Google needs them to render and understand your pages.
  • Your uploads folder – otherwise your images can’t appear in Google Images.
  • Pages you want to rank – obvious, but accidental blocks are common.

When Blocking Makes Sense

  • Internal search result pages (for example Disallow: /?s=) on large sites
  • Admin and login areas
  • Endless filter or parameter URLs on large e-commerce sites

Small blogs rarely need anything beyond the default rules.

The Most Dangerous Line

User-agent: *
Disallow: /

This blocks crawlers from your entire site. It is sometimes left over from development or staging sites. If your traffic disappears after a launch or migration, check robots.txt first.

How to Check Your Robots.txt

  1. Visit yourdomain.com/robots.txt in a browser.
  2. Use Google Search Console’s robots.txt report (under Settings) to see the version Google fetched and any errors.
  3. Use the URL Inspection tool to confirm important pages are “allowed”.

Final Thoughts

For most blogs, robots.txt should be short and simple: block the admin area, allow everything else and point to your sitemap. When in doubt, block less, not more. Pair it with a clean XML sitemap for best results.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *