Content and technology

What is robots.txt?

A file that tells automated crawlers which paths on a website they may access. It manages crawling; it does not guarantee search visibility or security for a page.

Definition

robots.txt is a plain-text file at the root of a website that tells automated crawlers which paths they may and may not fetch. Rules can be written per crawler, so a site can allow search crawlers while restricting training crawlers.

It manages crawling, nothing more. It does not protect a page from being seen, does not affect ranking, and a rule left over from years ago can quietly block an AI platform that did not exist when it was written.

Why it matters

A single Disallow line can decide whether a platform ever reads your site. Reviewing it takes minutes and prevents the most expensive kind of invisibility.

How it works

Read the file for rules that name AI crawlers or apply to all user agents, confirm the paths of your key pages are allowed, and keep the file in sync with your policy on search versus training access.

See where your brand stands in AI search Perplexity ChatGPT Claude Gemini Microsoft Copilot DeepSeek Google AI Overviews

A first reading from one website address, with the questions behind the number.