Glossary
robots.txt
robots.txt is a site-root file that tells compliant crawlers which URL paths they may request. It controls crawling rather than indexing by itself, and different search and AI providers publish distinct user agents that site owners can allow or disallow.
In plain terms
It is the public rule file that tells named bots where they may crawl.
Why it matters
A broad block can remove important content from retrieval, while allowing one AI crawler does not automatically allow every AI product.
How to apply it
- Review rules for search, training and user-triggered agents separately.
- Do not use robots.txt as the only method for removing indexed content.
- Test the exact public file after changes.
Example
A company allows a live-search crawler while blocking a separate training crawler according to the provider’s documented user agents.
Sources
Related reading
Related terms
Back to the full glossary (75 terms).
Ranking is no longer enough
You need to be cited, mentioned, and recommended.
Being cited, mentioned, and recommended are three different outcomes, and most brands only ever achieve the first one. Ranking is no longer enough because AI engines answer buyers directly and name only a short list of vendors as the recommendation — everyone else is cited in passing, if at all. Get a free AI Visibility Report to see exactly where your brand appears today across ChatGPT, Google AI Overviews, Gemini, Perplexity, and Copilot, where competitors are winning the recommendation instead, and what's keeping you from moving up the shortlist.