Name the user agent
Target a documented user agent instead of relying on broad labels or assumptions about every AI system.
robots.txt · AI crawler reference
robots.txt is a crawler-access protocol. Use documented user agents and directives, then keep the intended crawler and purpose explicit.
Direct answer
Robots.txt can communicate access preferences to documented crawlers, but it is not a universal AI-search switch or a guarantee of crawling, indexing, inclusion, or citation.
Reference guide
Target a documented user agent instead of relying on broad labels or assumptions about every AI system.
Robots.txt tells compliant crawlers which paths they may request; it is not a ranking or source-management control.
Test directives against the documented user agent, confirm the file is reachable and valid, and check server logs for the crawler you intended to address. A directive only expresses a preference; verification shows whether compliant crawlers honor it in practice.
Changes to access preferences do not establish whether a page will be crawled, indexed, included, or cited. Access is a precondition you can inspect; answer behavior is a provider decision you can only observe.
Primary sources
Google Search Central. Platform guidance can change; recheck the source before acting.
OpenAI. Platform guidance can change; recheck the source before acting.
Related guides
Return to the educational GEO reference hub.
Identify documented crawlers before setting controls.
Keep the emerging proposal separate from robots controls.
Questions
No. It is an access-control convention, not an outcome guarantee. Outcomes belong to each provider’s systems and cannot be forced from a directives file.
Have a real research question?
Use supported workflows and keep their boundaries visible.
Start free trial