robots.txt is a plain text file at the root of a domain that tells automated clients which paths they should not fetch. It is a convention, published in 1994 and still the closest thing the web has to a standard signal of intent.
It is not an access control, and it is not a contract.
What it actually is
User-agent: *
Disallow: /private/
Crawl-delay: 10
Three things worth understanding:
- It is advisory. Nothing enforces it. A client that ignores it still receives the page.
- It is not authentication. Anything genuinely private needs a login, not a
Disallowline. Listing a path in robots.txt arguably advertises it. - It expresses intent. Which is precisely why ignoring it is hard to defend afterwards, whatever the legal position.
Where it sits legally
Breaching robots.txt is not in itself illegal, and the file is not a term of service you accepted. The distinct legal questions are the platform's terms of service — a contractual matter, and Meta v. Bright Data (2024) held that logged-off scraping of public data is not bound by terms you never agreed to — and data protection, which is separate again and covered under personal data.
Those three get conflated constantly, including by vendors citing one to imply another. They are independent: you can be in the clear on all three, or on none.
Why it still matters
Two practical reasons beyond principle:
Crawl-delayis information. A site telling you its tolerated interval has told you something useful about its rate limits before you discover them by being blocked.- It is the cheapest signal of how a site will treat you. A domain that disallows everything and fingerprints its clients is not going to become friendlier at volume — see TLS fingerprinting for what that looks like in practice.
The honest position
A data product's relationship with robots.txt is a real choice with real trade-offs, and the right thing is to state it rather than leave it implied. Our own reasoning about which sources we collect and which we leave alone — several of them declined for licence reasons rather than technical ones — is written source by source under sources.