HonestHook

Sign in

Glossary

robots.txt

robots.txt is a plain text file at the root of a domain that tells automated clients which paths they should not fetch. It is a convention, published in 1994 and still the closest thing the web has to a standard signal of intent.

It is not an access control, and it is not a contract.

What it actually is

User-agent: *
Disallow: /private/
Crawl-delay: 10

Three things worth understanding:

Where it sits legally

Breaching robots.txt is not in itself illegal, and the file is not a term of service you accepted. The distinct legal questions are the platform's terms of service — a contractual matter, and Meta v. Bright Data (2024) held that logged-off scraping of public data is not bound by terms you never agreed to — and data protection, which is separate again and covered under personal data.

Those three get conflated constantly, including by vendors citing one to imply another. They are independent: you can be in the clear on all three, or on none.

Why it still matters

Two practical reasons beyond principle:

The honest position

A data product's relationship with robots.txt is a real choice with real trade-offs, and the right thing is to state it rather than leave it implied. Our own reasoning about which sources we collect and which we leave alone — several of them declined for licence reasons rather than technical ones — is written source by source under sources.

Trend data with a memory

Every social API answers what’s trending now, then throws it away. HonestHook keeps the hourly archive, so you can ask what gained traction.

Free key, 1,000 credits a month, no card →