HonestHook

Sign in

Glossary

Web scraping

Web scraping is extracting data from a web page instead of from an API. It exists because the data you want is often visible to any visitor and absent from any official endpoint.

The technique matters more than the label. Four approaches, in descending order of reliability:

1. Embedded structured data

Many pages ship their own data as JSON inside the HTML — application/ld+json, a hydration blob, an Apollo cache. This is the best target available: it is the data the page itself renders from, it has named fields, and it changes far less often than markup.

2. The page's own JSON API

Some sites fetch their content from an internal JSON endpoint the browser can see. Reading that directly is cleaner than parsing HTML, though it is the surface most likely to require a login.

3. DOM parsing

Selecting elements by structure or class. Workable, but class names change without notice and a redesign breaks everything at once.

4. Regex on raw HTML

The worst option, and it fails in a specific way: it matches the first thing that looks right. On one platform we measured, the string follower_count appears 42 times on a profile page, and most of those are board counters, not the profile's. A regex takes whichever came first and returns a plausible number from the wrong object — a silent failure.

What makes scraping unreliable

Obstacle What it does
Login walls Puts the data behind terms you would have to accept
Client fingerprinting Returns a page without the data to non-browser clients
Rounded display values Shows 83K where the embedded JSON holds 83754
200-with-no-data Serves an error page with a success status

That last one is why every parse needs an anchor: a field that must be present for the document to be what you think it is. Without one you will extract data from a not-found page and never notice.

Bandwidth is the cost nobody prices

A profile read on a small protocol is under a kilobyte. The same read on a large commercial platform can be 270 KB decoded. That difference is the actual cost of scraping at volume, and we worked it out per platform.

Legality is a separate question from technique — see personal data for the part that survives even when the terms-of-service argument is won.

Trend data with a memory

Every social API answers what’s trending now, then throws it away. HonestHook keeps the hourly archive, so you can ask what gained traction.

Free key, 1,000 credits a month, no card →