HonestHook

Sign in

Blog ·

Three platforms stopped in the same hour, and my log blamed three platforms

Someone typed this into Google and landed on a page of mine, at position 91: "our scrapers keep breaking every time instagram or tiktok updates their platform - what's a better solution?" One impression, zero clicks, measured 28/09/2026. That is one person's sentence, not a keyword. I am answering it because I measured the same situation in my own database this month, and the measurement says the premise is usually wrong: the platform update is rarely what broke.

What my own collector recorded

I run an hourly collector over public platforms. Every run writes one row per platform into a table called coleta, carrying a state: ok, vazio ("ran, found nothing"), or erro ("the source said it failed"). Counted 29/09/2026:

platform ok vazio erro first row last row
instagram 0 162 0 07/09 09/09
tiktok 0 162 0 07/09 09/09
reddit 0 162 0 07/09 09/09
bluesky 27 131 13 07/09 12/09
hackernews 1,938 21 0 07/09 29/09
lemmy 1,346 12 0 07/09 29/09
lobsters 1,953 6 0 07/09 29/09
youtube 1,954 5 0 07/09 29/09
cratesio, devto, github, npm, vscode 1,959 each 0 0 07/09 29/09

Read that table the way the log invites you to. Instagram, TikTok and Reddit each ran 162 times and found nothing: three platforms, three problems, zero errors anywhere. That is exactly the story the query tells — the platforms changed, the parsers rotted, here we are.

It is not what happened.

The three stopped in the same window

A second table, lacuna ("gap"), records known holes in the archive with their reason. Instagram, TikTok and Reddit each have exactly one row, and all three carry the same desde: 2026-09-06 23:00:00+00. The reason field, translated from my original Portuguese:

"Source produced in the 22:00 window of 06/09 and nothing after. The three sources from the same supplier stopped in the same window, which indicates a single cause at the supplier and not at the platform. No record of an attempt in the period: collection instrumentation only came into existence on 07/09."

Three platforms do not change their layout in the same hour. One supplier has one outage in one hour. Those 486 rows of vazio describe a single event, and every one of them names the wrong culprit.

Why the log could not say so

The supplier adapter, on any non-OK response — including the 402 meaning "no credit left" — logged to the console and returned an empty list. An empty list is a successful return. The promise resolves. The Promise.allSettled upstream swallowed nothing, because no rejection ever existed. The collector saw zero rows, and zero rows with no error attached is, by its own definition, vazio.

Your scrapers may be fine. What is broken is that a failure and a quiet afternoon are the same row.

What is fixed, and what is not (29/09/2026)

The decision point now reads a diagnostic off the response instead of guessing from the row count:

const d = itens?.diagnostico ?? null;
const falhou = d && d.termos_falhos > 0;
coletas.push({ /* ... */ estado: falhou ? "erro" : "vazio" });

A real fix, and a partial one. I counted the adapter files: 16, and exactly one publishes diagnostico today. The other fifteen still return an empty list when they fail, so for those the collector cannot tell silence from emptiness. The root fix lives in the adapter layer and is written down as owed, not done.

Bluesky is the same shape with a different swallow — a burst of 403s consumed by an if (!resposta.ok) continue, which is what produced its 131/13 split above. I wrote that one up on its own; the point here is only that the same question found it.

The maintenance the question is actually about

The part of the query that is real: platforms do change, and each identifies a profile its own way. My endpoint catalog has 14 entries, and the field recording what anchors a profile holds 12 distinct values across those 14 (counted 29/09/2026). did on Bluesky, userInfo on Instagram, __sc_hydration on SoundCloud, ld+json ProfilePage on Pinterest — and on TikTok, a note that the challenge page returns HTTP 200 with no data in it at all.

Twelve answers to one question is the maintenance surface, and no amount of better scraper code makes it smaller.

Four options, and we are not the first one

1. Instrument what you already have. Make failure and emptiness different rows, with the reason attached. Where it fails: it removes nothing. The 486 rows still get written; they just stop lying. It is the cheapest option and prerequisite for the other three — you cannot choose between them while your log files a supplier outage as a quiet platform.

2. Use the official API, where one is open. Say it out loud: if the platform you need publishes a documented, open API on terms you can live with, use it. You need no supplier there, and you do not need us. Where it fails: coverage, only coverage. Reddit's terms changed in 2026, X's free tier is what it is, and the public surfaces of Instagram and TikTok were never in this category.

3. Pay a supplier. One integration, one response shape, someone else absorbing layout changes. Where it fails: the failure does not disappear, it moves behind a boundary you cannot see through and cannot fix. That is not hypothetical: it produced the 486 rows above, on our own infrastructure, invisible to us for three days. We are option 3 for our customers and also option 3's customer, so we get to fail from both sides.

4. Keep an archive. Store what you collected and serve history instead of re-fetching. Where it fails: an archive answers "what was", never "what is now" — and it is worth exactly as much as its gaps are honest, which lands you back on option 1.

Run this on your own collector

select plataforma,
       min(janela) filter (where estado <> 'ok')  as primeira_falha,
       count(*)    filter (where estado = 'vazio') as vazios,
       count(*)    filter (where estado = 'erro')  as erros
from coleta
group by plataforma
order by primeira_falha;

(Mine are Portuguese: plataforma platform, janela the hourly window, estado the state.) If two or more platforms share a primeira_falha hour, stop reading it as N platform problems. And if a platform shows a large vazios with erros at exactly zero, distrust the zero: it is likelier that your adapter cannot report failure than that a source failed 162 times politely.

What we couldn't confirm

We sell the third option — one key, ten platforms, one response shape — and the second option is genuinely better wherever it exists.

Trend data with a memory

Every social API answers what’s trending now, then throws it away. HonestHook keeps the hourly archive, so you can ask what gained traction.

Free key, 1,000 credits a month, no card →