Someone typed this into Google and landed on a page of mine, at position 91: "our scrapers keep breaking every time instagram or tiktok updates their platform - what's a better solution?" One impression, zero clicks, measured 28/09/2026. That is one person's sentence, not a keyword. I am answering it because I measured the same situation in my own database this month, and the measurement says the premise is usually wrong: the platform update is rarely what broke.
What my own collector recorded
I run an hourly collector over public platforms. Every run writes one row per platform into a table called coleta, carrying a state: ok, vazio ("ran, found nothing"), or erro ("the source said it failed"). Counted 29/09/2026:
| platform | ok | vazio | erro | first row | last row |
|---|---|---|---|---|---|
| 0 | 162 | 0 | 07/09 | 09/09 | |
| tiktok | 0 | 162 | 0 | 07/09 | 09/09 |
| 0 | 162 | 0 | 07/09 | 09/09 | |
| bluesky | 27 | 131 | 13 | 07/09 | 12/09 |
| hackernews | 1,938 | 21 | 0 | 07/09 | 29/09 |
| lemmy | 1,346 | 12 | 0 | 07/09 | 29/09 |
| lobsters | 1,953 | 6 | 0 | 07/09 | 29/09 |
| youtube | 1,954 | 5 | 0 | 07/09 | 29/09 |
| cratesio, devto, github, npm, vscode | 1,959 each | 0 | 0 | 07/09 | 29/09 |
Read that table the way the log invites you to. Instagram, TikTok and Reddit each ran 162 times and found nothing: three platforms, three problems, zero errors anywhere. That is exactly the story the query tells — the platforms changed, the parsers rotted, here we are.
It is not what happened.
The three stopped in the same window
A second table, lacuna ("gap"), records known holes in the archive with their reason. Instagram, TikTok and Reddit each have exactly one row, and all three carry the same desde: 2026-09-06 23:00:00+00. The reason field, translated from my original Portuguese:
"Source produced in the 22:00 window of 06/09 and nothing after. The three sources from the same supplier stopped in the same window, which indicates a single cause at the supplier and not at the platform. No record of an attempt in the period: collection instrumentation only came into existence on 07/09."
Three platforms do not change their layout in the same hour. One supplier has one outage in one hour. Those 486 rows of vazio describe a single event, and every one of them names the wrong culprit.
Why the log could not say so
The supplier adapter, on any non-OK response — including the 402 meaning "no credit left" — logged to the console and returned an empty list. An empty list is a successful return. The promise resolves. The Promise.allSettled upstream swallowed nothing, because no rejection ever existed. The collector saw zero rows, and zero rows with no error attached is, by its own definition, vazio.
Your scrapers may be fine. What is broken is that a failure and a quiet afternoon are the same row.
What is fixed, and what is not (29/09/2026)
The decision point now reads a diagnostic off the response instead of guessing from the row count:
const d = itens?.diagnostico ?? null;
const falhou = d && d.termos_falhos > 0;
coletas.push({ /* ... */ estado: falhou ? "erro" : "vazio" });
A real fix, and a partial one. I counted the adapter files: 16, and exactly one publishes diagnostico today. The other fifteen still return an empty list when they fail, so for those the collector cannot tell silence from emptiness. The root fix lives in the adapter layer and is written down as owed, not done.
Bluesky is the same shape with a different swallow — a burst of 403s consumed by an if (!resposta.ok) continue, which is what produced its 131/13 split above. I wrote that one up on its own; the point here is only that the same question found it.
The maintenance the question is actually about
The part of the query that is real: platforms do change, and each identifies a profile its own way. My endpoint catalog has 14 entries, and the field recording what anchors a profile holds 12 distinct values across those 14 (counted 29/09/2026). did on Bluesky, userInfo on Instagram, __sc_hydration on SoundCloud, ld+json ProfilePage on Pinterest — and on TikTok, a note that the challenge page returns HTTP 200 with no data in it at all.
Twelve answers to one question is the maintenance surface, and no amount of better scraper code makes it smaller.
Four options, and we are not the first one
1. Instrument what you already have. Make failure and emptiness different rows, with the reason attached. Where it fails: it removes nothing. The 486 rows still get written; they just stop lying. It is the cheapest option and prerequisite for the other three — you cannot choose between them while your log files a supplier outage as a quiet platform.
2. Use the official API, where one is open. Say it out loud: if the platform you need publishes a documented, open API on terms you can live with, use it. You need no supplier there, and you do not need us. Where it fails: coverage, only coverage. Reddit's terms changed in 2026, X's free tier is what it is, and the public surfaces of Instagram and TikTok were never in this category.
3. Pay a supplier. One integration, one response shape, someone else absorbing layout changes. Where it fails: the failure does not disappear, it moves behind a boundary you cannot see through and cannot fix. That is not hypothetical: it produced the 486 rows above, on our own infrastructure, invisible to us for three days. We are option 3 for our customers and also option 3's customer, so we get to fail from both sides.
4. Keep an archive. Store what you collected and serve history instead of re-fetching. Where it fails: an archive answers "what was", never "what is now" — and it is worth exactly as much as its gaps are honest, which lands you back on option 1.
Run this on your own collector
select plataforma,
min(janela) filter (where estado <> 'ok') as primeira_falha,
count(*) filter (where estado = 'vazio') as vazios,
count(*) filter (where estado = 'erro') as erros
from coleta
group by plataforma
order by primeira_falha;
(Mine are Portuguese: plataforma platform, janela the hourly window, estado the state.) If two or more platforms share a primeira_falha hour, stop reading it as N platform problems. And if a platform shows a large vazios with erros at exactly zero, distrust the zero: it is likelier that your adapter cannot report failure than that a source failed 162 times politely.
What we couldn't confirm
- The supplier's side. Every number here comes from our own tables. We never saw their status page, ticket or log, and cannot say what the outage was.
- Two of our own dates disagree. The
lacunarow says the last production was the 22:00 window of 06/09. A separate alert, opened 15/09 and closed 16/09, says the three sources "began failing in the same run (07/09 21:00)". Both are our own records, about 23 hours apart, and we have not reconciled them — nor are we going to quietly pick the prettier one. - "Same supplier" is inferred, not recorded. The alert says so in its own text: there is no supplier field in the source layer yet. The grouping came from the synchrony of the stop.
- Nothing before 07/09 was measured. Instrumentation only started that day, so "162 empty runs" counts what was recorded, not what was attempted.
- The 12 distinct anchors are our catalog's, not the platforms'. They measure how many ways we identify a profile — not how often any platform changes.
- Whether the fix holds. One adapter of 16 publishes a diagnostic today; the other fifteen are untested against this failure, because none has failed this way since.
We sell the third option — one key, ten platforms, one response shape — and the second option is genuinely better wherever it exists.