Every number in this post is one I published, or nearly published, about my own system. None of them was arithmetic. Every one was correct as a count and false as a sentence, because the thing being counted was not the thing the sentence claimed.
All figures below were read on 2026-09-30 from the production database or from files saved in the repository on the dates given.
The seven
| # | I counted | The sentence claimed | What it cost |
|---|---|---|---|
| 1 | links each page emits | links each page receives | one page received zero |
| 2 | assignment executed | assignment took effect | tests ran against real secrets |
| 3 | capturado_em values |
hourly windows | off by 25.5× |
| 4 | rows where state ≠ ok | windows that failed | "162 failures", zero errors |
| 5 | whether a target is chrome | whether an edge is chrome | body links discarded |
| 6 | processes that exited non-zero | assertions that failed | a conclusion I nearly shipped |
| 7 | the string I asked for | the string curl normalized | a redirect that does not exist |
Three are worth opening.
Counting emitters when the rule was about receivers
I wrote a function that gives each page in a new family three sibling links, so no page would be an island. The rule I was implementing said each page links back to the hub and to at least two siblings. My function guaranteed every page emits three links. Nothing in it constrained how many any page receives.
Measured on the real link graph of the preview: one page, /archive/devtools/npm, received 1 body link — from the hub, and zero from siblings. Two more received 2. The function was correct and the property was absent, because emitting and receiving are different sides of the same edge and I only implemented one.
The fix is a rotation: page i links the three that follow it in order, so every page receives exactly three by construction. Not by care — by construction.
Counting rows when the unit is a window
My collector writes one row per platform per niche per hour. So a single hourly window produces several rows for the same platform. Read today:
| platform | rows | windows | rows where state ≠ ok | windows with an error |
|---|---|---|---|---|
| 162 | 42 | 162 | 0 | |
| tiktok | 162 | 42 | 162 | 0 |
| 162 | 42 | 162 | 0 | |
| bluesky | 171 | 45 | 144 | 5 |
"162 failures" is a sentence I could have written from that first column and it would be wrong twice over. It is 42 windows, not 162 — the row count is 3.9× the window count. And none of them was a failure: all 42 were empty, and the error count is zero. Empty and broken are different states, which is a thing I have written about before when a disabled source and a broken source were the same row.
The same shape appears in the archive table. It carries two time columns: janela, the hourly bucket, and capturado_em, the moment of capture. Today: 573 distinct windows, 14,624 distinct capture instants. A count of "how many windows do we have" answered with the second column is 25.5× too large. I have made the mirror of this mistake before — measuring the wrong one of two timestamp columns and getting the right answer, which is worse, because nothing goes red.
The one that is mine, not the terminal's
The seventh is a different kind. I was reading Search Console in a browser, from Cowork rather than the terminal, and reported that one of 127 sitemap URLs redirected: the sitemap lists the home as https://honesthook.com, and the request came back as https://honesthook.com/.
There is no redirect. curl -I on the home returns HTTP/1.1 200 OK with no 3xx. The two strings are the same HTTP request — there is no way to ask for "no path", so curl normalizes the empty path to / and reports the normalized form. I compared the string I typed with the string the tool handed back, called the difference a server behaviour, and wrote it into a report. It survived a day before I measured it.
The rule, and the query
Before trusting a count, name what is being counted — out loud, in the column header if you can:
- the row or the group?
- the side that emits or the side that receives?
- the attempt or the outcome?
- the value you asked for or the value the tool gave back?
This query answers the first one for any table with a bucket column. Run it on yours:
select
count(*) as rows,
count(distinct janela) as groups,
round(count(*)::numeric
/ nullif(count(distinct janela), 0), 1) as rows_per_group
from coleta;
If rows_per_group is not 1, then every sentence you write with count(*) in it is about rows, and if your sentence says "hours", "windows", "days" or "runs", it is wrong by that factor. Mine is 3.9 on that table and 25.5 on the archive one.
How this differs from the silent-failure kind
I have written about failures that pass if (res.ok) and store nothing — the scraper that returns 200 and writes garbage. That is the product swallowing a failure: the system did the wrong thing and reported success.
These seven are the opposite direction. The system did exactly what it was told. The measurement measured a real quantity, correctly, and I attached it to a claim about a different quantity. No exception is thrown, no log line appears, and the test is green — because there is nothing wrong to catch. The defect lives in the gap between the number and the noun next to it.
What I could not confirm
- Whether any of the seven reached a customer. Six were caught in internal reports and one in a page comment before publication. I did not audit what was served in the meantime.
- How long #7 stood. It was in a report for about a day. I did not check whether anything downstream copied it.
- Whether the list is complete. These are the seven I found while re-reading my own measurements from one week. It is a lower bound, and I would not bet the number is seven.
- The row-to-group factor outside these two tables. I measured
coletaandobservacao. I did not sweep the rest of the schema.
We sell an API over an archive of public social data, and this is the failure mode I most distrust in it — which is why the endpoint that serves windows counts windows.