HonestHook

Free API key (1,000/mo)Sign in

Blog ·

The flag that prices my API also decides whether my cron can run

I sell an API over public social data. Inside its database there is a catalog table called endpoint: one row per route, with a price in credits and a boolean called ativo. A migration on 2026-09-07 wrote down exactly what that boolean means, as a column comment:

Only true when an HTTP route is answering. The catalog feeds the public price table. Active without a route becomes a false promise.

That is a reasonable rule, and it is a catalog rule. ativo answers: may this appear on the price page. Two of those rows are not routes at all. They are the internal billing legs of a scheduled job, and both are ativo = false, which is correct under that rule, because they have no HTTP door. (The migration I quoted set the rule and the comment. It deactivated a different row, and I cannot date when these two were set, for the reason under What I am not claiming.)

The scheduled job reads the same column as a precondition for running.

All readings below were first taken from the production database on 2026-10-01, between 13:19 and 13:26 UTC, and re-run on 2026-10-02 against the same window. They reproduced. Where the catalog has moved since, I say so and give the value I read on 2026-10-02. (My code comments are in Portuguese; the quotes are my translations.)

The first statement in the job

select e.creditos into v_custo_orq
from public.endpoint e
where e.id = 'monitor/check' and e.ativo;

if not found then
  return jsonb_build_object('erro', 'orquestracao_indisponivel',
    'nota', 'endpoint monitor/check ausente ou inativo -- nenhum monitor rodou');
end if;

monitor/check is ativo = false. So the function returns that object before reaching its loop. The note it returns is accurate and says so plainly: no monitor ran.

The job is on pg_cron, enabled, and its schedule today is 1-59/5 * * * *. Up to 2026-10-01 13:21 UTC it had run 7,143 times; the first run recorded in cron.job_run_details is from 2026-09-06 18:10 UTC, which is not on today's schedule, so the schedule itself has changed at least once. Runs that did not report succeeded: zero. On 2026-10-02 the count was 7,546, and still zero.

Why zero

Because the function returns its error instead of raising it. A select over a function that hands back a jsonb object describing a failure is still a successful select. cron.job_run_details records status = succeeded and return_message = 1 row. That is true at every layer it passes through, and it is useless.

The audit trail does not help either. The job's own execution table carries this comment: "one row per check attempted, charged or not. It is the statement that makes the billing auditable." It has zero rows. An error returned before the loop writes nothing, so the absence of rows is exactly what a healthy idle job looks like.

So three independent places agreed that nothing was wrong: the cron status, the audit table, and the alert table, which had no open alert. Nothing disagreed, because nothing was asked. (The one alert built for a silent monitor watches a different system; more on that below.) I have written before about a billing reconciliation that compares zero rows and reports zero problems; this is the same silence, one layer up, in the scheduler.

Two different things called monitor

The confusion has a second half, and it is why I did not notice from the outside. There are two subsystems in this project with the word monitor in them.

One is the scheduled job above. The catalog gates it with a second boolean, monitoravel. Three rows carry it, out of 26 active rows when I re-read the catalog on 2026-10-02 (the catalog has grown since the first reading).

The other is an internal uptime checker that calls each of its routes every hour and writes a row per attempt into a table called checagem_rota. Up to the first reading, that table held 4,655 rows across 12 distinct routes, from 2026-09-15 16:16 to 2026-10-01 13:03 UTC. It is healthy. Its most recent row then was from 13:03 UTC, inside the hour I ran those queries.

Here is the part worth the post:

select
 (select count(*) from endpoint where ativo and monitoravel) as marcadas,
 (select count(distinct endpoint_id) from checagem_rota) as checadas,
 (select count(distinct endpoint_id) from checagem_rota
   where endpoint_id in (select id from endpoint where monitoravel)) as nas_duas;

Marked monitoravel: 3. Routes the checker had ever checked: 12. In both: zero, and still zero on 2026-10-02. The sets are disjoint. The column whose name means monitorable marks exactly the rows nothing has ever monitored, and the twelve routes being monitored every hour are all marked monitoravel = false.

Neither flag is wrong. monitoravel means "can be subscribed to by the scheduled job", and the uptime checker is a different system that never consults it. Both statements are defensible. Together they produce a database where "which routes are monitored" has two answers that do not intersect, and I had been reading one name as if it answered both.

There is a third piece of irony I will report without dressing it up. A migration on 2026-09-15 added an alert for precisely this failure mode, with a comment calling it "the watchman that dies in silence" and noting that a checker which stops writing produces the same state as everything is fine. That alert watches max(em) of checagem_rota. It cannot see the scheduled job at all. It has fired once, in the whole history of the table, for the checker it does watch.

Run this on your own schema

The test is mechanical and needs to know nothing about my tables. For each boolean your code branches on, ask who reads it:

select p.proname
from pg_proc p join pg_namespace n on n.oid = p.pronamespace
where n.nspname = 'public'
  and pg_get_functiondef(p.oid) ilike '%sua_coluna%';

Then read each hit and sort it into two piles: the ones asking a question about presentation, and the ones asking a question about execution. A flag that appears in both piles is a flag whose next edit is going to surprise someone. Mine appears in a column comment about a price page and in the first statement of a job on a five-minute cron.

And if a scheduled function returns errors as values, its success in the scheduler means the process did not crash. It does not mean it worked. Put the check on something the work itself writes. The same lesson, from the uptime side, is in my monitor that was green 284 times.

What I am not claiming

What is actually done

Done: the measurement, and the reads behind it. Decided, not built: the separate precondition column, and raising instead of returning. Does not exist: the alert that would have caught this, the history on that column that would let me date it, and any customer this could have reached.

Trend data with a memory

Every social API answers what’s trending now, then throws it away. HonestHook keeps the hourly archive, so you can ask what gained traction.

or get a key with just your email

Free, no card. One key per address — we store your email to attach the key to it and to warn you if something breaks.

Free, 1,000 credits a month, no card. Read the docs →