‹ Ship Your Tool Lesson 10 of 16
Contents Lesson 10 of 16

3 min read · professional

Knowing it broke before someone tells you

Your tool will break while you are not looking. The question is only whether you find out from a monitor or from opening it three weeks later and realising it has been dead the whole time.

Check the thing, not the process

A process can be running and serving errors. Check behaviour instead: hit /api/health from outside your infrastructure, on a schedule, and expect a specific response body rather than merely a 200.

From the previous unit, that endpoint reports the store and upstream separately and costs no quota — which is what makes it safe to call every few minutes.

Watch four things

Is it up? The health check, from outside.

Is it fresh? A cockpit serving hour-old prices during the session is not "down" and is broken. Include the age of the newest data in the health response and alert on it — this is the failure that monitoring most often misses, because everything is technically working. Give the alert the market calendar and the feed's delay first, though: outside the session an old timestamp is the correct answer, and on a fifteen-minute feed fifteen minutes is the floor even at ten in the morning. An alert that fires every Friday evening is one you will learn to ignore.

Is it within quota? calls_used against the limit. This is the wall arriving, and you want to see it approaching rather than at the moment it lands.

Is it erroring? A rate of 5xx, not individual ones.

The watchdog has to be somewhere else

A monitor running on the machine it watches tells you nothing when that machine is gone. It has to run somewhere independent — a scheduled job elsewhere, a free uptime service, anything that is not the thing being watched.

And the alert has to reach you somewhere you actually look. An email you check weekly is not monitoring; it is a diary.

The failure this lesson exists to prevent

The specific shape, and it is common: something breaks quietly, nothing tells you, and you notice weeks later.

The reason it is so common in personal projects is that a silent failure has no cost until the moment you need the tool — and that moment is exactly when you cannot afford to spend an evening on repairs. A watchdog converts a three-week outage into a ten-minute one, and it is a few lines plus a free service.

Watch the watchdog

A monitor that stops running looks exactly like everything being fine.

So make it report success, not only failure — a heartbeat you would notice the absence of. Most uptime services do this by design, which is a good reason to use one rather than write your own.

Try it now

Set up the external check, then stop your deployment and confirm you are told within minutes. Then stop the monitor and see whether you notice — that second experiment is the one that teaches the lesson.