Skip to content
← Back to blog

Published on September 9, 2026 · 4 min read

My monitoring stays silent. I built it so that silence would mean something

My backup-checking script usually says nothing. It runs, checks a few dozen things, and finishes without a word. Silence means: everything is fine.

That is convenient, and at the same time it is the riskiest design decision in the whole setup. Because silence has exactly two causes: either there really is nothing to report, or the check never ran. From the outside they look identical.

A check that checks itself

That is why the very last thing the script checks is itself. Two questions: are the system's scheduled tasks even loaded, and when did this check last run. If it hasn't run in roughly a day, it says so directly — because silence during that time meant nothing.

The threshold is 48 hours. Below that it's usually just a computer switched off for the weekend, and an alarm about a switched-off computer teaches you to ignore alarms.

That's the whole rule, and it fits in one sentence: silence is proof only when something confirms that the silent system is alive. Without that, no alarm is indistinguishable from no monitoring.

What this check actually verifies

Backups: the age of the system backup, how much of the allotted limit is used, the error code reported by the backup mechanism, the depth of history — because a shortened chain means the backup restarted and older file versions stopped existing. On top of that, the age of the cold backup and of the off-site copy.

Hardware: SMART status of the drives, free space on every backup target, a stuck network-share mount. On the server side: the status of each storage pool, mismatched blocks, free space in the pools, the presence of the correct backup file, and the size of the network recycle bin that receives files deleted during history pruning.

The thresholds aren't round for looks. 36 hours for the system backup is the hourly schedule plus a margin for a night without network. 10 days for the cold backup, made weekly, is already a slip. 14 days for the off-site copy is a week plus a margin for travel. 50 GB free on the target, because sync-with-deletion needs room to write before it deletes old files.

What I deliberately do not monitor

This is the more important half of the list, because every notification without a decision behind it teaches you to ignore the rest.

I don't watch server load, memory usage, whether temperatures are within range, or uptime without a restart. None of those numbers changes any decision of mine. I also don't watch the CPU, GPU, power supplies or RAM — not because they can't fail, but because they give no advance signal. A drive announces failure a month ahead, which is why I watch it. Everything else either works or it doesn't.

None of my machines has error-correcting memory, so a memory error is invisible until it causes a failure. Instead of pretending to measure something I don't have, I count system crashes in a 14-day window and spontaneous server restarts. One crash is information, several are an alarm. That's the only signal I have about memory, power and cooling — and I state it as exactly that, instead of showing a green indicator based on nothing.

What I did not do

I did not add a reminder about one of the storage pools that has no redundancy. It was tempting, because the script sees this state, and the operating system reports that pool as healthy — it's defined as a mirror made of a single drive. It looks healthy and technically it is healthy.

I checked what lives there: recoverable data. A failure of this drive is a cost in time, not a loss. So instead of a monthly warning I would ignore anyway, the decision is written down in the documentation along with the condition that would overturn it: if anything unrecoverable ever ends up on this pool, revisit the topic. Documentation remembers better than a notification I've trained myself to scroll past.

Three things I take from this

Silence has to be earned. A system that stays quiet has an obligation to prove it's alive — otherwise its silence is worthless.

A list of unmonitored things is part of the monitoring. Written down and justified, it stops being an oversight and becomes a decision you can revisit.

An alarm without a decision behind it is a cost, not a gain. Before adding a check, answer what you will actually do when it fires. If the answer is "nothing," you have just weakened every other alarm you have.


The thresholds, numbers and decisions described above come from a configuration measured and organized on 16 August 2026 on my own home hardware.

Patryk Piecyk

Patryk Piecyk

Warsaw · junior implementation consultant · available now

For seven and a half years I worked at a German company, five of them running its office: orders, invoices, complaints, ERP. Since June 2026 I have been building my own tools for that same work — I am not a programmer by training; the code is written together with an AI assistant, while the design, the decisions and the testing are mine. These notes describe what broke in those systems and what came out of it.

Got a question about this piece?

Write to me →