Published on August 13, 2026 · 4 min read
My system stopped at 1 a.m. I found out in the morning, from a late report
On 1 July 2026 my system stopped working during the night and said nothing about it. I found out in the morning — not from an alert, but because the daily report arrived late and was shorter than it should have been.
I am writing this up because it is the most instructive thing that has happened to me in this project, and because the mechanism behind it repeats everywhere something waits in a queue.
What actually happened
I run a set of nightly processes that all pass through a single queue to a local language model: market analysis, compliance checks, security monitoring. They all enter through the same gate, one after another.
After a series of unusually long requests, the process in the middle hung internally. The word "internally" is the important one: it was still running. It listened on its port, it answered that it was alive, it left no error entry behind. It simply stopped handling new requests.
This is the worst kind of failure I know. A process that crashes leaves a trace: a log line, a restart, an alert. A process that stands still and answers "everything is fine" leaves nothing — and the only symptom is that something further down the chain did not arrive on time.
The cause
Two things at once, and only together do they make a failure:
- A dead, orphaned network connection. The other side had disappeared, but the socket stayed open. From the process's point of view the connection was still "in progress", so it waited for an answer that would never come.
- No timeout of any kind on the critical path. This is the real cause. The first point is an ordinary thing that happens on networks every day — only the missing time limit turns it into a full stop of the entire system.
The queue let one job through at a time. It was enough for the first one to get stuck, and everything behind it stopped.
What I did not do
I did not restart the process, and I did not add a scheduled job to restart it every hour. That was the first idea and it is a bad one: a restart removes the symptom and makes sure the cause stays in the code forever, except that nobody looks for it any more, because "it works, doesn't it".
The repair
Three changes, all of them in the code:
Releasing the queue on EVERY path. Previously the queue was released after a job had been processed successfully. Now it is released always: on success, on error, and on timeout. This is the one change that removes the whole class of problem rather than this particular case. If I had to remember one sentence from this write-up, it would be: cleanup has to hang on the exit from a function, not on its happy ending.
Hard time limits on every step. Every call now has an upper bound. A job that exceeds it is aborted and recorded — instead of waiting forever.
An e-mail alert on every self-repair. Not on failure — on repair. This is an important distinction: a system that quietly heals itself is a system that quietly hides from you that something is wrong. I want to know how many times it picked itself up, because a rising number is a signal that the cause has come back somewhere else.
What happened on the night of 2–3 July
The repair got a practical test one day later. The queue backed up under an unusually heavy load — three consecutive jobs waited longer and longer (times in UTC):
2026-07-02 23:28 → waited 51 min → queue released automatically
2026-07-02 23:48 → waited 65 min → queue released automatically
2026-07-03 00:08 → waited 78 min → queue released automatically
Nobody got up. These are night processes, so an hour of delay means nothing here — what means something is that in the morning everything was done, and there were three messages in my inbox saying exactly what had happened.
Over seventeen days of observation I counted 33 such events. Without that change, every one of them would have been a silent halt of the whole chain, up until the moment somebody noticed a missing report.
What this means beyond my own server
Three things I take with me into every system that follows:
"Alive" and "working" are two different questions. Checking whether a process answers says nothing about whether it is doing anything. Since I learned that, my checks measure progress, not presence.
Silence is not good news. If the only signal that a system is working is the absence of an alarm, you do not have monitoring — you have hope. The same mechanism later caught a backup that had not changed for 169 days: not because something shouted, but because the daily report states the age of the backup as a number.
Fix the class, not the case. I could have added a timeout in the one place where it got stuck. Instead I went through every exit path of that function. It took half a day longer and it has not come back since.
The numbers in this write-up come from monitoring data for 1–17 July 2026 and are the same ones I show on the home page.

Patryk Piecyk
Warsaw · employed since 10.2026 · growing towards ERP/SAP and AI
For seven and a half years I worked at a German company, five of them running its office: orders, invoices, complaints, ERP. Since June 2026 I have been building my own tools for that same work — I am not a programmer by training; the code is written together with an AI assistant, while the design, the decisions and the testing are mine. These notes describe what broke in those systems and what came out of it.
Got a question about this piece?
Write to me →