BugsRadar

Background jobs and queues that fail quietly

Updated 1 October 2026 · BugsRadar team

Nothing fails more quietly than a background job. A worker throws, the queue retries, the retry throws, the message lands in a dead-letter queue, and the metric that would show it is on a page nobody opens. The user who ordered the thing the job was processing finds out first.

Worker services and consumers

Job runners and schedulers

If the job runner logs a failed job at ERROR through the logger you connected - ILogger, winston, pino, Python's logging - the failure arrives without any code of yours. If it swallows exceptions or logs them elsewhere, report in the job itself: one SendException, sendException or send_exception in the catch, with the job's name as the module. The message then says which job, on which host, with the stack.

Retries become a count

A job that fails, is retried and fails again is the same error repeating. The first failure arrives in full; the retries are counted in that message, so "×24 since 03:00" tells you the queue has been retrying for an hour without a flood of notifications. Group by what matters: log with a message template, Job {JobName} failed, and every failed order is one error with a count, not a message per order. See What you receive.

Crash loops

A worker that dies and is restarted by its supervisor every five seconds is the worst case for any alerting tool. The package folds repeats before they leave the process, BugsRadar counts them in one message, and if the chat still can't keep up the waiting errors arrive bundled. One message with a rising count, not thousands. See Keeping the chat quiet.

Short-lived jobs

A job that exits right after its work can exit before the report has left. Call FlushAsync, flush() or bugsradar.flush() before the process ends; in Python the exit hook does it for you. A job that isn't an application at all - a script - reports with one HTTP request: see the guides.

Where to go next

Cron, backups and services on one VPS is the same problem outside the application. Internal tools nobody monitors covers the tools the jobs serve.

Next: When their outage becomes your bug →