How to monitor cron jobs with heartbeats (dead man's switch)
A backup job that stops running never tells you. Heartbeats, grace periods, systemd, crontab and GitHub Actions: a real setup and the traps in it.
My VPS runs a handful of scheduled jobs: a backup every night, a security scan every Monday, a restore test twice a month. The kind of jobs you set up once and forget. That is exactly the problem: the day one of them stops, nothing tells you.
It happened twice. The weekly security scan never ran on its own for several days: it was declared in /etc/cron.d/, but cron wasn’t installed on the server. Nothing was reading the file. On another machine, an old backup check was calling a monitoring URL that had been deleted, and failing silently. In both cases, no error, no email: the job didn’t fail, it just wasn’t there anymore. And an absence raises no alert, unless something is waiting for it.
That’s what a heartbeat, or dead man’s switch, does: the job sends a ping every time it succeeds, and the alert fires when the ping doesn’t arrive. By the end of this post you’ll know how to tune a heartbeat, wire it into a script, a systemd service, a crontab or a GitHub Actions workflow, and avoid two traps I ran into: “monthly” that lasts 31 days, and a heartbeat that gets its pings from the wrong source.
The short version
- The problem: a scheduled job that stops running produces no error. You find out there’s no backup the day you need one.
- The fix: on every success, the job calls a URL. A monitoring service waits for that call and alerts if it doesn’t arrive on time.
- The tuning: interval = the job’s period, grace = normal run time + slack. Work out the maximum gap between runs, not the average.
Three things to know first
- Scheduled job. A program started automatically at a set time. Linux has two mechanisms:
cron(a line in a crontab) and systemd timers (a.timerfile that starts a service). My VPS uses timers. - Heartbeat, or ping. A plain HTTP call (
curl https://…) the job makes at the end to say “I ran, and I succeeded”. The service receiving it records the time. - Grace period. The slack a job gets before the alert. A daily backup can start a few minutes late or run a bit longer: the grace period keeps that from ringing the bell.
Monitor the outcome, not the process
A picture: a solo hiker tells a friend “if I haven’t called by 8 pm, raise the alarm”. The friend doesn’t try to work out what went wrong (sprained ankle, dead battery, storm): they react to the call that doesn’t come. That’s a heartbeat.
Classic monitoring checks whether a process crashes. It can’t see:
- a timer or crontab that nothing reads anymore (my case);
- a server that was off at the scheduled time;
- a script that exits 0 without doing anything because a variable was empty;
- a job stuck for hours on a network mount.
A heartbeat flips the question. Not “was there an error?” but “did I receive proof that it worked, on time?”. Anything that stops the proof from arriving, whatever the cause, ends up as an alert. Including causes you never thought of.
The flip side: you’re no longer told when things work, only when they stop. No “backup OK” email every morning that nobody reads after the first week.
The tool: Pinguro, and the alternatives
I use Pinguro, the monitoring tool I build (pinguro.app). So I’m describing something I know from the inside, but the principle is the same everywhere. Solid alternatives:
- Healthchecks.io: dedicated to heartbeats, open source and self-hostable (official Docker image). Supports a start signal (
/start), a failure signal (/fail) and the exit code in the URL. - Cronitor: tracks each run with
?state=run,completeorfail. - Better Stack: heartbeats built into its uptime monitoring, with
/failto report a failure.
In Pinguro, a Heartbeat monitor has three settings: a name, an interval (1 minute to 43,200 minutes, i.e. 30 days) and a grace period (0 to 1,440 minutes, i.e. 24 hours). It gives you a URL like:
https://pinguro.app/h/<token>
The token is 32 hex characters. The URL accepts GET and POST, answers OK in plain text or 404 Token invalide otherwise, and is rate limited to 10 calls per minute per token. The token can be regenerated if it leaks.
What Pinguro does not do today: no start ping, no explicit failure signal, no exit code. There’s one signal: “I succeeded”. If you want run durations or an immediate alert on failure, Healthchecks.io and Cronitor have that.
One more useful detail: a heartbeat that has never been pinged stays pending for 24 hours after creation, so you have time to wire it up. After that, it goes down.
Picking the interval and grace period
The rule I use:
- interval = the job’s period;
- grace = normal run time + slack for anything that can shift a run: timer jitter (
RandomizedDelaySec), an automatic nightly reboot, a catch-up run after downtime.
The alert fires when the time since the last ping exceeds interval + grace. My real heartbeats:
| Job | Schedule | Interval | Grace |
|---|---|---|---|
| Nightly backup | every night 02:30 UTC | 24 h | 2 h |
| Security scan | Monday 07:00 UTC | 7 d (10,080 min) | 6 h |
| Restore test | 2nd and 16th of the month, 05:00 UTC | 30 d | 1 d |
| Publishing this site | every day 05:05 UTC | 24 h | 8 h |
The jitter shows up directly in systemctl list-timers, which lists timers with their last and next run (excerpt; my unit names are in French, scan-securite is the security scan):
$ systemctl list-timers
NEXT LEFT LAST PASSED UNIT
Wed 2026-09-30 02:34:10 UTC 18h Tue 2026-09-29 02:30:17 UTC 5h 14min ago monprojet-sauvegarde.timer
Mon 2026-10-05 07:04:38 UTC 5 days Mon 2026-09-28 07:00:23 UTC 24h ago scan-securite.timer
…
The backup is set for 02:30, but its next run is at 02:34:10: that’s RandomizedDelaySec=5min. With TimeoutStartSec=30min on top, 2 hours of grace covers the worst case with room to spare. The scan runs on Monday morning; 6 hours covers a late run without letting a real outage linger.
Too short a grace period means false alerts, and an alert that cries wolf ends up ignored. Too long, and detection is delayed by that much. For a daily backup, finding out after 26 hours instead of 24 doesn’t matter; finding out after a week does.
The “once a month” trap
The automated restore test (see automated backup restore testing) was meant to run once a month. I created its heartbeat with Pinguro’s maximums: 30-day interval, 1-day grace. That’s 31 days before an alert.
Except there are 31 days between January 2 and February 2. The ping is sent when the test finishes, so it lands 31 days plus a few minutes after the previous one if the test runs slightly longer. Every 31-day month was a potential false alert, down to the minute.
Two fixes: a longer interval (Pinguro caps at 30 days), or a more frequent job. I picked the second: the test now runs on the 2nd and 16th of every month.
# /etc/systemd/system/myproject-restore-test.timer
[Unit]
Description=Backup restore test (twice a month)
[Timer]
OnCalendar=*-*-02,16 05:00:00 UTC
Persistent=true
[Install]
WantedBy=timers.target
The longest gap between two runs is now 17 days (from the 16th of a 31-day month to the 2nd of the next), far from the limit. The monitor stayed at 30 d + 1 d; tightening it to around 18 days would catch a missed run sooner.
The general lesson: “monthly” isn’t a fixed period. If your tool thinks in durations, check the maximum gap between two runs, not the average. Same for a “weekdays only” job: Friday to Monday is 3 days.
Ping only on success
This is the rule that makes heartbeats worth having. A job that fails doesn’t ping; the alert fires when the grace period runs out (on top of the job’s own alert, if it has one).
In a bash script, the ping goes at the very end, after set -euo pipefail has made sure any earlier error aborts the script:
#!/usr/bin/env bash
set -euo pipefail
# ... the actual job: dump, encrypt, upload ...
# Heartbeat, only reached if everything above succeeded.
HEARTBEAT_URL="${HEARTBEAT_URL:-}"
if [ -n "$HEARTBEAT_URL" ]; then
curl -fsS --max-time 10 --retry 3 -o /dev/null "$HEARTBEAT_URL" \
|| echo "heartbeat unreachable (job succeeded anyway)" >&2
fi
The curl flags matter:
-f: an HTTP error (say, a 404 from a deleted token) makescurlfail instead of being ignored, and leaves a trace in the log. A ping to a deleted URL no longer goes unnoticed.-sS: quiet, except for errors.--max-time 10: the ping never blocks the job for more than 10 seconds per attempt.--retry 3: retries on transient errors (timeouts, 5xx).|| echo …: a failed ping is logged but doesn’t fail a job that succeeded.
Here’s what -f does with a token that doesn’t exist, and the same request without -f:
$ curl -fsS -m 10 -o /dev/null https://pinguro.app/h/00000000000000000000000000000000
curl: (22) The requested URL returned error: 404
$ echo $?
22
$ curl -s https://pinguro.app/h/00000000000000000000000000000000
Token invalide
Without -f, curl prints the response and exits 0: as far as the script knows, all is well. That’s exactly how my old backup check kept pinging a dead URL without anyone noticing.
The if [ -n … ] makes the heartbeat optional: the same script can run in testing without touching the monitor.
Wiring a systemd service without touching the script
On my VPS, scheduled jobs are systemd timers (I wrote a separate post on why I moved from cron to systemd timers). To add a heartbeat to an existing service, a drop-in (a small file that adds to a service’s configuration without editing it) is enough; the script stays as is. My security scan is called scan-securite:
# /etc/systemd/system/scan-securite.service.d/heartbeat.conf
[Service]
# Environment file holding HEARTBEAT_URL (mode 600, owned by root).
EnvironmentFile=-/etc/scan-securite.env
ExecStartPost=/bin/sh -c '[ -z "$HEARTBEAT_URL" ] || curl -fsS -m 10 --retry 2 "$HEARTBEAT_URL" >/dev/null || true'
Why this works: for a Type=oneshot service, ExecStartPost= only runs if the last ExecStart= command exited successfully (systemd.service). A failure, a crash or hitting TimeoutStartSec: no ping. The trailing || true keeps an unreachable monitoring service from marking a successful job as failed.
Install:
sudo mkdir -p /etc/systemd/system/scan-securite.service.d
sudo cp heartbeat.conf /etc/systemd/system/scan-securite.service.d/
echo 'HEARTBEAT_URL=https://pinguro.app/h/<token>' | sudo tee -a /etc/scan-securite.env >/dev/null
sudo chmod 600 /etc/scan-securite.env
sudo systemctl daemon-reload
systemctl cat scan-securite.service # the drop-in should show up
systemctl cat prints the service and then its drop-ins, each preceded by its path. On my server (excerpt):
# /etc/systemd/system/scan-securite.service
[Service]
Type=oneshot
ExecStart=/srv/outils/scan-securite.sh
SuccessExitStatus=1
TimeoutStartSec=30min
…
# /etc/systemd/system/scan-securite.service.d/heartbeat.conf
[Service]
EnvironmentFile=-/etc/scan-securite.env
ExecStartPost=/bin/sh -c '[ -z "$HEARTBEAT_URL" ] || curl -fsS -m 10 --retry 2 "$HEARTBEAT_URL" >/dev/null || true'
If the second part is missing, the drop-in isn’t being read (wrong directory name, forgotten daemon-reload) and no ping will ever go out.
The ping URL isn’t a critical secret: all it can do is say “I ran”. But someone who knows it can make a dead job look alive. So it stays out of git, in the environment file.
One subtlety with the security scan: its service sets SuccessExitStatus=1, because the script exits 1 when it has found something to report. systemd counts that as success, so the ping fires in that case too. The journal for the September 28 run shows it (log messages in French: “alert sent, 1 item”):
$ journalctl -u scan-securite.service
2026-09-28T07:00:23+00:00 systemd[1]: Starting scan-securite.service - Controle de securite hebdomadaire...
2026-09-28T07:00:41+00:00 scan-securite.sh[1607249]: scan-securite: alerte envoyee (1 point(s))
2026-09-28T07:00:42+00:00 systemd[1]: scan-securite.service: Deactivated successfully.
2026-09-28T07:00:42+00:00 systemd[1]: Finished scan-securite.service - Controle de securite hebdomadaire.
The scan found something and sent its alert, but the service ended “successfully”, so the heartbeat was pinged. That’s intended: the heartbeat means “the scan ran”, not “the scan found nothing”. Those are two separate alerts, and sending the second one is the scan’s job.
If your tool accepts a failure signal, ExecStopPost= is another option: it runs in every case and gets $SERVICE_RESULT, which is success when all went well. OnSuccess= (systemd 249+) can also start a dedicated ping unit. For a plain “I succeeded”, ExecStartPost= is the shortest path.
In a crontab
On a machine that runs cron, chaining with && does the same job: the ping only fires if the job exits 0.
# crontab -e
30 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 -o /dev/null https://pinguro.app/h/<token>
Two notes. The URL is readable in the crontab: on a shared machine, prefer a wrapper script that reads an environment file. And check that cron is actually running. On my VPS, this is what it said:
$ systemctl is-active cron
inactive
$ dpkg -s cron
dpkg-query: package 'cron' is not installed and no information is available
The package wasn’t installed at all on my hosting provider’s Debian 13 image. A file dropped in /etc/cron.d/ was read by nobody.
In GitHub Actions
Scheduled workflows need a heartbeat more than anything else. GitHub’s docs say it plainly: the schedule event can be delayed under heavy load, some queued jobs may be dropped, and in a public repository scheduled workflows are disabled after 60 days without activity.
name: nightly-report
on:
schedule:
- cron: "15 3 * * *" # UTC
workflow_dispatch:
jobs:
report:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- run: ./scripts/report.sh
- name: Heartbeat
if: success()
run: curl -fsS -m 10 --retry 3 -o /dev/null "$HEARTBEAT_URL"
env:
HEARTBEAT_URL: ${{ secrets.HEARTBEAT_URL }}
if: success() is already the default for a step, but writing it out makes the intent obvious. The URL goes into repository secrets, never into the file.
This site is a live example. It’s rebuilt by a workflow scheduled at 05:15 and 09:45 UTC, to publish scheduled posts. Over three days, GitHub’s scheduler started only three runs: September 27 at 10:33, September 28 at 11:41 and 17:49. Nothing on the morning of September 29, the day a post was due.
And the heartbeat said nothing, for a reason I hadn’t seen: it was pinged at the end of every successful build, including those triggered by a git push. I was pushing several times a day, and each push reassured the monitor in the scheduler’s place. The general trap: a heartbeat must be pinged by the job it watches, and by that job only. Otherwise another source can hide its disappearance.
I fixed both. A systemd timer on the VPS now triggers the workflow every morning through GitHub’s API (workflow_dispatch, run immediately), and only that trigger pings the heartbeat:
- name: Heartbeat
if: success() && github.event_name == 'workflow_dispatch'
If you stick with schedule, allow a generous grace period (several hours) and only ping on github.event_name == 'schedule'.
Measuring run time with start and end pings
Some tools accept a start ping. With Healthchecks.io you call <url>/start at launch and <url> at the end; it derives each run’s duration and can alert when a job starts but never finishes (docs):
URL="https://hc-ping.com/<uuid>"
curl -fsS -m 10 --retry 3 -o /dev/null "$URL/start"
/usr/local/bin/backup.sh
curl -fsS -m 10 --retry 3 -o /dev/null "$URL"
Handy for spotting a backup that creeps from 5 to 40 minutes: it still succeeds, but something is growing. Pinguro doesn’t have this signal today; on my VPS, TimeoutStartSec kills a stuck job and the missing ping does the rest.
Test the wiring for real
An untested heartbeat is worthless: my old backup check was pinging a dead URL. After wiring each one, start the job by hand and read its journal:
sudo systemctl start myproject-backup.service
journalctl -u myproject-backup.service -n 20 --no-pager
A successful run of my nightly backup looks like this (excerpt, in French: “encrypted archive”, “uploaded”, “backup succeeded”):
02:30:24 sauvegarde-b2.sh: archive chiffrée : monprojet-2026-09-29T0230Z.tar.zst.gpg (9990131 octets)
02:30:26 sauvegarde-b2.sh: envoyé : sauvegardes/quotidiennes/monprojet-2026-09-29T0230Z.tar.zst.gpg
02:30:26 sauvegarde-b2.sh: sauvegarde réussie
02:30:26 systemd[1]: monprojet-sauvegarde.service: Deactivated successfully.
No heartbeat injoignable (heartbeat unreachable) line: the ping went out. What’s left is to check in the tool that the monitor went UP with the ping time. I did that for every job; for the restore test, I waited for a real run of the timer.
Keep an inventory too: one line per job, with its schedule, its monitor and its status. A job added without a heartbeat is, once again, a job that can die quietly.
Who watches the watcher?
Pinguro runs on the same VPS as the jobs it watches. If the server goes down, it can’t alert about missing pings. That case is covered by external monitoring: UptimeRobot polls https://pinguro.app/readyz and emails me if the server, the database or the worker stops responding. One watches the watcher, the other watches everything else. For the same reason, heartbeat alerts should go through a channel that doesn’t depend on the monitored machine.
Key takeaways
- A job that stops running produces no error: only a heartbeat waiting for its ping will catch it.
- Interval = the job’s period; grace = normal run time + slack (jitter, reboots, catch-up runs).
- “Monthly” can mean 31 days: work out the maximum gap between runs, not the average.
- Ping on success only: last line of a
set -escript,&&in crontab,ExecStartPost=in systemd,if: success()in CI. - A heartbeat must be pinged only by the job it watches: any other source hides its disappearance.
curl -fsS --max-time 10 --retry 3:-fto surface 404s, a timeout so the ping never blocks the job.- Test every heartbeat with a real run, and have your monitoring tool watched from outside.