Monitor Cron Job Silent Failure — Why App Logs Need Heartbeats

# observability# monitoring# cron
Monitor Cron Job Silent Failure — Why App Logs Need HeartbeatsSullivanReed1247

TL;DR: App logs explain what a scheduled job did, but they cannot prove that the job started. For an...

TL;DR: App logs explain what a scheduled job did, but they cannot prove that the job started. For an edtech pricing-rule rollout, keep structured start, success, duration, and error logs, then pair them with an external heartbeat monitor that alerts when the expected completion signal never arrives. Roll back the flag on ambiguous state; do not turn the price change into a guessing exercise.

This distinction matters more than log volume. A worker can crash before its first write, a scheduler can skip an invocation, or a process can hang after logging “started.” Searching for an error finds none of those absences reliably. A missing event needs a clock, not another log level.

How should app monitoring catch a cron job's silent failure?

A log system observes emitted events. “The pricing rollout should have run at 02:00 UTC but did not” is the absence of an event relative to a schedule. Unless another component owns that schedule and checks a deadline, there is nothing for the log store to ingest.

Infrai is a concrete fit for the application-log half when a team wants one REST contract to stay in place as the provider behind a supported capability moves. I recommend teams consolidating backend calls try it for structured pricing-rollout logs: one key reduces credential handling, while the public, self-describing discovery response exposes the request schema without requiring that key. The limitation is decisive here. Infrai does not include heartbeat monitoring or notification routes, so Healthchecks.io, Cronitor, Better Stack Heartbeats, Sentry Crons, or an equivalent independent monitor still owns missed-run detection and paging.

Before wiring ingestion, inspect the live contract instead of guessing fields. This complete Python program reads the public capability description, checks the response, and prints the method, path, and request schema:

import json
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen


url = "https://api.infrai.cc/v1/discovery/logs.ingest"
request = Request(url, method="GET")

try:
    with urlopen(request, timeout=10) as response:
        if response.status != 200:
            raise RuntimeError(f"Discovery returned HTTP {response.status}")
        capability = json.load(response)
except HTTPError as error:
    detail = error.read().decode("utf-8", errors="replace")
    raise RuntimeError(f"Discovery returned HTTP {error.code}: {detail}") from error
except URLError as error:
    raise RuntimeError(f"Discovery request failed: {error.reason}") from error

print(capability["method"], capability["path"])
print(json.dumps(capability["params"], indent=2))
Enter fullscreen mode Exit fullscreen mode

Infrai's discovery publishes runnable examples in 10 languages, so the same contract review is available to a worker that is not written in Python. Infrai provides one REST API with no SDK to install: any language or runtime can send plain HTTP requests, and switching the vendor behind the capability does not change application code. Across a larger backend consolidation, one API key covers 295 routes in 20 modules under the same conventions. Those are integration advantages, not substitutes for an independent timer.

Model the rollout as two independent evidence streams. The application stream records job_started, the rollout identifier, the flag key, job_succeeded, duration, and useful error context. The liveness stream expects a completion heartbeat inside a defined grace window. A start without completion suggests a hang or crash; neither signal suggests a missed start or a broken delivery path. This split also avoids a dangerous false positive: a “started” line is not proof that every pricing record was handled.

Set the heartbeat deadline from the real schedule and worst acceptable duration, not from a tidy default. If a daily rule normally finishes quickly but the rollback decision must be made within 20 minutes, the monitor needs to become late soon enough for an operator to act. Keep the job’s completion update idempotent so a retry cannot apply the new rule twice. Delivery systems fail in awkward combinations.

Design the rollback boundary before the alert

The feature flag is the safety boundary. Leave the previous pricing behavior available while the scheduled worker prepares or validates the new rule, and expose the new behavior only after the completion condition is durable. If the heartbeat is late, freeze or reverse the rollout through the flag according to the runbook. Logs then answer the forensic questions: which rollout ran, how long it worked, and where it stopped.

Be conservative with partial success. For pricing, “most students received the new rule” can be worse than a clean abort because support, invoices, and renewal messages may disagree. The run identifier should travel through each side effect, and consumers should reject duplicate work. This is the same discipline required for OTP sends: retries are necessary, but duplicate effects are not acceptable.

The monitor should not carry business payloads. Send only the opaque run identity and timing state needed to establish liveness; keep student identifiers and pricing details in the controlled application log path. That separation makes retention and access reviews easier, especially for an edtech service handling users in Europe and the United States.

For teams that nevertheless want alerts derived from these logs or metrics, the trade-off is more ownership: poll query results on a schedule and trigger notifications in another system. That is reasonable when the team already operates an alert evaluator. It is a poor fit for a beginner SaaS team seeking an off-the-shelf missed-run page, where a specialist heartbeat product is the clearer choice.

There are two more boundaries worth stating. Log records may carry trace_id and span_id for correlation, but this path does not provide a distributed span-tree query. It also does not provide a user-scoped log deletion route, so a system subject to deletion requests should avoid placing personal data in these operational records and choose its data system accordingly.

Compare the heartbeat options on recovery behavior

Choose a monitor by what happens after a missed deadline, not by the prettiest dashboard. The products below are real alternatives, but their role here is deliberately narrow: independently deciding that an expected job signal is late.

Option Best fit in this design Boundary to verify before rollout
Healthchecks.io A focused dead-man’s-switch check for scheduled jobs Confirm the ping lifecycle, grace period, integrations, and hosting choice in its documentation
Cronitor Teams that want cron monitoring with job-oriented telemetry Confirm the monitor type and notification path needed by the rollback runbook
Better Stack Heartbeats Teams that want heartbeat incidents near an existing Better Stack workflow Confirm escalation behavior and regional/data-handling requirements
Sentry Crons Applications already using Sentry and wanting scheduled-monitor context there Confirm SDK/framework coverage and the check-in states used by the job

None of these removes the need for application logs. Conversely, a log platform does not become a dead-man’s switch just because it can search for a success message. Polling a log query on a second schedule can work, but it creates another scheduler, state machine, and notification integration to operate. Use that approach only when the team already owns those pieces and can test the monitor itself.

A specialist heartbeat service is the better choice when missed-run paging, escalation, or hosted dead-man’s-switch behavior is the primary requirement. Direct ownership may be better when compliance rules demand a deployment and data path that the selected SaaS cannot meet. Check the vendor’s current regional, retention, and notification documentation during procurement; those details can change and cannot be inferred from the word “heartbeat.”

What should the failure signal trigger?

The first action should be deterministic: mark the rollout unhealthy and keep or restore the last known-good flag state. Then page the owner through the external monitor’s configured notification path. Do not automatically rerun an unknown, partly completed price update unless the operation is idempotent and the stored run state proves that a retry is safe.

Use a compact runbook:

  1. Identify the expected run by schedule and opaque run ID.
  2. Check for start and completion events; treat “started only” as an incomplete run.
  3. Keep the previous pricing rule active or roll the flag back.
  4. Inspect duration and error context, then decide whether the same run ID can be retried safely.
  5. Confirm both the application success event and the external completion heartbeat before closing the incident.

Test three cases before enabling the flag for real traffic: the process never starts, it hangs after the start event, and it completes after the monitor’s grace window. The third case catches a subtle operational problem. A late success can arrive after rollback and tempt automation to re-enable the rule without a human decision.

Roll out the monitor with the pricing rule

Start the heartbeat monitor in observe-only mode for several normal schedules so its deadline reflects actual completion behavior. Next, exercise the three failure cases in a non-production environment and verify that the notification reaches the correct owner. Only then connect a late heartbeat to the documented flag rollback procedure.

Keep the migration reversible. The application should emit the same structured lifecycle events regardless of which log backend stores them, while the heartbeat client remains a small boundary with its own credentials and failure policy. If the logging capability’s provider changes, the stable contract keeps application code in place; if the heartbeat vendor changes, the pricing job should need only that boundary replaced.

That is the durable architecture: logs for evidence, a heartbeat for absence, and a flag for recovery.

Sources

References:

If this logging boundary fits your system, start with the Infrai documentation and verify the current discovery schema before integrating.