Node.js Health Endpoint, Cron Job, and Uptime Monitoring: 3 Heartbeat Reconstruction Checks

# node# observability# sre
Node.js Health Endpoint, Cron Job, and Uptime Monitoring: 3 Heartbeat Reconstruction ChecksHaelion14

TL;DR: For a marketplace AI agent, monitor three evidence joins: the public request to its agent run,...

TL;DR: For a marketplace AI agent, monitor three evidence joins: the public request to its agent run, each scheduled occurrence to its terminal result, and every alert to the release and feature state that produced it. A Node.js health endpoint, a cron heartbeat, polling, and metrics are useful inputs, but none is an uptime verdict by itself. Optimize first for reconstructing an incident without the original dashboard; tune polling frequency and storage ownership only after that test works.

This framing makes latency and cost accountable at the same boundary as success. The agent run record needs stable phase names, timestamps, outcome, and reported usage evidence; health remains a cheap statement about process state; cron completion is recorded after the intended side effect. Three joins are enough to expose most gaps without turning a monitoring schema into a copy of the application database.

How should a Node.js health endpoint check cron job uptime?

Begin at the buyer-visible outcome and walk backward. A marketplace request can receive an HTTP response while its agent work remains queued, can finish its model-related work while a later business side effect fails, or can complete successfully while a scheduled reconciler is late. A green process check does not resolve any of those distinctions. It proves only the narrow condition encoded by that check.

The first join connects an opaque request identifier to one agent-run record. That record should preserve accepted, started, and finished timestamps; a terminal outcome; the deployed release; and stable phase observations for queueing, retrieval, tool execution, and other phases the application actually implements. Record upstream usage only when the operation returns auditable usage evidence. Do not estimate cost inside a probe, and do not place prompt text, buyer data, listing content, or an unbounded run identifier in metric labels.

The second join connects a scheduled occurrence to all of its attempts. Its identity represents the intended occurrence, while attempt identities distinguish retries. "Missing," "late," "failed," and "unknown" must survive as separate states in durable evidence even if a compact status display renders all four as unhealthy. Otherwise an eventual retry can overwrite the event that explains the alert.

The third join connects an observation to execution context: release identity and any feature state that can alter the path. Feature toggles are operational state, not decoration. Martin Fowler distinguishes toggle categories with different lifetimes and operational characteristics; that distinction matters because a release identifier alone cannot explain behavior selected by a runtime toggle.

This is the acceptance test: choose one run, hide the dashboard, and reconstruct those three joins from retained records. Fail fast here. More charts cannot repair absent identity or overwritten history.

Keep probes narrow and success boundaries explicit

A health handler should be bounded, cheap, and meaningful to its caller. Liveness answers whether a process should be restarted. Readiness answers whether it should receive traffic. Making a remote dependency control liveness can turn a dependency problem into restart churn, so expensive end-to-end checks belong in a separate synthetic path or a bounded readiness assessment rather than in an unconditional restart signal.

Cron monitoring has a different contract. An arrival heartbeat establishes that some code reached the reporting point; it does not establish that the intended side effect committed. Put the completion record after the business success boundary, preserve the failure attempt, and treat evidence-delivery failure as its own state. Short code can still make that ordering visible:

package evidence

import (
    "context"
    "time"
)

type Completion struct {
    OccurrenceID string
    AttemptID    string
    Release      string
    ScheduledAt  time.Time
    FinishedAt   time.Time
    Outcome      string
}

type Recorder interface {
    Store(context.Context, Completion) error
}

func Run(
    ctx context.Context,
    recorder Recorder,
    occurrenceID, attemptID, release string,
    scheduledAt time.Time,
    work func(context.Context) error,
) error {
    if err := work(ctx); err != nil {
        return recorder.Store(ctx, Completion{
            OccurrenceID: occurrenceID,
            AttemptID:    attemptID,
            Release:      release,
            ScheduledAt:  scheduledAt.UTC(),
            FinishedAt:   time.Now().UTC(),
            Outcome:      "failure",
        })
    }

    return recorder.Store(ctx, Completion{
        OccurrenceID: occurrenceID,
        AttemptID:    attemptID,
        Release:      release,
        ScheduledAt:  scheduledAt.UTC(),
        FinishedAt:   time.Now().UTC(),
        Outcome:      "success",
    })
}
Enter fullscreen mode Exit fullscreen mode

This sample deliberately leaves a hard production decision exposed: if the business action succeeds and storing its completion fails, rerunning the whole occurrence may repeat the action. The implementation needs an idempotency policy for the business operation and a separately retryable evidence path. Returning the recorder error prevents a false confirmed heartbeat, but it cannot decide whether the caller should replay evidence or replay work.

That distinction is small in code and large during an incident.

Make latency and cost explainable, not merely visible

An end-to-end duration says that a run was slow. It does not say where time accumulated. Use a stable, bounded phase vocabulary and retain per-run detail in logs, traces, or event records; aggregate metrics by dimensions such as phase, outcome, route class, or worker pool. The separation protects the metric series budget while leaving high-cardinality evidence where incident queries can use it.

The same rule applies to cost. Attribute usage to the run and phase only when a durable accounting record or the upstream response supplies it. A health poll should never trigger an agent turn merely to prove the agent can run: that mixes synthetic activity with buyer work, adds variable latency and usage, and makes the resulting signal harder to interpret. A synthetic transaction can be appropriate when the SLO explicitly covers the full path, but it needs its own identity and outcome so responders can separate it from marketplace traffic.

Polling closes the outside-in gap. An external observer can detect DNS, routing, TLS, or process reachability symptoms that an in-process metric cannot see, while internal evidence explains the work after acceptance. Set the interval from the detection objective and error-budget policy. For example, checking 100 targets once every 60 seconds produces 144,000 checks per day before retries; changing to 30 seconds doubles that load. It still does nothing to improve missing causal records.

Capacity planning follows from these units:

  • Poll traffic follows target count and polling frequency.
  • Completion-event volume follows scheduled occurrences and retry attempts.
  • Detailed run storage follows retained runs multiplied by recorded phases.
  • Metric series follow the product of bounded label cardinalities.

A run identifier in a metric label breaks the last assumption. Avoid it.

Choose ownership after an export drill

The useful buy-versus-build question is not which interface looks better. It is whether the team can preserve and retrieve the three joins within its incident-review window while meeting the on-call and capacity obligations created by the chosen transport.

Decision pressure Managed transport Self-hosted transport
Reconstruction Verify correlated raw records, timestamps, and export behavior Design indexes and queries around the three joins
On-call load Provider operates the service; the team still owns signal semantics Team owns upgrades, backup, recovery, and monitoring
Lock-in Test schema export and deletion behavior before committing Control the schema while accepting component migration work
Capacity Treat ingestion, retention, and cardinality limits as design inputs Forecast storage growth, query concurrency, and series count

Run an export drill before choosing either column: export one completed marketplace run, its related scheduled occurrence, and its execution context, then reconstruct the incident narrative without the normal UI. The ownership model passes only if the evidence survives that exercise. Price can be an input, but it cannot compensate for evidence that disappears at the first serious review.

A public status page belongs downstream of this process. It communicates assessed service state; polling it as the primary detector makes presentation the source of truth and discards the internal timestamps needed to explain the change.

Where does the three-join model stop helping?

The limitations and trade-offs matter. For a single synchronous process with no scheduled work, no agent phases, and no per-run usage requirement, this model isn't suitable; a bounded health response, an external path check, and request metrics are the simpler alternative, with less storage and less cognitive load.

It also does not measure the buyer's browser experience. Core Web Vitals cover loading, responsiveness, and visual stability through LCP, INP, and CLS, and the guidance evaluates the 75th percentile. If the marketplace SLO includes browser interaction, field measurements need route and release context; server polling cannot substitute for them.

Privacy sets another boundary. Opaque correlation identifiers and minimal operational fields are sufficient to prove ordering and outcome. Content is not required.

The final decision rule is blunt: start from an alert, recover one affected agent run and one relevant scheduled occurrence, establish the release and feature state, and identify which latency or cost evidence supports the SLO decision. If those joins hold, the monitoring design can explain the incident. If they do not, a faster poll only reports confusion sooner.

Sources