Node.js Duplicate Alerts — Idempotency Keys for Polling Retry Webhook Errors

# observability# node# webhooks
Node.js Duplicate Alerts — Idempotency Keys for Polling Retry Webhook ErrorsNielsChristensen4981

A page fires during an edtech pricing-rule rollout. The on-call sees duplicate alerts after webhook...

A page fires during an edtech pricing-rule rollout. The on-call sees duplicate alerts after webhook retries: one checkout failure has acquired several identities because the Node.js polling sender generated a fresh dedupe key for each attempt. The loudest number is twelve. The useful number is one: one logical pricing evaluation has not reached its notification destination, and idempotency must preserve that identity until resolution.

TL;DR: Treat evaluation, delivery attempt, and alert notification as different events. Give the logical failure a stable dedupe key, record every attempt as a metric or trace event, and page only on the sustained number of unresolved logical failures. Roll back the pricing flag when the user-facing error-budget policy says to, not when retry traffic makes an attempt counter look frightening.

That separation is the fix.

Without it, a webhook retry policy becomes an alert multiplier, a polling loop can reopen a resolved incident, and the on-call cannot tell whether rollback will protect learners or merely hide a noisy sender.

How should idempotency keys stop duplicate alerts across polling retries?

The page should identify the affected pricing rule and flag cohort, count distinct unresolved evaluations, show the age of the oldest one, and link those failures to delivery attempts. It should not page once per HTTP request. A request is transport activity; the pricing evaluation is the unit of user impact.

Suppose evaluation eval_7f3 fails, and a sender makes four attempts before delivery succeeds. Four attempt failures are legitimate observations. They belong in a counter such as webhook_delivery_attempts_total, partitioned by outcome and bounded destination class. They do not constitute four incidents. The open-failure gauge, or an equivalent query over durable state, moves from zero to one and then back to zero after acknowledged delivery.

Keep identity out of labels.

Avoid putting evaluation IDs, learner IDs, URLs, or raw error strings in metric labels. Those values create unbounded cardinality and belong in logs or traces. Metrics should retain dimensions an operator can aggregate: pricing-rule version, flag cohort, destination class, and a small outcome vocabulary. The trace or structured log carries the dedupe key that joins a page to the exact attempts.

This is where rollback safety becomes concrete. A rollout policy might halt exposure on a fast signal and roll back only after a slower, user-impact signal breaches its budget; the actual windows and thresholds must come from the service's SLO and traffic shape. There is no defensible universal percentage. Capacity planning matters too: at peak enrollment, retries multiply outbound work even when distinct failures remain flat, so queue depth and oldest-item age should be visible before the sender exhausts its concurrency.

Work backward from notification fan-out

Start with the repeated page and walk upstream. Did the paging system receive the same logical alert several times, did the monitor evaluate the same time window repeatedly, or did the application mint a new identity on every attempt? Those failures look identical on a phone and require different repairs.

The application should persist one record per logical notification and separate it from an append-only attempt history. Pollers claim eligible records; senders append attempts; an acknowledgement closes the logical record. A retry changes scheduling metadata, not identity. If the process restarts after sending but before persisting the acknowledgement, the receiver may still observe a duplicate, because a network timeout cannot prove whether the remote side committed the request. End-to-end idempotency therefore requires the stable key to cross the webhook boundary and the receiver to honor it.

Use a key derived from immutable business identity and effect, not from time or attempt count. For this rollout, pricing-evaluation-id + notification-kind + destination-id is a reasonable tuple if each field is stable and the intended effect occurs once. Adding the retry number defeats deduplication. Omitting the notification kind can suppress a later, distinct effect.

package delivery

import (
    "crypto/sha256"
    "encoding/hex"
    "strings"
)

type Event struct {
    EvaluationID string
    Kind         string
    Destination string
}

func DedupeKey(e Event) string {
    canonical := strings.Join([]string{e.EvaluationID, e.Kind, e.Destination}, "\x00")
    sum := sha256.Sum256([]byte(canonical))
    return hex.EncodeToString(sum[:])
}
Enter fullscreen mode Exit fullscreen mode

A digest makes the key compact; it does not repair unstable inputs.

Keep a unique constraint on the canonical logical identity, persist the generated key, and send that same value on every attempt. The database transition that schedules retry should be conditional on the current state, so two pollers cannot both turn the same pending item into independent work.

Instrument the state machine, not the loop

Polling errors need two layers of telemetry. The loop itself reports health: poll duration, claim errors, claimed records, queue depth, and age of the oldest eligible record. The delivery state machine reports outcomes: logical failures opened, attempts made, acknowledgements recorded, items resolved, and items exhausted under the configured policy. Combining those layers into a single errors_total counter makes diagnosis nearly impossible.

A minimal sender can make the distinction explicit. The storage methods below represent atomic state transitions; their implementation must use the datastore's compare-and-set or transactional primitive.

package delivery

import (
    "context"
    "errors"
    "time"
)

type Store interface {
    Claim(context.Context, time.Time) (Job, error)
    Acknowledge(context.Context, string) error
    ScheduleRetry(context.Context, string, time.Time) error
}

type Sender interface {
    Post(context.Context, Job) error
}

type Metrics interface {
    Attempt(outcome string)
    PollError()
}

type Job struct {
    DedupeKey string
    DueAt     time.Time
}

var ErrNoWork = errors.New("no work")

func PollOnce(ctx context.Context, store Store, sender Sender, metrics Metrics, now time.Time) error {
    job, err := store.Claim(ctx, now)
    if errors.Is(err, ErrNoWork) {
        return nil
    }
    if err != nil {
        metrics.PollError()
        return err
    }

    if err := sender.Post(ctx, job); err != nil {
        metrics.Attempt("failed")
        return store.ScheduleRetry(ctx, job.DedupeKey, now)
    }

    metrics.Attempt("acknowledged")
    return store.Acknowledge(ctx, job.DedupeKey)
}
Enter fullscreen mode Exit fullscreen mode

This sketch deliberately does not claim exactly-once delivery. The claim expires or is otherwise recovered after a worker dies, and the same key survives that recovery. Backoff, jitter, attempt ceilings, and dead-letter handling are policy decisions that need tests around clock boundaries and concurrent claims; inventing constants in a reusable example would disguise the capacity assumptions.

Test the ugly interval.

Force a timeout after the receiver accepts a request but before the sender reads the response, restart the poller between Post and Acknowledge, and run two claimers against one due record. Then repeat the case with an acknowledgement that arrives after the claim lease expires, because that ordering exposes a distinction a happy-path test misses: the second worker may be correct to retry even though the first delivery took effect. The invariant is stable through all three cases: many attempts may exist, but one logical record remains open and one idempotency key crosses the boundary. The assertions should inspect both durable state and emitted telemetry; checking only the final HTTP status can pass while the monitor still counts the attempts as separate failures.

Rollback and ownership are coupled

A feature flag is an operational control, not a substitute for deployment or observability discipline. Fowler's feature-toggle taxonomy distinguishes release decisions from other toggle categories and warns that toggle management carries complexity. For a pricing rollout, record which rule version and cohort produced each evaluation, so rollback can stop new exposure without erasing evidence about requests already in flight.

The SLO should describe the learner-facing outcome, such as successful pricing evaluations within an agreed latency, while the webhook sender gets supporting indicators. If user-visible evaluation failures consume the rollout budget, freeze or reverse exposure. If evaluations succeed but a secondary notification destination is slow, contain delivery load and investigate that path; rolling back the rule may add risk without changing the failing dependency.

A page must map to an owned action. The application team can own malformed payloads and rule-version regressions. The platform team can own poller availability, queue saturation, and the delivery substrate. A destination outage may require a rate-limit or retry-policy response. Put the owner and action in the alert annotation, then verify during review that disabling the flag actually changes the signal being paged.

The buy-versus-build choice does not change these invariants, but it changes who operates them:

Concern Managed delivery or alerting Self-hosted path Acceptance test
Idempotency boundary Confirm where keys are stored and how long they remain effective Define schema, uniqueness, retention, and recovery Replay the same logical event across a worker restart
Retry control Inspect configurable backoff, ceilings, and visibility Operate schedulers, queues, and dead-letter handling Simulate destination timeout and recovery
On-call load Provider runs part of the control plane; integration ownership remains Team owns the full failure domain Trace one page to one runbook action
Lock-in Export events, state, and telemetry through documented interfaces Preserve portable schemas and standard telemetry Rebuild an unresolved-failure view from exported data
Capacity Validate quotas and overload behavior Model peak claims, attempts, storage, and worker concurrency Load test retry amplification, not only steady state

No row selects a winner. The deciding evidence is whether the system preserves logical identity, exposes state transitions, and behaves predictably during the failure modes the team can afford to exercise.

There are real limitations. A durable dedupe store is a poor fit for fire-and-forget signals whose duplication has no material effect, because the state, retention policy, and recovery path add operational weight. For those signals, an at-most-once send with a loss-tolerant dashboard may be the honest choice. The design is also insufficient when the receiver cannot accept an idempotency key: the sender can suppress its own concurrent attempts, but it cannot prove that a timed-out remote write did not take effect. In that case, choose a receiver-side lookup or reconciliation workflow before claiming duplicate-free delivery. This is the central trade-off: stronger suppression reduces repeated effects while requiring longer-lived identity and more coordinated ownership.

The threshold has an operational price

An alert on any failed attempt catches transient transport noise early and trains responders to distrust the pager. An alert only after retries are exhausted can arrive after the rollout has spent its error budget. The useful design is multi-window: attempts and polling errors feed dashboards or lower-urgency notifications, while a sustained rate or age of unresolved logical failures drives the page.

Short windows detect abrupt regressions. Longer windows keep one brief dependency wobble from waking someone. Both must be evaluated against real traffic, including low-volume periods where ratios become unstable; absolute counts, minimum-volume guards, and oldest-failure age can provide context. Review the resulting pages against two questions: would rollback reduce current user harm, and can the named owner act before the SLO budget is gone?

False positives are not free. They consume on-call attention, encourage broad flag rollback, and can make a pricing experiment harder to reason about because cohort assignment changes while old deliveries remain in flight. False negatives carry the opposite cost: learners may receive the wrong checkout behavior without a timely intervention. Threshold tuning is therefore a rollback decision with capacity and human-load consequences, not a cosmetic monitor setting.

One page, one logical failure set, one clear action. Keep every attempt for diagnosis, but never let attempts vote independently on incident severity.

Further reading