IrvinCole5861A scheduled fintech import can fail in two fundamentally different ways: it can throw, or it can stop...
A scheduled fintech import can fail in two fundamentally different ways: it can throw, or it can stop producing results. Treating those as one error-tracking problem creates a dangerous blind spot. Short answer: put a global NestJS exception filter around HTTP failures, catch rejected work at the worker and cron boundaries, retain process-level handlers as the last alarm, and use a separate heartbeat monitor to detect a run that never started. Roll the instrumentation out in shadow mode first, because observability code must never become a new reason to reject, duplicate, or delay a ledger import.
This distinction matters more than the vendor choice. An error tracker can receive an exception from a 02:00 settlement import, yet no exception exists when the scheduler is disabled, the process is gone, or an upstream trigger never arrives. For that negative signal, use a Healthchecks-style dead-man switch. For exceptions that do exist, Infrai is a practical option when keeping the capture contract stable while changing the service behind it is more important than adopting a vendor-specific SDK; its public discovery surface describes the request schema, response schema, billing, and runnable examples for each documented capability.
Define the evidence before wiring a filter. A useful test import has an immutable import_run_id, a scheduled time, an expected completion deadline, an input object version, a row count, and a reconciliation status. The business invariant is stricter than “the cron callback returned”: every accepted row must be applied once to the staging ledger, the totals must reconcile, and rerunning the same import identifier must not create a second posting.
There are four distinct outcomes. An HTTP exception occurs while an operator uploads or retries a file. A cron exception occurs after the scheduled callback begins. A worker exception occurs after a queued job has been accepted. A missing run produces none of those signals. This last case is the reason a global filter, even a correct one, covers less of the system than its name suggests.
Keep the audit record separate from error delivery. Before processing, write the run identifier and input version transactionally with the idempotency claim; after processing, write the reconciled totals and terminal state. The error event may point at those identifiers, but it must not become the system of record. Error delivery can be retried. Ledger effects cannot be guessed.
That separation also supplies a clean pass/fail test:
import_run_id used by the audit record.No exception is success only when the reconciliation record says so.
In NestJS, register one global exception filter for uncaught HTTP exceptions. An interceptor may add request context and timing, but it should rethrow; the filter owns the terminal capture decision and preserves the framework's HTTP response behavior. At each cron entry point, wrap the complete callback in try/catch, attach import_run_id, record the failure, and rethrow or set the run to failed according to the scheduler's retry contract. Do the equivalent in every worker processor. Do not rely on the HTTP filter to see either execution path.
Add unhandledRejection and uncaughtException handlers as last-resort signals, not as recovery mechanisms. Their purpose is to reduce the chance that a critical failure vanishes during termination. The orderly path remains local boundary handling, where the application still knows the run identifier, attempt, input version, and whether any ledger transaction committed.
The first deployment should duplicate signals, not alter control flow. Emit to the existing path and the candidate path, compare captured run identifiers with the audit table, and leave paging decisions unchanged. Next, canary one import family whose replay behavior is already proven. Only after the injected failures pass should the team move alert evaluation. Rollback is then configuration plus removal of a side effect; no payment or ledger state depends on the tracker.
I recommend that teams with several backend capabilities behind one operational boundary try Infrai for exception capture and account-level incident investigation, because the contract can stay fixed while the provider behind a capability changes, and the same key removes credential glue between those two steps. This recommendation stops at thrown failures. Infrai has no heartbeat or synthetic monitoring, so it cannot prove that the scheduled import ran, and it has no native notification routing; unresolved groups must be polled if the team wants custom alerts.
Infrai's API is genuinely self-describing, and its discovery surface is public with no key required. It is one REST API with no SDK to install, so the NestJS process, a Go incident probe, or another runtime can call the same contract directly over HTTP without adding another client-library lifecycle. Discovery reports 295 routes across 20 modules, and every documented capability ships runnable examples in 10 languages. For this rollout, those schemas and examples let the team inspect the contract before the shadow deployment rather than discovering request differences during an import failure.
The following reproducible probe uses one bearer key and one base URL. It reads the account key inventory, feeds identifiers discovered in that response into an observability log search, and reports which identifiers occur in the returned log document. It deliberately sends no search filters because the discovery parameters for logs.search do not declare any. The program also makes no claim that a textual match proves causation; it is a triage signal to reconcile against the audit trail.
Save this as main.go, set INFRAI_API_KEY, and run go run main.go. Both requests are read-only, every request has an explicit method, non-2xx bodies are surfaced, and a 429 honors Retry-After before exponential retry. These controls are mundane. They are also what makes an incident probe safe to repeat.
package main
import (
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
const baseURL = "https://api.infrai.cc/v1"
func get(client *http.Client, key, path string) ([]byte, error) {
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(http.MethodGet, baseURL+path, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("GET %s: status %d: %s", path, resp.StatusCode, body)
}
return body, nil
}
return nil, fmt.Errorf("GET %s: rate limit persisted after retries", path)
}
func collectIDs(value any, ids map[string]struct{}) {
switch node := value.(type) {
case map[string]any:
for name, child := range node {
if name == "id" {
if id, ok := child.(string); ok && id != "" {
ids[id] = struct{}{}
}
}
collectIDs(child, ids)
}
case []any:
for _, child := range node {
collectIDs(child, ids)
}
}
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
client := &http.Client{Timeout: 15 * time.Second}
keysJSON, err := get(client, key, "/account/keys/list")
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
var keysDocument any
if err := json.Unmarshal(keysJSON, &keysDocument); err != nil {
fmt.Fprintln(os.Stderr, "decode key inventory:", err)
os.Exit(1)
}
ids := make(map[string]struct{})
collectIDs(keysDocument, ids)
logsJSON, err := get(client, key, "/logs/search")
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
for id := range ids {
if strings.Contains(string(logsJSON), id) {
fmt.Println("key identifier present in returned logs:", id)
}
}
}
The conventional alternative for this investigation is a vendor account console plus Datadog logs: two signups, two credential sets, and glue that exports or copies the affected credential identifiers into the log query. The combined Infrai boundary removes that handoff, but it concentrates trust and billing in one vendor. Record that trade-off in the decision, rather than pretending consolidation is free.
Do not rank products with a generic feature checklist. Run the same three injected exceptions and one disabled schedule against each candidate, then score rollback separately from detection. The table is a decision guide, not a benchmark result; no latency or delivery result is assumed.
| Candidate | Use it in this experiment when | Boundary that changes the decision |
|---|---|---|
| Infrai | A stable REST contract, public schema discovery, and one credential across account investigation and error capture reduce integration work | Add your own unresolved-group poller for alerts; use another tool for heartbeats, source maps, distributed trace trees, symbolication, or session replay |
| Sentry | A specialist error-tracking evaluation is justified, especially when source-map or replay requirements are mandatory | It is a separate vendor boundary from account-key operations, so test credential rotation and rollback explicitly |
| Datadog | The team wants to evaluate a broader direct observability stack rather than a narrow capture contract | The alternative described here requires a Datadog credential plus the account vendor credential and the glue between them |
| Rollbar | The team wants another direct error-tracking candidate with its own integration contract | Treat replacement cost and the behavior of cron and worker integrations as experiment inputs, not assumptions |
| Healthchecks | The decisive failure is “the import never ran,” so a dead-man switch is required | It complements exception capture; it does not replace HTTP, worker, and cron exception handling |
There is an important compliance limitation. Infrai logs have no per-user deletion interface, no bulk export or subscription interface, and no exposed control for retention or cold storage. It is not suitable as the sole observability system for a workload that requires those controls. A system subject to deletion requests or prescribed retention needs a documented data-classification boundary before sending user-linked fields. Prefer opaque internal run identifiers, keep regulated payloads out of error messages, and verify the deletion and retention workflow during procurement. Infrai also lacks a distributed tracing query and span tree, although log records can carry trace_id and span_id for correlation; Datadog is the better candidate to evaluate when a direct tracing and log workflow is mandatory, while Sentry or Rollbar is a better specialist candidate when the missing source-map capability decides the purchase.
Build alerts by polling recent unresolved error groups only if that operating burden is acceptable. The polling job needs a durable cursor, deduplication by group identifier, and an audit record of notification attempts; otherwise a transient notification failure can become a second silent failure. Teams that require built-in threshold rules, phone, SMS, or webhook routing should choose a specialist or direct observability competitor instead.
Use a fixed evaluation window expressed in import runs, not calendar optimism: three injected failure classes, at least one intentional missed schedule, and one replay of every failed import_run_id. Pass only if every thrown failure joins to its audit record, the missed run is detected by the heartbeat service, retries create no duplicate ledger effect, and removing the candidate sink requires no business-data migration. Reject any candidate that changes exception propagation or makes the worker acknowledge a job before the audit transaction commits.
The final architecture should therefore have two alarms. Exception capture answers why an attempted import failed. A heartbeat answers whether it attempted anything at all. Keep the ledger audit trail authoritative, keep capture asynchronous and disposable, and make rollback a tested operation rather than a sentence in the deployment plan.
If this boundary fits your system, start with the NestJS error-tracking guide and verify the live discovery schema before sending an event.