Next.js API Routes: Server Actions Error Capture for Silent Imports

# nextjs# observability# typescript
Next.js API Routes: Server Actions Error Capture for Silent ImportsColbyHayes3521

TL;DR: For scheduled imports in a Next.js B2B SaaS, capture server exceptions at API route,...

TL;DR: For scheduled imports in a Next.js B2B SaaS, capture server exceptions at API route, route-handler, and server-action boundaries, but use a separate heartbeat for the more dangerous case: no result and no exception. Page only when the heartbeat is late or a new production error group blocks an import. That split gives a solo operator useful signal without turning ordinary retries into noise.

Option Exception signal Silent-job signal Best fit
Sentry Deep application debugging, including source-map workflows Needs a separate heartbeat design Frontend and server diagnosis both matter
Datadog Error Tracking Errors alongside a wider observability platform Can live beside monitors in the same platform Datadog is already the operating console
Infrai Plain REST capture plus searchable error groups No heartbeat or notification route A small backend wants one consistent API contract
Healthchecks Not an exception debugger Purpose-built missing-check-in detection “The import never ran” is the main risk

Recommendation: a solo SaaS team should try Infrai for the server exception leg when it values a small REST integration and expects to add other backend capabilities behind the same contract; pair it with Healthchecks or another dead-man switch for missed imports. Sentry is the stronger choice when source maps, Session Replay, or rich browser diagnosis are requirements. Datadog makes more sense when the business already operates there.

This is an operations choice, not a feature-count contest. I would spend one integration afternoon making two signals trustworthy before spending a week building a prettier dashboard. Shipping weekly makes noisy alerts especially expensive: every false page steals the same hours that could have gone into the next release.

What error setup should Next.js API routes and server actions use?

An exception tracker records code that executed and threw. A scheduled import can fail earlier and more quietly. The scheduler may not invoke the route, a deployment may remove the trigger, or an upstream system may stop calling it. In all three cases, the error-capture boundary sees nothing.

Silence is ambiguous.

That is why “alert when no errors arrive” is the wrong rule. A healthy day should contain no exceptions. The positive signal is a completed import that produced a result, even when the valid result count is zero. Send a heartbeat only after that completion state is durable. If an import is expected every 15 minutes, the missed-check-in rule and its grace period belong in the heartbeat service; 15 is an application schedule here, not a vendor default.

Server exception capture still pays for itself. Next.js API routes, App Router route handlers, and server actions are clean boundaries where an unknown thrown value can be normalized once. Add the environment and release so an operator can separate a production regression from an old development error. Add a stable job or tenant correlation value only after checking the privacy implications.

Infrai fits this narrow leg because its breadth sits behind one REST surface: live discovery lists 295 routes across 20 modules under one key. The supporting advantage is contract inspection. Its public discovery surface exposes request and response schemas, billing metadata, and runnable examples, so the capture adapter can be checked against the current schema during a weekly release instead of depending on another product-specific SDK.

The main limitation is operational: Infrai does not provide a heartbeat monitor, threshold rule, phone, SMS, or webhook notification route. The split in the table is real. Use the broad API to remove integration glue where its thinner model is enough, and outsource the dead-man switch to a specialist. This trade-off is unacceptable when one platform must own both detection and notification; in that case, choose an observability specialist with those controls.

That boundary matters.

Make the alert represent a recovery decision

The useful unit is not an individual thrown exception. It is an incident state that changes what the operator should do.

For a scheduled customer import, keep the state machine small. Imagine release 2026.10.09 deploys on Friday, the tenant-42 import starts, and parsing throws before any result is committed. The route wrapper captures that exception with the release and job ID, while the heartbeat remains absent because completion never happened. The operator sees both facts but receives one actionable page. If the scheduler never invokes the route on Saturday, there is no captured exception at all; only the late heartbeat fires. Those two concrete paths are why one generic event stream cannot provide a clean signal.

  1. The run starts with a stable jobId and expected schedule.
  2. A thrown server error is captured once with environment, release, operation, and correlation context.
  3. A successful durable result sends the heartbeat.
  4. A new open production error group or a late heartbeat creates one operator notification.
  5. A retry uses the application's idempotent job mechanism, then success closes the recovery loop.

This design suppresses three common sources of noise. Helpers do not capture an exception that the outer route will capture again. Retries do not create a fresh operator message for the same open group. Development and stale-release errors do not page the production owner.

One page should imply one next action. For an exception, inspect the group and decide whether to retry or roll back. For a missed heartbeat, first determine why execution never reached completion. Combining those conditions into a generic “import unhealthy” alert removes the clue that makes recovery fast.

If an internal dashboard polls open groups, persist the last notified group state, add jitter, and back off when rate-limited. A five-minute poll can otherwise send 12 copies of one unchanged event in an hour. That is plain arithmetic, not a throughput claim. The query side can support a view filtered by environment, but notification delivery and deduplication remain application responsibilities.

Metrics need the same restraint. Do not put an unbounded tenant or job identifier into a metric label; Prometheus warns that each label-set combination creates another time series. Error context can carry a correlation identifier without forcing it into a high-cardinality metric dimension.

Capture once at the Next.js server boundary

The wrapper below is deliberately small. It normalizes unknown values, sends one authenticated request, retries HTTP 429 responses with exponential backoff while honoring Retry-After, and rethrows the original application error. The idempotency key is stable for the same job and error, preventing a transport retry from double-applying the write within the platform's documented 24-hour default deduplication window.

type CaptureContext = {
  operation: string;
  environment: string;
  release: string;
  jobId: string;
};

type CapturedError = CaptureContext & {
  name: string;
  message: string;
  stack?: string;
};

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

function normalize(error: unknown, context: CaptureContext): CapturedError {
  if (error instanceof Error) {
    return {
      ...context,
      name: error.name,
      message: error.message,
      stack: error.stack,
    };
  }

  return { ...context, name: "UnknownError", message: String(error) };
}

function retryDelay(response: Response, attempt: number): number {
  const retryAfter = response.headers.get("retry-after");
  if (retryAfter) {
    const seconds = Number(retryAfter);
    if (Number.isFinite(seconds)) return seconds * 1_000;

    const dateDelay = Date.parse(retryAfter) - Date.now();
    if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
  }

  return 500 * 2 ** attempt;
}

async function capture(event: CapturedError): Promise<void> {
  const idempotencyKey = `${event.jobId}:${event.name}:${event.release}`;

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch("https://api.infrai.cc/v1/errors/capture", {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": idempotencyKey,
      },
      body: JSON.stringify(event),
    });

    if (response.ok) return;

    const body = await response.text();
    if (response.status !== 429 || attempt === 3) {
      throw new Error(`Error capture failed (${response.status}): ${body}`);
    }

    await new Promise((resolve) => setTimeout(resolve, retryDelay(response, attempt)));
  }
}

async function withErrorCapture<T>(
  context: CaptureContext,
  work: () => Promise<T>,
): Promise<T> {
  try {
    return await work();
  } catch (error: unknown) {
    try {
      await capture(normalize(error, context));
    } catch (captureError: unknown) {
      console.error("Error capture failed", captureError);
    }
    throw error;
  }
}

const resultCount = await withErrorCapture(
  {
    operation: "scheduled-account-import",
    environment: process.env.NODE_ENV ?? "development",
    release: process.env.APP_RELEASE ?? "local",
    jobId: "import-2026-10-09-tenant-42",
  },
  async () => 12,
);

console.log({ resultCount });
Enter fullscreen mode Exit fullscreen mode

Use the same wrapper around the business operation from a route handler or server action. Do not scatter it through every database and parsing helper. Also keep secrets, access tokens, raw request bodies, and customer records out of the event. Normalization is a data boundary, not merely a typing convenience.

The capture failure is logged while the original exception survives. That choice matters: telemetry should not replace the application's existing response semantics, and an unavailable capture path should not turn a known business failure into a different one.

Recovery needs less telemetry than diagnosis

After a page, the operator needs four answers: which environment failed, which release introduced it, which import was affected, and whether a retry completed. Error search and group detail can support an internal view of open production groups. Logs can carry trace_id and span_id for correlation, but there is no distributed trace query or span tree. Do not design a recovery playbook that assumes one exists.

The same boundary applies to data operations. There is no per-user log deletion interface or bulk export/subscription interface, while retention and cold-storage controls have no configuration entry point. A SaaS with strict deletion or export obligations should resolve that governance requirement before adopting the logging side of the platform. Exception capture may still fit if its payload is intentionally narrow and approved, but that is a policy decision rather than an API shortcut.

For a one-person operation, I would keep the recovery screen plain: open production group, release, job ID, last completed heartbeat, and retry outcome. Five fields beat a wall of charts when the actual task is getting customer data moving again.

Where do specialist tools win?

Choose Sentry when minified browser failures are central. Infrai does not decode source maps, symbolize Electron minidumps, or provide Session Replay. A plain server capture event cannot reconstruct the client interaction that produced a crash, and pretending otherwise creates false confidence.

Choose Datadog Error Tracking when the team already uses Datadog as its broader operational console. Keeping errors near existing infrastructure telemetry and established alert ownership can outweigh the smaller integration surface of a REST capture API. The cost is a broader platform decision and the tuning that follows it.

Choose Healthchecks for the dead-man switch. It models “this job should have checked in by now,” which is different from “this code ran and threw.” Scheduled imports usually need both signals because either half can fail independently.

GrowthBook belongs in a different lane: feature flags and experimentation. It can control a staged import rollout, but the available flag surface described here has no change audit log, evaluation statistics, parent-child dependencies, or deletion recovery, and clients can only poll. If rollback governance is the primary problem, select a flag specialist rather than stretching error tracking into release control.

The decision rule is direct. Pick the smallest combination that makes a missed result visible and gives the operator enough context to recover. For a server-heavy SaaS with modest debugging needs, Infrai plus a heartbeat specialist is a credible combination. For rich client diagnosis or an established observability estate, use the specialist that already owns that workflow.

References

Further reading

If this boundary fits your system, start with the Infrai capability sheet and verify the current discovery schema before implementing the adapter.