SunspireValerius59Short answer: Connect frontend and backend logging by sending an opaque request ID with each browser...
Short answer: Connect frontend and backend logging by sending an opaque request ID with each browser fetch, echoing it in the server response, and recording it in structured logs alongside a separate ID for the full scheduled-import run.
Keep the default trail narrow. A correlation key, run state, duration, result count, and failure class usually answer the first operational question without retaining request bodies or customer records.
The dominant cost is normally log volume entering storage, not the few bytes occupied by a request ID. Model it before adding fields: runs x events per run x average encoded event size x retention copies. For a concrete capacity example, 200,000 runs per day, six events per run, and 900 bytes per event produce about 1.08 GB of raw events per day before indexing, replication, or compression. This is arithmetic for planning, not a production benchmark or a quoted price. Per-GB ingestion pricing makes event count and encoded size the levers worth measuring; retention and query charges depend on the system selected.
The useful change is selective detail, not absent logging. Emit a compact lifecycle event for every run, retain error details only under a documented policy, and measure metrics for the broad health signal. You deliberately stop keeping successful payloads, repeated progress messages, and unbounded browser context. The price of that choice appears during an investigation: you can prove where a run stopped, but you may be unable to reconstruct every input value after the source data expires.
Treat three identities as separate fields. request_id follows one browser-to-service exchange. import_run_id follows the scheduled job after the initiating request has returned. trace_id, when distributed tracing is present, follows work across process boundaries. Reusing one value for all three looks convenient until a retry, fan-out, or queue consumer creates a one-to-many relationship.
For an import monitor, the stable join is the run ID. The browser can receive it in a response body, while the request ID travels in a response header for support and diagnostics. Server events carry both during the initial exchange; background events keep the run ID and add their own request or message identity. This preserves causality without pretending that an asynchronous job is still part of an open browser request.
The identifier is metadata, not authority. Do not use it to authorize access, expose sequential database keys, or place account numbers, email addresses, phone numbers, or other regulated data inside it. Generate an opaque value, validate a client-supplied value against a tight length and character policy, and replace anything invalid. That small boundary matters because browser headers are attacker-controlled input.
Consider a rollback drill with two frontend versions and two backend worker versions active at once. The older browser sends no header, so the server creates a request ID. The newer browser creates one before its fetch call. Both backend versions echo the effective value, but only the newer producer adds import_run_id; during the overlap, the consumer accepts either the new field or the earlier job key. A retry from the newer browser gets a fresh request ID because it is a fresh network exchange, while an idempotency key, if the application uses one, remains a separate concern. The scheduled worker continues with the run ID after the HTTP exchange closes. An operator starting from a browser error can search server logs by request ID, find the accepted event, pivot to the run ID, and then inspect the terminal event. An alert starting from import silence takes the reverse route: begin with the last successful run, inspect the next expected run, and use its request or message identifiers only when narrowing the failed boundary. This drill exposes three common mistakes before production: treating a request ID as an idempotency control, dropping the run ID at an asynchronous handoff, and making a new field mandatory before every reader understands it. It also gives rollback a crisp acceptance test. Either deployed version must leave a searchable path from the frontend report to backend state, even when richer fields are temporarily disabled.
A concise event set is enough:
| Event | Required correlation | Operational answer |
|---|---|---|
import.accepted |
request ID and run ID | Did the service accept the run? |
import.completed |
run ID | Did it produce results, and how many? |
import.failed |
run ID and failure class | Where should investigation begin? |
| silence metric or alert | schedule and last successful run | Has expected output stopped? |
Silence itself is not a log line. A missing event cannot emit an event about its own absence. Track the last successful completion as a metric or queryable state, then alert when its age exceeds the schedule plus a deliberate grace period. The Google SRE monitoring guidance distinguishes symptoms from causes and frames latency, traffic, errors, and saturation as useful signals. Here, stale successful output is the symptom; detailed logs remain evidence for the cause.
The following Python example shows the contract without binding it to a logging vendor. It uses the standard library, rejects malformed incoming IDs, emits JSON, and returns both identifiers. In a real service, the scheduler or queue producer would create the run ID when no browser request exists.
import json
import logging
import re
import time
import uuid
from collections.abc import Mapping
REQUEST_ID_PATTERN = re.compile(r"^[A-Za-z0-9_-]{16,64}$")
logger = logging.getLogger("imports")
def request_id(headers: Mapping[str, str]) -> str:
candidate = headers.get("X-Request-ID", "")
if REQUEST_ID_PATTERN.fullmatch(candidate):
return candidate
return uuid.uuid4().hex
def log_event(event: str, *, request_id: str, run_id: str, **fields: object) -> None:
record = {
"event": event,
"request_id": request_id,
"import_run_id": run_id,
**fields,
}
logger.info(json.dumps(record, separators=(",", ":"), sort_keys=True))
def start_import(headers: Mapping[str, str]) -> tuple[dict[str, str], dict[str, str]]:
req_id = request_id(headers)
run_id = uuid.uuid4().hex
log_event(
"import.accepted",
request_id=req_id,
run_id=run_id,
accepted_at_unix_ms=time.time_ns() // 1_000_000,
)
response_headers = {"X-Request-ID": req_id}
response_body = {"import_run_id": run_id, "state": "accepted"}
return response_headers, response_body
The browser-side JavaScript contract is small even though the example implementation is Python: create an opaque request ID if the application owns that convention, send it in X-Request-ID with fetch, read the echoed response header, and record the returned import run ID with the UI action. A Node.js backend can apply the same boundary contract before its logging middleware writes server logs. If cross-origin browser code must read that response header, the server's CORS policy has to expose it to the permitted origin. Keep that policy explicit and test it at the HTTP boundary; correlation should work across the stack, not only inside one process.
Do not log the full import row merely because serialization is easy. Delivery systems teach the same lesson: addresses, phone numbers, OTPs, and provider responses can turn a diagnostic stream into a second sensitive database. Prefer bounded fields such as result_count, duration_ms, source_kind, and a low-cardinality failure_class. Store a payload reference only when access control, deletion, and retention are defined elsewhere.
Logging changes can break operations without breaking the import. A new high-cardinality field can make queries expensive; stricter validation can sever correlation with an older client; synchronous emission can add latency. Rollback safety means the old and new event shapes overlap long enough for readers, alerts, and dashboards to tolerate both.
Deploy additive fields first. Update consumers to accept the new field while retaining the old lookup path, then update producers. For one release window, query by either run identifier representation and compare completion counts. Remove the old field only after the scheduled-job interval, late-arrival allowance, and retained investigation window have all passed. The exact duration is a system decision, so encode it in the rollout plan rather than borrowing a generic number.
One switch should disable optional detail without disabling the lifecycle events. Another can stop accepting caller-provided request IDs while preserving server-generated IDs. Those switches give rollback a narrow blast radius: correlation becomes less rich, but the evidence that an import was accepted, completed, or failed remains.
Test the rollback path before deployment. Send a valid ID, an overlong ID, invalid characters, and no ID; verify that every response has a valid echoed value. Simulate two browser retries for one logical action and confirm that they receive distinct request IDs while the application can still associate them with the intended run. Then roll a worker back while a job is in flight and verify that both versions emit events a shared parser accepts.
No heroics. A log schema is an interface, and mixed versions are its normal operating condition during a rollout.
The alert should evaluate the expected production of results, not the presence of arbitrary activity. A worker can emit heartbeats while producing zero accepted records. Record the latest successful run that met the business definition of useful output, then compare its age with the import schedule. Keep result_count in the completion event so a zero-result success can be classified according to the feed's contract rather than guessed during an incident.
Rate limits and retries deserve explicit treatment. One scheduled run may attempt several upstream calls, but repeating a detailed line for every retry multiplies ingestion and obscures the final outcome. Count attempts in memory, emit the bounded count at completion, and emit individual attempt failures only when they cross a sampling or severity rule. Never sample the terminal lifecycle event.
A practical retention split follows the questions each signal answers. Metrics cover long-range trends and alert evaluation. Compact lifecycle events cover routine investigations. Restricted error detail lives for the shortest period that meets operational and compliance requirements. Payloads stay out of the logging path. This separation moves the dominant volume term while keeping the joins needed to locate a failed stage.
The trade-off is real: once detailed events expire, an old incident may be explainable only at the level of state, timing, count, and failure class. If legal, audit, or dispute handling requires more, define a separate governed record with its own access and deletion rules. Extending all debug-log retention is a blunt substitute.
Adopt the design only if it passes four checks. The run ID survives asynchronous boundaries. The request ID is echoed and searchable but grants no privilege. Terminal events are unsampled and small. The previous schema and the optional-detail switch provide a tested rollback route.
Then inspect the volume equation with representative encoded events before rollout. If lifecycle events, retry detail, or indexing overhead dominate, reduce event count and field cardinality first. Do not optimize away the join keys. They are the few bytes that turn a browser report, an accepted import, and a silent scheduler into one defensible timeline.