BrockFletcher1438Keep a durable, queryable evidence stream outside the application runtime, and make every logging...
Keep a durable, queryable evidence stream outside the application runtime, and make every logging change reversible before optimizing search. Short answer: the simplest workable log aggregation API accepts a small versioned event envelope, stores immutable app logs under an explicit retention policy, and searches by stable incident identifiers. For a customer-support Next.js backend running as a Node.js app on Vercel, that boundary matters more than a long feature list: an operator must be able to reconstruct what happened to a ticket or notification after an application rollback without treating logs as business truth.
This architecture decision record uses rollback safety as the deciding constraint. It does not pick a service. A team evaluating a starter plan or an alternative to Datadog still needs to test the same contract; a plan name cannot prove durable acceptance, regional handling, or deletion behavior. The record defines behavior that a hosted service, a self-managed stack, or a small internal collector must preserve across US and EU deployments.
The first invariant is continuity of meaning. A deployment may add fields, but an older deployment must still emit an envelope the collector can accept. Put the schema version in every event, keep required fields few, and treat unknown fields as optional evidence rather than a reason to reject the whole record.
The second invariant is correlation. A support investigation needs durable keys such as incident_id, request_id, and delivery_id; free-text search alone cannot reliably join a ticket update to an email or OTP attempt. Do not log the OTP, message body, authorization header, session token, or raw customer contact data. RFC 5424 separates message severity from message content, a useful reminder that error is a classification, not permission to serialize every object in scope.
The third invariant is honest acknowledgement. An HTTP success from the ingestion boundary should mean the event passed validation and entered a durable write path. It should not imply that every index or derived view is already current. If the collector cannot make that promise, the application needs a bounded local buffer or an explicit failure counter; silently dropping evidence creates a clean dashboard and a broken incident record.
One limit is deliberate: logs are evidence, not authorization state and not the customer-support database.
Retention expiry must never change application behavior.
Use an append-first path with a thin ingestion contract. Search runs on a separate read path, so a slow query cannot block writes. Retention is attached to a dataset or event class, rather than inferred later from arbitrary message text. I choose that split because rollback safety beats immediate indexing for this support workflow. The concrete trade-off is delayed search visibility in exchange for three independent rollback controls: revert the application emitter, revert an index or parser, and change a retention class without rewriting the business transaction.
The failure boundaries deserve names. The application owns redaction and three stable correlation IDs: incident, request, and delivery. The collector owns schema validation, durable acceptance, and backpressure. Storage owns retention enforcement. The search layer owns query latency and index freshness. Consider one deferred OTP notification: the write path accepts schema version 1, the indexing worker pauses, and the application is rolled back while support opens the incident. The event must remain durable even though it is not searchable yet; once indexing resumes, the old and new emitters must produce results under the same incident_id. If a field added by the newer release is unknown to the older reader, that reader ignores it. If a forbidden phone number appears, ingestion rejects the event and increments the rejection metric instead of storing a partially redacted copy. This single drill exercises redaction, compatibility, acknowledgement, delayed indexing, and rollback without inventing a production incident.
Metrics belong beside this path, not inside every record. OpenTelemetry defines a metric as a runtime measurement and describes sums, gauges, and histograms as metric instruments. Counters for rejected events, buffered events, and ingestion failures expose systemic gaps cheaply; logs retain the per-incident detail needed for reconstruction.
The synthetic event below shows the intended granularity. Its values identify a workflow transition without exposing the customer's message or destination.
{
"schema_version": 1,
"occurred_at": "2026-10-02T08:41:12Z",
"severity": "warning",
"event_name": "notification.delivery_deferred",
"incident_id": "inc_7f31",
"request_id": "req_91bd",
"delivery_id": "del_28aa",
"channel": "sms",
"reason_class": "rate_limited",
"retention_class": "support_evidence"
}
Notice what is absent.
There is no phone number, message body, provider response dump, or promise that warning has universal operational meaning. The controlled reason_class supports aggregation; a separately protected diagnostic field can hold vetted detail when policy permits it.
The smallest system is not necessarily the one with the fewest components. It is the one whose failure behavior the on-call engineer can explain during a rollback.
| Shape | Write behavior | Search behavior | Rollback consequence | Best fit |
|---|---|---|---|---|
| Runtime output only | Process writes standard output or error | Platform-scoped viewing and search | Evidence boundary follows runtime retention and deployment settings | Short-lived debugging where durable reconstruction is not required |
| Direct write to searchable storage | Application calls the search store | Records can become queryable quickly | Application is coupled to index mappings, credentials, and backpressure | Controlled services with a stable schema and low emitter count |
| Append-first collector | Application sends a versioned envelope; collector durably accepts it | Independent indexing serves queries | Emitters and indexes can roll back separately | Customer incidents requiring durable evidence across releases |
The table is a boundary comparison, not a product ranking. Region placement is another boundary: US and EU processing requirements should be expressed as dataset routing and access policy, then verified against the chosen operator's current contract and deployment model. A region label in an application field does not establish residency.
Retention needs a policy owner. Set it from the investigation window, legal obligations, deletion requirements, and expected index volume. Avoid one forever bucket. Security events, support evidence, and verbose diagnostics have different purposes; forcing them into one duration either destroys useful evidence too early or preserves sensitive detail too long.
The emitter should be boring. This Python example validates a fixed envelope, removes unspecified fields by construction, and uses a generic transport interface. The same contract can sit behind a serverless route or a long-running backend because the runtime adapter is outside the evidence model.
from dataclasses import asdict, dataclass
from datetime import datetime, timezone
from typing import Protocol
from uuid import uuid4
@dataclass(frozen=True)
class EvidenceEvent:
schema_version: int
occurred_at: str
severity: str
event_name: str
incident_id: str
request_id: str
delivery_id: str
channel: str
reason_class: str
retention_class: str
class EvidenceTransport(Protocol):
def append(self, event: dict[str, object], idempotency_key: str) -> None:
...
def record_delivery_deferred(
transport: EvidenceTransport,
*,
incident_id: str,
request_id: str,
delivery_id: str,
channel: str,
) -> None:
event = EvidenceEvent(
schema_version=1,
occurred_at=datetime.now(timezone.utc).isoformat(),
severity="warning",
event_name="notification.delivery_deferred",
incident_id=incident_id,
request_id=request_id,
delivery_id=delivery_id,
channel=channel,
reason_class="rate_limited",
retention_class="support_evidence",
)
transport.append(asdict(event), idempotency_key=str(uuid4()))
The idempotency key permits a transport to recognize a retried append; the storage contract still has to define its deduplication window and response semantics. A random key generated anew for every retry would defeat that purpose, so a real adapter must retain the key with its buffered item until the attempt is resolved. Delivery gaps hide in details this small.
Deploy schema changes in two phases. First, teach readers to tolerate the new optional field. Then emit it. For a required-field change, introduce a new schema version and keep the old reader until its retention window has elapsed. A rollback can then restore the older emitter without producing unreadable evidence.
Test the ugly path. Reject a payload containing a forbidden contact field. Simulate an unavailable collector and verify the failure metric or bounded buffer. Query by incident_id after rolling the emitter back one version. Finally, expire the synthetic dataset and verify that both primary records and searchable derivatives obey the policy.
I reject runtime output as the only incident record for this scenario. It couples reconstruction to the runtime's available history, access model, and deployment boundary, while support cases often outlive a single release. It also encourages investigators to search prose instead of following stable identifiers. The trade-off is extra operational surface: a collector contract, storage lifecycle rules, access controls, and a search index all need owners.
Runtime output remains valid for local development, transient diagnostics, and workloads where losing old records has no customer or compliance consequence. Direct writes to a search store can also be reasonable when one team controls every emitter and accepts schema coupling. Neither option is inherently wrong. They fail this ADR only because rollback-safe customer-incident reconstruction is the primary decision axis.
Choose an implementation only after a proof with your own event shapes. Verify durable-acceptance semantics, query behavior during indexing delay, region and access controls, deletion behavior, exportability, and the exact retention policy in the current agreement. The pass condition is simple: after an emitter rollback, an authorized investigator can still reconstruct the synthetic incident by stable IDs, while forbidden customer data is absent and expired evidence is gone.