UrbanDonovan1576TL;DR: Cheap hosted application logging for a Postgres SaaS is useful only when its searchable...
TL;DR: Cheap hosted application logging for a Postgres SaaS is useful only when its searchable records can connect the Node.js API, workers, and cron jobs for one AI tutor run. Test that path in a European region before choosing a service. Emit one structured event per boundary, calculate estimated cost from recorded token counts and a versioned rate table, and keep latency as both a total and named step durations. A low ingestion quote is irrelevant if retention, query scans, or uncontrolled payloads make attribution unreliable.
That gives a small team a useful unit of work: cost and latency per tutoring run, not a monthly pile of bytes. It also keeps the logging backend replaceable. A European region, searchable JSON, explicit retention controls, export, and predictable limits are requirements to verify during a trial; they are not reasons to organize application code around a provider SDK.
A learner asks for help through the API. The agent retrieves course context, calls a model, may call it again after a tool result, writes the final state to Postgres, and queues grading work. Cron jobs later expire abandoned sessions or aggregate usage. The data flow is plain: every boundary emits JSON to standard output, the runtime forwards it to a hosted log store in the required European region, and a search on run_id returns the ordered story.
Three identifiers do different jobs. request_id follows the inbound HTTP request. run_id survives queues and retries for the logical tutor run. step_id distinguishes retrieval, generation, persistence, and grading. A retry can create a new attempt while retaining the same run, which is exactly what cost attribution needs.
Events need service, environment, event, duration_ms, status, and a timestamp. Model events add input and output token counts plus a stable internal model key. Queue events add attempt number and wait time. Database events record an operation name, never raw SQL parameters or student text.
Keep secrets out.
No exceptions.
The runnable core is a narrow logger interface. This TypeScript example records a model step without importing a logging vendor, and it injects the rate table rather than pretending that prices never change. Rates are currency units per million tokens; the application owns the effective date and review process.
import { randomUUID } from "node:crypto";
import { performance } from "node:perf_hooks";
type Usage = { input: number; output: number };
type Rates = { inputPerMillion: number; outputPerMillion: number };
type Result = { text: string; usage: Usage };
type Generate = (prompt: string) => Promise<Result>;
type StepEvent = {
timestamp: string;
service: "tutor-api" | "grading-worker" | "maintenance-cron";
event: "agent.model.completed" | "agent.model.failed";
run_id: string;
step_id: string;
model_key: string;
duration_ms: number;
input_tokens?: number;
output_tokens?: number;
estimated_cost?: number;
status: "ok" | "error";
error_kind?: string;
};
const writeEvent = (event: StepEvent): void => {
process.stdout.write(`${JSON.stringify(event)}\n`);
};
const estimateCost = (usage: Usage, rates: Rates): number =>
(usage.input * rates.inputPerMillion +
usage.output * rates.outputPerMillion) /
1_000_000;
export async function runModelStep(
generate: Generate,
prompt: string,
modelKey: string,
rates: Rates,
runId = randomUUID(),
): Promise<Result> {
const stepId = randomUUID();
const started = performance.now();
try {
const result = await generate(prompt);
writeEvent({
timestamp: new Date().toISOString(),
service: "tutor-api",
event: "agent.model.completed",
run_id: runId,
step_id: stepId,
model_key: modelKey,
duration_ms: Math.round(performance.now() - started),
input_tokens: result.usage.input,
output_tokens: result.usage.output,
estimated_cost: estimateCost(result.usage, rates),
status: "ok",
});
return result;
} catch (error) {
writeEvent({
timestamp: new Date().toISOString(),
service: "tutor-api",
event: "agent.model.failed",
run_id: runId,
step_id: stepId,
model_key: modelKey,
duration_ms: Math.round(performance.now() - started),
status: "error",
error_kind: error instanceof Error ? error.name : "UnknownError",
});
throw error;
}
}
There is no prompt, completion, email address, or database row in that event. Debug payloads feel convenient until they expose learner data, multiply ingestion volume, and make deletion obligations harder. Store an approved template identifier and content hash when diagnosis needs provenance. Put full content in a separately governed system only when the product requires it.
That boundary is a trade-off. A hosted store is a poor fit when policy forbids third-party processing, when the team needs control over every storage layer, or when sustained volume makes operating a dedicated pipeline the more predictable choice. In those cases, send the same JSON schema to a self-hosted collector and store instead. Self-hosting adds capacity planning, upgrades, backups, and on-call ownership; hosted search gives up some infrastructure control in exchange for delegating those jobs. Neither option fixes a careless event schema.
Record the token counts returned for the completed call, then calculate cost from the rate-table version active for that event. Preserve raw counts even if you also emit estimated_cost. Recomputing later is then possible when a mapping was wrong, and changing commercial terms does not rewrite historical events silently. Failed calls may lack usage, so absence must remain distinct from zero.
Logs and metrics answer related questions at different granularity. Logs should carry run_id, step_id, and error details because engineers search individual runs. Metrics should aggregate request counts, token totals, estimated cost, and duration distributions using bounded dimensions such as service, environment, operation, status, and an internally controlled model key.
Prometheus instrumentation guidance warns against labels with high cardinality and specifically advises against unbounded values such as user IDs. The same reasoning applies to student IDs, run IDs, prompt hashes, Postgres query text, and exception messages. A unique label series per tutor run makes the metric system do the job of a log index. Keep those fields searchable in logs and link the views with time, service, operation, and status.
This split controls spend. Sampling routine successful debug events can reduce volume, but never sample the accounting event that carries token usage unless another authoritative ledger exists. Error sampling needs a deterministic key, such as a hash of run_id, so related events are kept or dropped together. Always retain aggregate counters outside the sampled stream. Otherwise the dashboard looks cheaper precisely when traffic grows.
Postgres deserves its own boundary event. Record operation: "session.update", elapsed time, affected-row count, and a normalized outcome. Do not log parameter values. For a worker, add queue wait and execution durations separately; for cron, add the schedule name, rows examined, rows changed, and completion status. One query can then distinguish model latency from queue pressure or database time without collecting a transcript.
Replay representative synthetic events from the API, worker, and cron services. A useful trial can start with 1,000 synthetic tutor runs split across success, model failure, queue retry, delayed work, and database timeout cases. Include multiline exceptions normalized into JSON and retries sharing a run_id; cap each test event at an application-defined byte limit and verify that oversize data produces a visible truncation field. Synthetic learner identifiers avoid putting real student data into an evaluation account. The number 1,000 is a test fixture, not a capacity claim: increase it until the event-size and concurrency distributions resemble the planned workload, then document that distribution so two candidates receive the same input.
Measure the workflow that matters: search by exact run_id; filter by service, status, and time; order steps; export the result; and confirm that retention deletion behaves as documented. Verify where log data, indexes, backups, and support access are processed when a Europe-region requirement is contractual. A region selector alone does not answer those questions. Record the answers in the architecture decision, along with the date checked.
For cost, feed the trial with the expected event-size distribution and daily volume, then map every charged dimension you can verify: ingestion, retained storage, search or scan volume, users, and outbound export. Do not reduce this to one advertised number. Apply a burst case too. A runaway stack trace or accidentally logged response body should hit an application-side byte cap and truncation marker before it can distort the bill.
The pass condition can stay compact: an engineer can reconstruct a run across all three process types, accounting events are complete, searches finish within the team's incident budget, data handling meets the European requirement, and export preserves structured fields. If any result depends on manual parsing or a proprietary event shape, the switching cost is already visible.
Treat logging as a production dependency, even though application work must continue during a logging outage. Standard output plus a runtime collector gives the process a simple contract. Bound buffers, define overflow behavior, and monitor dropped-event counts locally. Never block a learner response indefinitely because a remote log endpoint is slow.
Before deployment, test serialization failures, circular error objects, oversized fields, and abrupt worker termination. In staging, assert that one synthetic run produces the expected API, model, Postgres, queue, and cron events, with no learner content. After deployment, compare aggregate token totals from logs with the model billing source on a fixed cadence; investigate gaps by service and status rather than assuming the estimate is an invoice.
The operational checklist is deliberately prose because ownership matters more than boxes. The application team owns the event schema and redaction tests. The platform owner watches forwarding lag, dropped records, and retention. Whoever updates model routing also updates the versioned rate table and validates a known token fixture. During an incident, search the run first, then move to aggregate metrics to learn whether it is isolated. Quarterly, test an export and deletion request.
Start with a small schema. Preserve raw accounting inputs, keep identifiers out of metric labels, and make the hosted store prove search, residency, retention, and export with your event shapes. That produces defensible cost per tutor run without tying the agent loop to a logging product.
References: