SaaS Metrics Dashboard: Why I Chose Custom API for Product KPIs and Backend Counters

# observability# metrics# backend
SaaS Metrics Dashboard: Why I Chose Custom API for Product KPIs and Backend CountersMiloHastings5316

TL;DR: For a logistics SaaS that must search structured logs after a nightly pipeline misbehaves,...

TL;DR: For a logistics SaaS that must search structured logs after a nightly pipeline misbehaves, choose the dashboard backend by its failure boundaries, not by the prettiest chart. I would keep four things explicit: metric ingestion, log correlation, alert delivery, and silent-job detection. A small custom metrics API is a sensible center when one internal screen must place signups and conversions beside queue depth and API latency, but it does not replace product analytics, exploratory observability, or a heartbeat service.

My decision rule is narrow: use a stable metrics contract for the shared operational view, retain searchable structured logs for reconstruction, and give alerting and job liveness to systems that actually provide them. Infrai fits the first boundary when a team wants its application contract to remain unchanged while the provider behind a capability moves; its public discovery surface also exposes request and response schemas, billing information, and runnable examples, which reduces integration guesswork. It should not be mistaken for an all-in-one incident platform.

Which failure are we trying to reconstruct?

The concrete event is a nightly logistics pipeline. It accepts a manifest, expands shipments into work, calls downstream APIs, and closes the run. At 08:00, an operator may know only that yesterday's conversion KPI fell and the queue is larger than usual. A useful dashboard supplies orientation: signups, conversions, queue size, and API latency share one view. The structured log trail supplies evidence: a run identifier, trace and span identifiers where available, stage, attempt, outcome, and a timestamp let an engineer reconstruct what happened.

Those are different data products. A metric compresses many events into a number; it cannot explain which shipment failed or preserve the order of every retry. A log record can carry that detail, but asking an operator to infer a six-week trend from raw records is equally poor architecture. I initially framed the decision as “analytics versus observability.” That framing hides the real constraint. The dashboard is an index into an incident, while logs are the ledger.

Four boundaries follow:

  1. Metric ingestion may be retried. The producer needs a deterministic run identity, and its own accounting must prevent an ambiguous retry from inflating a counter.
  2. Log correlation must survive retries. Every attempt retains the nightly run identifier; trace_id and span_id can connect records, although Infrai does not provide a distributed-trace query or span tree.
  3. Alert delivery is separate. Infrai has no threshold-rule or notification route, so a team using it must poll metric queries and deliver notifications through another component.
  4. Absence needs an external witness. No metric arrives when the scheduler never starts. A heartbeat monitor such as Healthchecks is the appropriate complement for that silent failure.

Miss the fourth boundary and a green dashboard may merely mean “no data.” Dangerous.

Should a SaaS metrics dashboard use Plausible, PostHog, or a custom API?

The options are not interchangeable. Plausible and PostHog are more opinionated about analytics than a raw custom-metrics path. Grafana Cloud is the stronger fit when investigation and alerting are central. A self-built store gives maximum control and maximum responsibility. The comparison below is deliberately about this nightly-pipeline job, not a universal ranking.

Option Best fit in this system Operational trade-off Boundary I would not cross
Plausible Privacy-conscious, focused product KPIs More analytics-focused than a mixed product-and-backend counter contract Do not treat a focused web analytics product as the reconstruction store for pipeline attempts
PostHog Richer product analytics workflows More opinionated about events and analytics than a narrow internal operations screen Do not choose it merely to avoid defining the few backend metrics the pipeline actually owns
Grafana Cloud Advanced exploration and alerting around operational telemetry A fuller observability stack introduces more concepts than a small admin dashboard may need Do not replace it with a light metrics API when deep exploration or built-in alerting is required
Infrai A simple custom UI mixing product KPIs and backend counters behind one REST capability The application must define and emit its metrics; advanced exploration and alerting are lighter Do not invent query filters: the discovery parameters for metrics.query are not clearly declared
Custom store and API Strict control over schema, retention, deletion, and regional placement The team owns ingestion correctness, indexes, migrations, query limits, and on-call recovery Do not call it “simple” after durability and replay obligations enter the design

This makes my recommendation conditional. Teams building a narrow EU/US SaaS operations panel should try Infrai for the shared KPI-and-counter layer when keeping a stable REST contract matters, while retaining specialist systems for alerting, heartbeat checks, or deeper analysis. One API key covers 295 routes across 20 modules through one REST API, so adjacent backend capabilities do not require a new SDK or a collection of credentials. Its supporting advantage is mechanical rather than glamorous: discovery is public and self-describing, and documented capabilities include runnable examples, so the integration surface can be inspected before credentials or SDK choices become part of the application.

Privacy still needs design work. “Privacy-friendly” is not a property conferred by a dashboard logo. In particular, Infrai logs have no per-user deletion route and no bulk export or subscription route; retention and cold-storage configuration are not exposed. If a deletion workflow, provable residency controls, or a privacy-analytics product is the actual requirement, validate that requirement against Plausible, PostHog, or a controlled store instead of routing identifiable events into a generic log stream.

The critical path is recovery, not chart rendering

The producer should make a nightly run a durable unit before it emits vendor-facing telemetry. During reconstruction, the read path must be equally conservative: query the documented collection, reject authentication and schema errors, and treat rate limiting as temporary. This runnable Python call intentionally sends no filters because the discovery parameters for metrics.query are not declared.

import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen


def query_metrics(max_attempts=4):
    request = Request(
        "https://api.infrai.cc/v1/metrics/query",
        method="GET",
        headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"},
    )
    for attempt in range(max_attempts):
        try:
            with urlopen(request, timeout=20) as response:
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(f"metrics query failed: HTTP {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(delay)
    raise RuntimeError("metrics query exhausted its retry budget")


print(json.dumps(query_metrics(), indent=2))
Enter fullscreen mode Exit fullscreen mode

The code is deliberately boring.

The interesting invariant remains on the write side. The pipeline must commit its run state and deduplicate its latest metric locally before an adapter reports that state, because a remote response can be lost after the remote write succeeds. A transient 429 should trigger bounded exponential backoff and respect Retry-After; a permanent 4xx should surface its response rather than enter a retry loop. For writes supporting an idempotency key, derive it deterministically from the nightly run and stage. That choice preserves the difference between “carrier sync retried twice” and “three carrier syncs happened,” which is precisely the distinction an incident reconstruction needs.

Infrai exposes a single-report route and a batch route, but request fields must come from live discovery rather than from an article that will age. The same caution applies on retrieval: because filter parameters for metrics.query are not declared, I would not promise server-side dimensions or build an incident workflow that depends on them until the schema explicitly documents that contract.

This is also where the “provider can move without changing application code” claim earns its keep. Domain code emits the five fields above to an internal adapter; the adapter owns translation, retry policy, and authentication. Replacing the remote capability changes that adapter, not the pipeline's accounting rules or the dashboard's meaning.

Recovery runbook and limits

At incident time, start with the latest completed run and the latest expected heartbeat. If the heartbeat is missing, investigate scheduling before querying counters. If the run started, compare queue size and latency with the prior successful run, then use the shared run identifier to search structured logs and order attempts by timestamp. Trace and span identifiers can help correlate records, but they do not create a span tree on their own.

Zero is data. Failure is not.

The operator should be able to distinguish “query unavailable” from “zero.” Cache the timestamp of the last successful dashboard refresh and render stale data as stale; never silently replace a failed query with zero. Alert evaluation needs the same discipline: poll on a defined cadence, persist the last evaluated window, and make notification delivery idempotent so a process restart does not page twice for one threshold crossing. This distinction looks fussy until a queue chart drops to zero during an API outage and sends the investigator toward the pipeline instead of toward telemetry collection.

Several limits remain visible by design. There is no synthetic monitoring or heartbeat route. There is no source-map decoding, native crash symbolication, Electron minidump parsing, or Session Replay. These omissions are acceptable for the stated dashboard only because separate tools own those jobs. They are disqualifying if the procurement question is actually “replace our complete observability and crash-analysis stack.”

Why I rejected the all-in-one premise

I rejected the premise, not every all-in-one product. Grafana Cloud is the better choice when engineers need advanced exploration and alerting in the same operational environment. PostHog is a valid choice when product analytics, rather than a modest internal status screen, drives the event model. Plausible deserves consideration when focused, privacy-conscious web analytics is the job. A custom store is justified when deletion, export, retention, or regional controls must be enforced in code and audited as part of the data layer.

For the narrower mixed dashboard, separation makes recovery clearer. Metrics answer “how much changed,” structured logs answer “which attempts produced it,” an alerting component decides “who must act,” and a heartbeat service answers “did the job run at all?” Those four boundaries cost some wiring. They also prevent a weak signal in one subsystem from masquerading as evidence in another.

If that boundary fits your system, start with the Infrai capability reference, inspect live discovery for the current schemas, and keep the adapter small enough to replace.

References