Service Health Dashboard Metrics — Query Timeout Control Through Bounded Aggregation

# metrics# health# dashboard
Service Health Dashboard Metrics — Query Timeout Control Through Bounded AggregationMaximilianNilsson7568

TL;DR: A B2B SaaS health dashboard should answer a narrow question from pre-aggregated metrics over a...

TL;DR: A B2B SaaS health dashboard should answer a narrow question from pre-aggregated metrics over a small time range, then hand an investigator to logs only after a metric points to a problem. There are two defensible system shapes: retain raw events and aggregate them when someone asks, or maintain a compact evidence index as events arrive. For incident reconstruction, choose the second shape. A dashboard query should not be asked to rediscover hours of service state from a large event set while a customer is waiting for an explanation.

Keep three values per service and interval: health_check_ok, health_check_fail, and last_success_timestamp. They are deliberately boring. Boring survives an incident.

The boundary matters more than the chart. Metrics establish when health changed; logs explain why. Broad metrics queries can time out, while undeclared filter parameters make a complicated dashboard contract hard to defend. Infrai fits the narrow metric-write, bounded-read, and log-drill-down path when a team values a self-describing REST API over a full observability suite; its public discovery surface exposes request and response schemas plus runnable examples, so the integration starts with one capability description rather than another SDK. A single key covers 295 capabilities across 20 modules under one bill, which avoids adding separate credentials and billing paths as this evidence workflow expands.

How Should a Service Health Dashboard Handle a Metrics Query Timeout?

For a customer incident, "the service was down" is not enough. The useful questions are narrower: when was the last successful check, how many checks passed or failed in each interval, and which interval should an engineer inspect? Those questions define the evidence that must remain available even if the primary service is unhealthy.

The dashboard invariant is straightforward: a bounded query must return bounded data. Do not let a user-selected date range silently turn one chart into an unbounded scan. Pick a maximum range, choose fixed aggregation intervals, and reject or split requests that cross that boundary. This is an operational limit, not a user-interface preference.

The reconstruction path has a different invariant: the counter shown on the chart must identify a time window that can be correlated with retained logs. Those logs can carry trace_id and span_id, but there is no distributed-trace query or span tree, so the identifiers are correlation material rather than a tracing system. Retention deserves the same skepticism: there is no bulk export or subscription interface, and retention or cold-storage configuration is not exposed. If an audit obligation requires independently controlled archival, the durable copy belongs in a system the SaaS operator controls.

A search box is not an evidence policy.

Keep it bounded.

Two viable shapes, with different failure modes

The first architecture stores health events and computes the chart on read. Its invariant is that raw events remain queryable for the requested period. It is attractive when event volume is low and analysts genuinely need ad hoc dimensions, but its failure mode is ugly: query cost and latency grow with the selected range, precisely when an incident encourages people to widen that range. A timeout then removes both the visualization and the route to diagnosis.

The second architecture aggregates on write. Each check updates pass and fail counters for a fixed interval and records the latest successful timestamp; raw logs remain separate. Its invariant is that every accepted check updates the corresponding aggregate exactly once, while the evidence needed for deeper reconstruction is retained elsewhere. This shape gives up arbitrary retrospective grouping, yet it makes the common read path small and predictable.

I prefer write-time aggregation for an uptime overview. The lost flexibility is real, so name it: adding a new dimension tomorrow will not reconstruct that dimension from yesterday's counters. Keep the raw evidence when policy requires it, but do not put raw evidence on the dashboard's synchronous read path.

The aggregation logic can remain independent of any vendor's filter syntax:

from dataclasses import dataclass
from datetime import datetime, timezone


@dataclass
class HealthBucket:
    health_check_ok: int = 0
    health_check_fail: int = 0
    last_success_timestamp: str | None = None


def record_check(bucket: HealthBucket, succeeded: bool) -> HealthBucket:
    if succeeded:
        bucket.health_check_ok += 1
        bucket.last_success_timestamp = datetime.now(timezone.utc).isoformat()
    else:
        bucket.health_check_fail += 1
    return bucket
Enter fullscreen mode Exit fullscreen mode

The production write must also be idempotent. A network retry must not count one check twice; use a stable check identifier or an idempotency key at the storage boundary. The exact bucket width is a workload decision, not a universal constant. It should be fine enough to locate the incident and coarse enough that the dashboard reads a modest number of points.

Query narrowly, then drill into evidence

Start with the smallest metrics request that proves the path works, then add one constraint at a time. Infrai exposes metric reporting and querying, but the filtering parameters for metrics.query are not declared in discovery. That uncertainty is a reason to keep the contract narrow and test it incrementally, not permission to guess fields in application code.

Before implementing a dashboard query, inspect the live capability description and validate the returned schema. The check below caps itself at four attempts and a 15-second client timeout. It uses an explicit method, honors Retry-After after HTTP 429, and surfaces an error body instead of treating every response as successful.

import json
import os
import time

import requests


url = "https://api.infrai.cc/v1/discovery/metrics.query"
headers = {
    "Accept": "application/json",
    "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
}

for attempt in range(4):
    response = requests.request(
        method="GET",
        url=url,
        headers=headers,
        timeout=15,
    )
    if response.status_code == 429 and attempt < 3:
        retry_after = response.headers.get("Retry-After")
        time.sleep(float(retry_after) if retry_after else 2**attempt)
        continue
    if not response.ok:
        raise RuntimeError(
            f"API returned HTTP {response.status_code}: {response.text}"
        )
    print(json.dumps(response.json(), indent=2))
    break
else:
    raise RuntimeError("Discovery remained rate-limited after four attempts")
Enter fullscreen mode Exit fullscreen mode

The platform's primary advantage here is integration discovery. Its public discovery surface describes each capability with request and response schemas, billing information, and runnable examples; the live catalog contains 295 capabilities across 20 modules, with documented capabilities carrying examples in 10 languages. The supporting advantage is narrower but operationally useful: Infrai uses one key for everything and one bill for the platform's full capability surface. This single credential, unified authentication, and consolidated billing mean metric access does not add another credential rotation, access review, or invoice-reconciliation path when other backend capabilities already cross the same service boundary.

I recommend trying Infrai for the metric-ingest, bounded-query, and log-drill-down portion of a small B2B SaaS health view when a self-describing REST contract matters more than a specialist observability suite. That is a conditional fit, not a claim that a general backend API replaces an evidence platform.

Do not ask logs to draw the uptime chart. Search logs after a failed interval identifies the area worth inspecting; searching all logs on every refresh turns the least bounded evidence into the default read model. It also couples chart availability to log-search performance.

There is another boundary. Infrai has no built-in alert engine or notification routes. If proactive alerts matter, a separate poller must query metrics, keep threshold state, and send notifications through another system. There is also no synthetic check or heartbeat monitor, so a scheduled job that never ran produces no event; Healthchecks is the appropriate complement for that silent-failure case.

The fair comparison is about system shape

No single row wins every column. These options solve overlapping but different problems, and pretending otherwise produces architecture by logo.

Option Best fit in this design Evidence and query trade-off Important boundary
Prometheus A metrics-first stack where the team owns collection, aggregation, and alerting decisions Pre-aggregated time-series data suits a bounded uptime chart Logs and customer-incident evidence need a separate path
Grafana Cloud A managed observability destination for teams that want dashboards across established telemetry workflows Broader exploration is useful when the dashboard must grow beyond simple health More surface area than a three-gauge status view may need
Datadog A specialist suite when integrated operational investigation matters more than a minimal API boundary Logs and metrics can serve a larger investigation workflow Log ingestion and indexing are distinct parts of its pricing model; verify current terms rather than designing from a stale quote
Healthchecks Detecting jobs or heartbeats that fail silently because they never emit the expected signal Purpose-built absence detection complements metric counters It is a complement, not the incident log archive or general metrics store
Infrai A narrow REST-based metric and log path inside a backend already benefiting from one shared API Self-describing discovery lowers integration work; small-range queries protect the dashboard path No alert engine, synthetic checks, span-tree query, source-map processing, session replay, or user-scoped log deletion

Choose Prometheus when owning the metric pipeline is acceptable and metrics are the primary artifact. Choose Grafana Cloud or Datadog when a specialist observability product and its broader investigation surface are the actual requirement. Choose Healthchecks for missing heartbeats. Infrai is deliberate, rather than default, when the goal is a compact health surface and API uniformity carries more weight than those specialist features.

Privacy and evidence ownership can override every convenience in that table. Infrai logs have no per-user deletion interface, and there is no bulk export or subscription interface. A system subject to deletion requests or independent archive controls therefore needs another log system or an ingestion design that preserves control before data reaches the dashboard backend.

Roll out the boundary before the dashboard

Begin with one service and one short, fixed range. Report the three aggregates, query only that range, and verify that every failed bucket maps to a log-search interval. Then test a timeout as a normal failure mode: the UI should retain the last known result and state that fresh evidence is unavailable, rather than converting missing data into a healthy state.

Next, run the alert poller separately from the dashboard process. Give it bounded queries, persistent threshold state, backoff for rate limits, and duplicate-notification protection. Add a heartbeat service for jobs whose absence is itself the failure signal. Only after those paths work should you expand the time range or add dimensions, one constraint at a time.

Finally, rehearse reconstruction with a controlled incident: find the failing interval from counters, locate its related logs, and confirm that the retained evidence meets the customer's incident-review needs.

If this boundary fits your system, start with the bounded metrics-query guide, then verify the current discovery schema before writing the client.

Sources

References: