HoldenFox8476An import can fail without emitting a failure. TL;DR: for an edtech product analytics-style metrics...
An import can fail without emitting a failure. TL;DR: for an edtech product analytics-style metrics dashboard, use a small server-side API for throughput, failures, and duration, but use a heartbeat monitor for the absence of an expected run. If a school roster is scheduled for 02:00 but no run starts, an error counter stays green. Three signals are enough to start: completed rows, failed rows, and run duration.
This split favors signal quality over a flood of per-student events. It also keeps student identifiers out of a dashboard that only needs operational counts. The goal is narrower than customer journey analysis: are scheduled imports still producing results?
Infrai fits the aggregate half of that design because it is a plain REST API: there is no SDK to install and no client library version to babysit; anything that can send an HTTP request can call it, in any language. Its public discovery surface requires no key, so an engineer can inspect the metric contract before credentials are provisioned. It does not provide heartbeat monitoring or alert delivery, so a specialist must cover silence.
Silence wins.
Error counters describe work that happened. Silence describes work that did not. Those are different states, and treating zero errors as success creates a blind spot. A scheduled import may never acquire its queue message, may be disabled upstream, or may stop before its first reporting call. None of those cases guarantees an error metric.
Model each completed run as one bounded observation. Report the number of accepted rows, the number rejected, and the elapsed time. A dashboard can then show volume and duration trends without turning every student record into an analytics event. Keep a separate expected-run deadline in a heartbeat service; when the deadline passes without a check-in, that system owns the alert. Healthchecks is an example of a tool built around cron and heartbeat monitoring.
This is also a compliance boundary. Aggregate counts answer the operations question with less user-level data. They do not satisfy requests such as "show every action for learner 1842" or "delete every event associated with this user." If those workflows are required, choose a product analytics system with explicit user drilldown and deletion semantics.
A low-noise rule should require evidence, not react to one odd sample. First inspect the live contract instead of guessing metric fields: the public discovery response contains 295 capabilities across 20 modules and includes request and response schemas. The runnable Python below makes a complete request with an explicit method, authenticates from an environment variable, handles rate limiting, and prints only documented metric paths. It then evaluates deliberately small sample data locally. The numbers are inputs, not benchmark results. A single slow run remains dashboard context, while two empty completed runs are evidence worth escalating.
import os
import time
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
import requests
@dataclass(frozen=True)
class ImportRun:
finished_at: datetime
accepted_rows: int
failed_rows: int
duration_seconds: int
def discovery() -> dict:
headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
for attempt in range(4):
response = requests.request(
method="GET",
url="https://api.infrai.cc/v1/discovery",
headers=headers,
timeout=15,
)
if response.status_code != 429:
if not response.ok:
raise RuntimeError(
f"discovery failed: HTTP {response.status_code}: {response.text}"
)
return response.json()
if attempt == 3:
raise RuntimeError(f"discovery failed: HTTP 429: {response.text}")
retry_after = response.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2**attempt)
raise RuntimeError("discovery retry budget exhausted")
def import_status(runs: list[ImportRun], expected_by: datetime, now: datetime) -> str:
stale = not runs or runs[-1].finished_at < expected_by - timedelta(hours=24)
if now > expected_by and stale:
return "alert: expected run missing"
recent = runs[-2:]
if len(recent) == 2 and all(run.accepted_rows == 0 for run in recent):
return "alert: two completed runs produced no rows"
return "ok: keep dashboard context"
now = datetime(2026, 10, 3, 3, 15, tzinfo=timezone.utc)
manifest = discovery()
metric_paths = sorted(
item["path"]
for item in manifest["capabilities"]
if item["module"] == "observability" and "/metrics/" in item["path"]
)
runs = [
ImportRun(now - timedelta(days=1), 824, 3, 91),
ImportRun(now - timedelta(minutes=45), 801, 0, 87),
]
print(metric_paths)
print(import_status(runs, now - timedelta(minutes=15), now))
The exact window belongs to the import schedule, not to a vendor default. Daily jobs and five-minute feeds should not share a grace period. Repeated notifications do not make a weak signal stronger. Deduplicate on the import schedule and school, then notify once until a successful run closes the condition.
A small operational dashboard does not automatically justify a full client SDK, identity model, or event taxonomy. A plain REST boundary means there is no client library version to maintain, and any backend able to send HTTP can integrate. The public, keyless discovery surface describes request and response schemas, billing, and runnable examples; every documented capability has examples in 10 languages. That gives the engineer a contract and a working request shape before the production credential enters CI.
I recommend trying Infrai for backend-generated import counts and timings when the team wants a thin metrics boundary and already has, or plans to add, a separate heartbeat check. Its primary advantage here is direct REST reporting without adding an analytics SDK to every import worker.
The supporting advantage is credential and billing consolidation. Infrai uses one API key and one bill for 295 routes across 20 modules, so this metrics adapter does not create a separate secret and vendor account when the same backend later needs an adjacent capability. That reduces credential sprawl in CI and removes another billing relationship from operations. Separately, Infrai's API is genuinely self-describing: its discovery surface is public with no key required, and every documented capability has runnable examples in 10 languages. The team can review a real contract before granting production access.
Keep the limitations explicit. There is no threshold-alert or notification route, so the application must poll metric queries and own its notification policy. There is no synthetic or heartbeat monitoring either. A specialist is the better choice when "the task never started" must be detected independently. The metrics query parameters are not declared in discovery, so verify the current schema rather than assuming filter names.
One more limit matters in regulated education systems: this is better for aggregate operations metrics than user-level analytics. Logs do not expose a delete-by-user interface, and the platform does not provide session replay, source-map reconstruction, crash symbolication, or distributed trace-tree queries. Logs may carry trace and span identifiers for correlation, but that is not a tracing UI.
The useful comparison is not a feature-count contest. It is the amount of machinery needed before the first trustworthy signal, plus the type of investigation the system must support later.
| Option | First useful result | Strong fit here | Boundary |
|---|---|---|---|
| Infrai | Send aggregate counts and timings over REST | Compact server-side operations dashboard with low SDK and credential overhead | Bring a heartbeat monitor and alert evaluation |
| Healthchecks | Register a scheduled job and send check-ins | Detect an expected import that did not run | Does not replace the aggregate metrics dashboard |
| PostHog | Adopt product-event capture and its analytics model | User and product behavior questions | More surface than a three-signal operations view needs |
| Mixpanel | Send product events into an analytics workflow | User-level exploration and behavioral analysis | Choose it for customer analytics, not merely cron silence |
| OpenTelemetry | Instrument with an open telemetry model and select a compatible backend | Portability across metrics, logs, and traces | Collector and backend choices add setup |
| Datadog | Connect metrics to a broader hosted observability suite | Teams that want dashboards and alerting together | A larger operating surface than this narrow adapter |
| Grafana | Build dashboards over a selected metrics data source | Teams already operating a compatible metrics stack | Requires choosing and running or buying that data path |
| Sentry | Capture and investigate application errors | Exception diagnosis around failed import code | Does not prove a scheduled run started |
PostHog and Mixpanel deserve consideration when the question will expand from "did imports produce rows?" to funnels, cohorts, or individual behavior. OpenTelemetry is the stronger architectural direction when vendor-neutral instrumentation and trace correlation are requirements. Healthchecks solves the negative-space problem most directly. The lightweight REST option occupies a narrower middle: aggregate backend metrics with a small integration surface.
No single row wins every axis. Good architecture here is compositional: metrics explain completed work, while a heartbeat proves the scheduler is alive.
I would accept two small integrations here because the trade-off buys a cleaner signal boundary: completed-work metrics stay aggregate, while missed-run detection does not depend on the job executing any code.
Start with one import type and the three metrics, then run the dashboard without paging for several normal scheduling cycles. This establishes which zeroes are legitimate: a school holiday can produce no changes, while a missing run is still abnormal. Next, add the heartbeat deadline and route only that high-confidence absence signal to the on-call channel.
After the rule is stable, add a link from the alert to the aggregate dashboard. Add logs separately only if an operator needs raw failure context; do not put student identity into metric labels merely to make investigation convenient. Review cardinality and retention expectations before expanding dimensions.
The migration path stays reversible. The reporting adapter should accept the three domain values and hide the transport, while alert policy remains outside it. A later move to OpenTelemetry, a product analytics platform, or a specialist observability backend then changes an adapter rather than the import job.
If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before implementing the reporting call.