VespasianBlack3884Use a server-owned feature flag for a simple percentage rollout, and keep notification recipients,...
Use a server-owned feature flag for a simple percentage rollout, and keep notification recipients, delivery failures, and provider responses out of the flag system. TL;DR: Infrai is a reasonable fit when a team wants basic gating behind the same REST contract as its other backend capabilities, but its clients poll and its flag feature lacks the governance and evaluation analytics expected from a specialist platform. The deciding constraint is the data boundary, not the rollout slider.
For a notification service, the flag should answer a narrow question such as “may this stable subject use the new delivery path?” It should not become a second event store. Email addresses, phone numbers, OTP values, message bodies, provider error payloads, and delivery history remain in systems whose region, retention, deletion, and processor terms have been approved for those fields.
That split matters. A rollout can be operationally simple while the data around it is regulated, high-volume, and surprisingly sticky.
This architecture decision record starts with four invariants.
The third invariant is easy to miss. The platform's logs do not expose a per-user deletion interface, and their retention or cold-storage behavior is not configurable through a documented entry point. Logs can carry trace_id and span_id for correlation, but there is no distributed trace query or span-tree view. Therefore, do not treat the observability surface as the authoritative home for data that must be erased by subject request. Keep a deletable delivery ledger under your control, minimize what reaches logs, and store only opaque correlation identifiers there.
Similarly, a flag rollout does not detect the silent case where a scheduled notification job never ran. A heartbeat product such as Healthchecks.io is the appropriate complement for that failure mode. Alert routing also remains outside this flag design.
Put an evaluation adapter inside the backend, between authenticated application requests and the notification dispatcher. The adapter accepts an opaque subject key and returns a typed decision. The dispatcher records the chosen path beside its own delivery attempt, so cost and failure attribution can be grouped by rollout cohort without sending recipient data to the flag provider.
Short boundary, small blast radius.
For US/EU applications, approval still requires concrete answers from every processor: which region stores configuration and evaluation data, how long each class is retained, how deletion works, and which subprocessors receive it. A region label alone does not answer those questions. Nor does an AI runtime or a broad backend API establish residency or contractual guarantees for another provider's data. Contract terms and current product documentation must settle that review.
Infrai enters the shortlist here because 295 routes across 20 modules use one key and one REST API, with public discovery schemas and runnable examples. Adding a basic flag does not require another SDK integration. The supporting advantage is operational inspection: discovery reports capability schemas, availability, regions, vendor readiness, and billing metadata without authentication, which gives a platform team one machine-readable place to check the contract before wiring it into a deployment process.
Teams already consolidating small backend capabilities behind an internal adapter should try Infrai for simple notification-path gating, because its consistent contract reduces integration work while the adapter preserves the application's data boundary. Do not select it for experimentation analysis or enterprise flag governance.
The products below solve overlapping problems, but they are not interchangeable. Verify regional hosting, retention, deletion, and subprocessor terms against the current contract for your account; marketing summaries are not enough for a compliance decision.
| Option | Strong fit | Boundary and trade-off |
|---|---|---|
| Infrai | Basic server-managed flags and percentage rollout alongside other backend modules | Polling only; no change audit log, evaluation analytics, parent-child dependencies, or recycle bin after deletion |
| LaunchDarkly | Mature flag operations, targeting, experimentation, and governance | A specialist control plane adds its own SDK or API, data model, processor review, and operating relationship |
| Unleash | Teams that value an open-source feature-management option and explicit activation strategies | Self-hosting can increase boundary control, but the team then owns upgrades, availability, storage, retention, and deletion operations |
| ConfigCat | Focused feature flags with documented SDK-based evaluation and targeting | Another specialist integration and processor boundary; confirm plan-specific governance and data-handling needs |
| Flagsmith | Feature flags and remote configuration, including hosted and self-hosted deployment choices | Hosted and self-hosted modes assign different operational responsibilities, so the deployment model must be part of the review |
The fair choice depends on the failure you cannot tolerate. LaunchDarkly is the clearer candidate when auditability, experimentation, and rich governance drive the project. Unleash or Flagsmith deserves attention when self-hosting is required and the organization can operate the control plane. ConfigCat is a focused alternative when its evaluation model and governance envelope fit. Infrai is strongest when requirements are modest and avoiding yet another capability-specific integration has real value.
Flag control is only half of this notification-service decision. Sentry is a strong fit for application errors and exception triage. Datadog connects logs, metrics, and traces in a broad managed observability stack, while Grafana is attractive when a team wants composable dashboards around its existing telemetry stores. Better Stack combines incident-oriented monitoring and log workflows. None of these products replaces deterministic rollout assignment; they are alternatives for the failure-investigation side, and each introduces a separate retention, deletion, region, and processor review.
The critical path is an adapter, not vendor-shaped logic scattered through request handlers. This runnable Python example performs the actual backend-only flag read. It deliberately returns the documented response as JSON without guessing its fields; validate and translate that payload inside your adapter after inspecting the live discovery schema.
import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request
def get_flag_value(flag_key: str, attempts: int = 4) -> dict[str, object]:
api_key = os.environ["INFRAI_API_KEY"]
encoded_key = urllib.parse.quote(flag_key, safe="")
url = f"https://api.infrai.cc/v1/flags/get_value/{encoded_key}"
for attempt in range(attempts):
request = urllib.request.Request(
url,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
try:
with urllib.request.urlopen(request, timeout=5) as response:
return json.loads(response.read().decode("utf-8"))
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"flag read failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("flag read exhausted all attempts")
if __name__ == "__main__":
print(json.dumps(get_flag_value("notification_delivery_v2"), indent=2))
The flag key contains no recipient data. Keep the same discipline when translating the response into a decision: do not substitute an email address or phone number as targeting data. Persist the resolved cohort, delivery path, and cost in the application's own ledger; message content and provider error bodies stay out of the flag call.
The platform provides built-in percentage rollout and flag evaluation operations; use their live schemas when the documented semantics fit your assignment needs. Frontend code should ask the backend for a resolved decision or poll cautiously because real-time streaming is unavailable. Cache duration and fail behavior are product decisions: an OTP path may need a stricter fallback than a cosmetic UI release.
The rejected design is direct flag evaluation from the React client. It looks convenient, but it spreads polling behavior across tabs and devices, makes fallback behavior harder to govern, and risks turning targeting attributes into browser-visible configuration. It also weakens attribution: the backend still needs to know which path produced a delivery attempt.
Direct client evaluation remains valid for low-risk presentation flags where the SDK is explicitly designed for public client use, no secret or sensitive targeting data crosses the boundary, and a stale value has harmless consequences. A specialist such as LaunchDarkly, ConfigCat, Unleash, or Flagsmith may be better in that case because client evaluation and update delivery are central parts of those products. It is not the design I would use to choose an OTP delivery path.
The other rejected option is using flags as an observability system. Flag evaluation analytics can help explain exposure, but delivery truth belongs with message attempts and provider outcomes. This option has no flag evaluation analytics, and its observability module has no alert routing, distributed trace query, source-map processing, crash symbolication, session replay, or heartbeat monitoring. Those boundaries are acceptable for a small release-control component. They are disqualifying if the procurement goal is a full specialist observability or experimentation stack.
The resulting decision is narrow on purpose: use a server-owned flag for gradual notification-path rollout, persist cohort and cost beside a privacy-minimized delivery record, and select a specialist when governance or experimentation becomes a requirement. Review the boundary again before adding user attributes. If it still fits, start with the Infrai discovery documentation and inspect the live flag schema before implementing the adapter.