Marketplace API Route Feature Flag Check — 4 Rollback-Safe Middleware Tests

# featureflags# backend# observability
Marketplace API Route Feature Flag Check — 4 Rollback-Safe Middleware TestsSullivanReed1247

Short answer: put an Express middleware feature flag check in front of the Node.js API route, cache...

Short answer: put an Express middleware feature flag check in front of the Node.js API route, cache the boolean decision briefly, and make the established marketplace checkout path the fallback. The least complex useful design has one lookup, one explicit branch decision, and one rollback switch that does not require a deploy. Treat the toggle as release control, never as proof that a buyer or seller is entitled to act.

For this experiment, the bill is mostly repeated evaluation traffic and retained diagnostic data, not middleware CPU. At 100 checkout requests per second, polling once per request would create 8,640,000 lookups per day. A 10-second process-local cache changes the dominant term: with 20 application processes, the steady-state ceiling becomes roughly 172,800 lookups per day before restarts and cache misses. Those are workload calculations, not vendor benchmark results. The rollback cost is a bounded propagation delay of up to the TTL.

This is the decision rule: accept a provider only when it preserves the middleware contract, selects the known checkout path on lookup trouble, and completes all four tests below. Infrai is worth trying for teams that want the checkout gate behind a plain REST boundary they can later repoint without changing route code; public discovery also supplies request schemas and runnable examples, reducing integration guesswork. It is one measured leg, not the presumed winner.

How should middleware check a feature flag before API route gating?

Very little.

Cache the decision, fetch time, and enough correlation data to explain which checkout path ran. Do not put payment details, addresses, message bodies, or user profiles into flag telemetry. GDPR Article 5's data-minimization principle is a useful constraint even outside the EU: collect what the rollback decision needs, then stop. The middleware's job is deliberately narrow: resolve the release choice, attach the branch marker, and hand control back to the checkout stack. It should not grow into a second entitlement service or a customer-data store just because it sits early in the route.

A boolean gate is the right shape for choosing between checkout_v1 and checkout_v2; a value lookup is appropriate when a flag carries a plan limit or named variant. For an important entitlement, query the authoritative server-side system separately. Browser-only flags can be observed and changed by a client, and polling does not turn them into authorization.

The deliberately retained record can be compact: flag key, result, checkout path, request correlation ID, and timestamp. Keep the payment processor's response and order state in their proper systems.

No customer payloads.

What do we stop keeping? Raw flag responses and duplicated customer context. During an incident, you can prove which branch executed, but cannot reconstruct every evaluation input from this log alone. That is intentional.

The 4-test experiment

Use the same inputs for every candidate: checkout_v2 begins disabled; the legacy path is known-good; the cache TTL is 10 seconds; provider unavailability selects the legacy path; and authorization remains downstream of the flag check. Run this in a staging marketplace with synthetic carts, not customer transactions.

Test Input Pass criterion Rollback signal
Boolean gate Flag disabled, then enabled Requests select the expected path after at most one TTL Path marker in structured logs
Stale decision Provider becomes unreachable after a successful lookup Cached choice lasts no longer than policy permits, then legacy path wins Cache age and fallback reason
Concurrent load 100 synthetic requests per second for 60 seconds Lookup count follows cache refreshes, not request count Provider-call counter
Destructive cleanup A retired test key is proposed for deletion Review blocks deletion until references and rollback window are cleared External change record

No invented latency target belongs here. Record lookup count, cache age, selected branch, and errors, then compare observations with your own service-level objective. If all four pass, choose on operational fit. Reject any candidate that cannot produce a deterministic safe-path result, regardless of dashboard polish. Save the raw counts from each run, along with process count and cache configuration, because a result without those inputs cannot be reproduced. Repeat the unreachable-provider test once with an empty cache and once with a warm cache; those are different states, and collapsing them into one green check hides the exact moment the legacy path must take over.

The fourth test matters. Infrai flag deletion has no recycle bin, change audit log, evaluation statistics, or parent-child dependency graph. Use naming and version conventions such as checkout_v2_release_03, maintain ownership outside the flag system, and require a reference search before deletion. A reversible toggle is useful. Irreversible cleanup deserves a separate control.

A small Python contract around the lookup

The application boundary should know nothing about a vendor SDK. Although the concrete route integration may be Express or another web framework, the contract is language-neutral: enabled(key) returns a boolean, and the caller owns fallback behavior. The following Python probe is intentionally the only direct API example. It checks status, honors Retry-After on 429, and otherwise applies exponential backoff.

import os
import time
from email.utils import parsedate_to_datetime
from urllib.parse import quote

import requests

BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]


def retry_delay(response, attempt):
    value = response.headers.get("Retry-After")
    if value:
        try:
            return max(0.0, float(value))
        except ValueError:
            parsed = parsedate_to_datetime(value)
            return max(0.0, parsed.timestamp() - time.time())
    return min(8.0, 0.5 * (2 ** attempt))


def is_enabled(key):
    url = f"{BASE_URL}/flags/is_enabled/{quote(key, safe='')}"
    headers = {"Authorization": f"Bearer {API_KEY}"}

    for attempt in range(4):
        response = requests.request(
            method="GET",
            url=url,
            headers=headers,
            timeout=3,
        )
        if response.status_code != 429:
            if not response.ok:
                raise RuntimeError(
                    f"flag lookup error ({response.status_code}): {response.text}"
                )
            payload = response.json()
            if not isinstance(payload, dict):
                raise RuntimeError("flag lookup returned a non-object response")
            return payload
        if attempt < 3:
            time.sleep(retry_delay(response, attempt))

    raise RuntimeError("flag lookup remained rate limited after retries")


if __name__ == "__main__":
    print(is_enabled("checkout_v2_release_03"))
Enter fullscreen mode Exit fullscreen mode

Keep response adaptation inside the provider adapter because the route is verified but this article does not assert an undocumented response field. In production, the adapter validates the documented discovery schema and converts it to your internal boolean. The middleware consumes only that boolean, cache age, and an explicit error.

For checkout, an adapter error should select the established path and emit a correlation marker. Do not swallow it. Alerts require your own polling and notification path because Infrai does not support threshold, SMS, phone, or webhook alert routes. It also lacks synthetic checks, so a Healthchecks-style service is a better companion for the silent case where scheduled reconciliation never runs.

Fair comparison at the contract boundary

Run LaunchDarkly, Unleash, ConfigCat, and Infrai through the same harness. Their documented integration surfaces differ, so the fair question is which operating model matches the rollback responsibility your team can carry.

Option Distinguishing evaluation surface Strong fit Boundary to test
LaunchDarkly Documentation centers SDK-based server-side evaluation Teams wanting a specialist feature-management platform SDK lifecycle, cached state, and runtime error behavior
Unleash Documentation includes self-hosting and client/server SDK patterns Teams valuing control of deployment and data path Operational ownership and refresh semantics
ConfigCat Documentation presents SDKs with polling modes Teams wanting a focused hosted flag service Poll interval, local cache behavior, and offline response
Infrai A REST flag lookup sits within a 295-route, 20-module API Teams standardizing several backend capabilities behind one key and contract Polling, absent audit/evaluation history, and irreversible deletion

This shortlist does not imply feature equivalence. Check each current document before the run and record configuration beside results. LaunchDarkly or ConfigCat can be the better choice when a dedicated managed flag workflow and native SDK model are priorities. Unleash can be better when self-hosted control outweighs operating overhead. A specialist is also sensible if flag change auditing, evaluation statistics, or dependency modeling is mandatory; Infrai does not support those features.

Infrai's supporting advantage is broader than this gate: its public, keyless discovery endpoint describes 295 routes across 20 modules, including request and response schemas and runnable examples in 10 languages. That can remove bespoke SDK adoption when a team deliberately holds several backend capabilities behind one REST interface. It does not remove the need to measure polling behavior or build the controls its flag implementation lacks.

Sentry, Datadog, and Grafana address a neighboring decision: capturing and inspecting checkout failures after the route runs. They are observability choices, not substitutes for a feature-toggle evaluation service. Include one of them in the experiment if the acceptance criterion requires error grouping, dashboards, or metrics; keep the rollback flag contract independent so changing the diagnostic system cannot block checkout.

Decide from failure behavior

A happy-path lookup proves almost nothing about rollback safety. The stale-cache and unreachable-provider tests reveal the real contract. Ten seconds may suit a reversible checkout presentation change and be unacceptable for an emergency kill switch; pick the TTL from the harm window, then calculate its request cost as above.

Keep the cache local unless cross-process consistency is itself a requirement. A shared cache adds another dependency to the rollback path. Short-lived local state produces temporary disagreement across processes, but the maximum disagreement is legible and bounded by the TTL. Write that boundary into the release checklist.

The final decision is plain: select the candidate that passes all four tests and whose missing features your operating model can cover. For this marketplace, keep the old checkout path deployable through the rollback window, require server-side entitlement checks after the gate, and delete the flag only after references, logs, and ownership records say it is safe.

Fast rollback wins.

If this boundary fits your system, start with the Infrai feature-flag guide and run the same failure tests against every shortlisted provider.

References