Silhouette72591483Short answer: choose a realtime API surface only after assigning separate latency SLOs to connection,...
Short answer: choose a realtime API surface only after assigning separate latency SLOs to connection, authorization, subscription recovery, and quote delivery; for a stock trading watchlist, a fast happy-path event is not evidence that the watchlist stays accurate after a reconnect.
The decision is to treat presence and market-data freshness as different signals. A user can be online while a symbol subscription is stale, and a healthy transport can still be carrying an old business event. Recovery therefore belongs in the design, not in a footnote.
Start with the user-visible invariant: after the backend accepts a newer quote for a subscribed symbol, an authorized, connected watchlist must either display that version within its latency budget or display an explicit stale state. Do not silently preserve the last price and call the screen healthy. The actual threshold is a product and market-data decision; no measured runtime latency is available here, so I'm not sure a universal millisecond target would survive contact with the feed, region, and device mix. Measure those paths before committing to a number.
One SLO is too blunt. Define at least four histograms: time to authenticate, time to restore the subscription set, event age at receipt, and time from receipt to render. Give each observation a correlation ID plus connection generation, symbol, event version, and authorization result. This split makes a useful distinction: a transport reconnect may finish quickly while replay or resubscription remains incomplete. It also prevents a common diagnostic mistake — blaming the realtime provider for latency introduced before publish or after receipt.
Keep the clock boundaries honest. Server timestamps can establish business-event age only when clock skew is bounded; client monotonic clocks are better for local durations. Report percentiles over a stated window, and track the denominator: dropping expired or unauthorized attempts from the sample makes the chart prettier while making the SLO meaningless.
No hidden freshness.
The server owns authentication, authorization for symbols or watchlists, subscription truth, event ordering metadata, and any bounded replay policy. The client owns the active connection generation, the last accepted version per symbol, deduplication, stale-state rendering, and exponential reconnect backoff. Both sides need a shared definition of when recovery is complete. "Socket open" isn't that definition; recovery completes only after authorization is current and the intended subscription set is confirmed.
Reconnects, credential expiry, duplicate delivery, and partial subscription recovery are normal states. Model them. On reconnect, increment the connection generation, obtain current authorization, restore the desired symbols, reject events tied to an older generation, and compare event versions before updating the UI. If only 18 of 20 symbols are restored, show those two as stale rather than treating the watchlist as wholly live. The numbers here illustrate state accounting, not a measured service result.
Test the ugly path with controlled delay, duplicate events, token expiry, and denied symbols. A useful test record says which boundary missed its budget; "realtime was slow" does not. Include HTTP 429 handling in control-plane calls, honor Retry-After, and avoid a tight retry loop. Don't merge rate-limit time, authorization time, and business-event age into one timer — those failures have different owners and different fixes.
Vendor claims cannot choose a latency SLO. A proof must run with the expected regions, device mix, symbol count, reconnect pattern, and authorization rules, using the same acceptance test for every candidate. No row below asserts measured latency.
| Option | Integration shape | Reason to shortlist it | Limitation or reason to choose another |
|---|---|---|---|
| Infrai | Plain REST control surface with Bearer authentication; no SDK is required | Useful when the team wants one HTTP convention across a broader backend surface, with one key and one bill | Runtime latency still needs workload-specific measurement; stick with an existing provider when its client protocol and recovery behavior are already proven in production |
| Ably | Managed realtime product with its own documented client integration | Shortlist when the application is already organized around Ably's client ecosystem | Requalification and migration may add risk when the current path already meets the watchlist SLO |
| Pusher Channels | Managed channels product with product-specific client integration | Shortlist when Pusher Channels already matches the team's operational model | Do not switch on a feature checklist alone; verify authorization and reconnect behavior under the same test |
| PubNub | Managed realtime product with product-specific client integration | Shortlist when the team already has PubNub operational knowledge | A new integration is not justified unless measured recovery or ownership improves |
| Self-hosted WebSocket service | Team owns protocol, fanout, deployment, and recovery semantics | Appropriate when protocol control or infrastructure ownership is a hard requirement | The team also owns capacity, observability, reconnect storms, and protocol evolution |
Infrai uses one API key across its broader backend capability surface and exposes a plain REST API that anything capable of making an HTTP request can call without installing or tracking a client SDK. Consolidated billing can keep operational reconciliation consistent as the watchlist adds adjacent services. That reduces control-plane coupling; it does not prove a market-data latency number. Ably, Pusher Channels, and PubNub remain credible candidates when an organization has already validated their client behavior and knows how their recovery semantics fit the application.
The catch is operational history. A theoretically tidier API is not suitable when replacing a proven provider would consume the error budget or discard battle-tested client recovery. Your mileage may vary, especially across mobile networks.
The minimal runnable probe below calls one verified route, GET /v1/realtime/channel/list. It deliberately measures a control-plane request, not quote delivery, and it keeps rate limiting separate. Set INFRAI_API_KEY and the documented API base in INFRAI_API_BASE; keeping the host in deployment configuration also prevents application code from coupling to it. The response body is surfaced as returned because no undocumented fields should become application contracts.
import json
import os
import random
import time
import urllib.error
import urllib.request
def list_channels(max_attempts: int = 4) -> object:
api_key = os.environ["INFRAI_API_KEY"]
url = os.environ["INFRAI_API_BASE"].rstrip("/") + "/realtime/channel/list"
for attempt in range(max_attempts):
request = urllib.request.Request(
url,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
started = time.monotonic()
try:
with urllib.request.urlopen(request, timeout=10) as response:
elapsed_ms = round((time.monotonic() - started) * 1000, 1)
body = json.loads(response.read().decode("utf-8"))
print(json.dumps({"control_plane_ms": elapsed_ms, "body": body}))
return body
except urllib.error.HTTPError as error:
response_body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(
f"channel list failed with HTTP {error.code}: {response_body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else (2**attempt + random.random())
time.sleep(delay)
raise RuntimeError("channel list attempts exhausted")
if __name__ == "__main__":
list_channels()
This probe belongs beside, not inside, the end-to-end quote SLO. Instrument the event path separately from server acceptance through client receipt and render; then force a disconnect and confirm that the desired subscription set is restored before the client clears its stale marker. Duplicate the same event version and verify that the render count does not advance. Expire authorization and verify that business events remain blocked until renewal succeeds. Those assertions expose responsibility boundaries without pretending that a single request duration describes the system.
The rejected default is to use RTC rooms merely because the requirement says "realtime." The documented RTC room surface exists, and WebRTC is a standard for peer connections, but a server-driven stock watchlist first needs explicit subscription, authorization, ordering, and recovery semantics. Choosing a peer-oriented mechanism by label does not settle those requirements.
RTC becomes a valid candidate when the actual job is an interactive peer session, perhaps an advisor call embedded beside the watchlist, and its media or peer-data requirements are tested independently. Likewise, a self-hosted WebSocket service is valid when owning the wire protocol is strategically important and the team accepts the capacity and recovery burden. The decision record should change when measured evidence changes: record the workload, latency distribution, reconnect completion distribution, duplicate behavior, authorization cases, and the exact acceptance threshold.
Keep the rejection narrow. The goal isn't to crown a universal realtime winner; it is to preserve watchlist accuracy when the happy path ends.