QuietDesk StudioA few weeks back I wrote about rate-limiting and backoff patterns for MCP servers under load. The...
A few weeks back I wrote about rate-limiting and backoff patterns for MCP servers under load. The feedback that came back wasn't "give me more backoff code" — it was "okay, but how do I know when I'm actually done? What's the full list I check before I say this server can take real traffic?"
That's a fair question, and it's also a trap. The honest answer is that "production-ready" isn't one checklist — it's several, covering different failure domains, and the posts that try to cram all of them into one listicle end up shallow on every single item. So this piece does one thing: it covers load and resilience — the specific question of whether your server degrades gracefully or falls over when traffic spikes, retries pile up, or a downstream dependency gets slow.
It does not cover authentication, OAuth token handling, tool permission scoping, or incident response after something breaks. Those are real, separate checklists with their own failure modes, and bolting them onto this one would make all of them worse. If you need that ground covered, that's exactly what our AI Agent Incident Postmortem & Permission-Scoping Template Pack is for — it's the companion piece for the "something went wrong, now what" and "who's allowed to call this tool" questions this post intentionally skips.
MCP sits in an unusual spot: the caller isn't a human clicking slowly, it's an agent that can retry a failed call in milliseconds and doesn't get bored. A load checklist for a normal REST API assumes human-paced traffic with occasional bursts. An MCP server has to assume an agent that, on a bad day, will call the same tool in a tight loop because its reasoning told it the first nine calls "almost worked."
That's the specific risk this checklist is built around. It's not a general "is my API fast" checklist — it's "will this server survive an agent that doesn't know when to stop."
1. Per-tool rate limits, not just per-connection. A global rate limit protects your server from one noisy client. It does nothing to stop a single agent from hammering your most expensive tool while staying under the global cap. Set limits per tool, per session.
2. A documented maximum concurrency per session. Decide how many in-flight tool calls one session is allowed to have before you start queuing or rejecting. If you haven't picked a number, the number is "however many the agent sends," which is not a number you want to discover in production.
3. Timeouts on every external call the tool makes, not just the top-level request. A tool that calls three downstream services needs three timeout budgets, not one umbrella timeout that masks which dependency actually hung.
4. Your server returns a retry signal the agent can actually use. A bare 429 or 503 with no Retry-After header just teaches the agent to retry immediately, which is the opposite of what you want under load.
5. You've tested what happens when the agent ignores the backoff signal. Some clients respect Retry-After. Plenty don't, especially home-grown agent loops. Your server's behavior under a client that retries immediately anyway is the real test, not the happy path where everyone cooperates.
6. There's a circuit breaker on your slowest downstream dependency. If one backend call starts timing out, the server should stop sending it traffic for a cooldown window instead of queuing requests behind a dependency that's already failing.
7. You've decided what "degraded" looks like, and it's not silence. When a rate limit kicks in, does the agent get a clear, structured error it can reason about, or does the connection just hang until it times out client-side? The second one looks like a bug to whoever's debugging it a week later.
8. Queue depth has a ceiling. An unbounded queue in front of a slow tool doesn't prevent failure, it just delays it and makes the eventual failure bigger. Pick a max queue size and reject past it with a clear error.
9. You've actually load-tested with concurrent sessions, not just concurrent requests from one session. These look similar on a chart and behave completely differently in practice — session-level state (locks, in-memory caches, connection pools) often breaks only under genuine multi-session concurrency.
10. You know your p95 latency under load, not just your average. An average that looks fine can hide a long tail where 5% of calls take 10x as long — and that tail is exactly where agent timeouts and retry storms start.
Here's the one that catches people who've done everything above correctly: backoff signals only work if your transport actually passes them through. If you're running MCP over a reverse proxy or a load balancer with its own retry logic, your carefully-tuned Retry-After header can get stripped or overridden before it reaches the client. Test the full path — agent to proxy to server and back — not just the server's response in isolation. This is the kind of thing that works perfectly in a local test and silently breaks the moment there's infrastructure in between.
To be specific about the boundary, this checklist does not cover:
If you're past the load question and into "how do we write down what happened and lock down permissions so it doesn't happen again," that's the gap our permission-scoping and incident postmortem template pack is built to close — it gives you a structured format for the postmortem itself and a worksheet for tightening scopes afterward, rather than starting from a blank doc at 2am.
Treat the 10 items above as pass/fail, not as "mostly handled." A server that passes 8 of 10 doesn't fail gracefully 80% of the time — it fails exactly the same way, just slightly less often, which is a worse debugging experience because the failure looks rare instead of systemic. Run through all 10 before you call a server load-tested, and keep the list next to your deploy checklist so the next server doesn't start the conversation from zero.
Written with AI assistance and reviewed for accuracy.