Your Readiness Probe Is Lying to You

# kubernetes# go# devops# grpc
Your Readiness Probe Is Lying to YouAvneet bansal

A readiness probe is a promise. It tells Kubernetes "route traffic to me, I can handle it." Most of...

A readiness probe is a promise. It tells Kubernetes "route traffic to me, I can handle it." Most of the time the promise is kept because the probe checks the thing that actually matters. The failures worth writing about are the ones where the probe keeps saying Ready while the component has quietly lost the one connection it needs to do any useful work. The traffic keeps arriving. None of it can be served. Nothing looks broken on the dashboard.

I hit exactly this in KEDA, the CNCF event-driven autoscaler, and the fix merged as kedacore/keda #8226. The bug is a good teaching case on its own. The part I did not expect was the second bug hiding inside the fix, which a maintainer caught in review.

TL;DR

  • KEDA's Metrics Server reported Ready even when its gRPC connection to the keda-operator was down.
  • In that state the HPA silently stops receiving scale signals, while the replica still looks healthy.
  • The fix makes the readiness probe observe the real gRPC connection state and report NotReady when the connection is unusable.
  • A naive version of that fix has a trap: once a replica goes NotReady it stops receiving RPCs, so an Idle connection never gets woken and the replica can stay NotReady forever.
  • The merged fix nudges an Idle connection to reconnect from inside the readiness check itself.

The setup: who talks to whom

KEDA has two moving parts here. The keda-operator does the actual work of talking to scalers and computing metrics. The Metrics Server (a custom metrics API server) is what the Kubernetes HPA queries when it wants to know "how many of these should be running." The Metrics Server does not compute metrics itself. It forwards the request to the operator over a gRPC connection and relays the answer.

That gRPC connection is the Metrics Server's entire reason to exist. If it is down, the Metrics Server is a shell. It can accept an HPA query and has nothing to answer it with.

The lie

The Metrics Server's readiness probe did not check that connection. It reported Ready as long as the process was up and its HTTP server was listening. So when the gRPC connection to the operator dropped, the following happened, all at once and all invisible: the replica stayed Ready, Kubernetes kept it in the APIService endpoints, the HPA kept sending it metric queries, and every one of those queries failed or hung because there was no live operator connection behind them. The autoscaler stops scaling, and the one signal you would look at, pod readiness, is green.

This is the specific shape of outage that is worst to debug. A crash is loud. A readiness probe that passes while the component cannot serve is silent, and it points you away from the real cause, because the first thing anyone checks is "are the pods ready," and they are.

The fix, version one

The fix is conceptually simple. Make readiness reflect the dependency. gRPC already tracks connection state through grpc.ClientConn.GetState(), which returns one of Idle, Connecting, Ready, TransientFailure, or Shutdown. Crucially, GetState() is a pure observer. It reports the current state and never initiates a connection, so it is safe to call from a probe with no side effects.

So version one added a small accessor on the metrics-service client that returned the connection state, and wired a readyz check in the adapter that reported NotReady unless the connection was Ready. When the operator connection drops, the probe fails, Kubernetes pulls the replica from the APIService endpoints, and the HPA stops sending it queries it cannot answer. When the connection recovers, the probe passes again. That is the behavior you want.

The second bug, caught in review

Here is the part I like. A maintainer looked at version one and pointed out a failure mode that only exists because the fix works.

Think through what happens after the probe correctly reports NotReady. The replica gets removed from the APIService endpoints. Now no RPCs flow through that client at all, because nothing is routed to it. And gRPC, by default, parks a connection in the Idle state after a period with no RPCs (GRPC_CLIENT_IDLE_TIMEOUT_MS, 30 minutes by default). An idle gRPC connection does not reconnect on its own. It waits for someone to use it.

So the naive fix can deadlock the recovery. The connection goes bad, the probe reports NotReady, the replica stops getting RPCs, the connection drifts to Idle, and nothing ever wakes it, because the only thing that would wake it is an RPC, and RPCs only resume once the probe reports Ready again, which requires the connection to be awake. The replica can sit NotReady permanently, needing a manual restart to recover. The fix for a silent hang would have introduced a silent hang.

The resolution is to make the readiness check active, not passive, in exactly one narrow case. When the accessor observes an Idle connection, it triggers a non-blocking reconnect before returning:

// GetConnectionState returns the current connectivity state of the underlying
// gRPC connection to the KEDA metrics service.
//
// When the connection is Idle it also triggers a non-blocking reconnection
// attempt (moving it towards Connecting) before returning. Once a non-Ready
// state removes the replica from the APIService endpoints, no RPCs flow through
// this client, so nothing else would ever wake an Idle connection.
func (c *GrpcClient) GetConnectionState() connectivity.State {
    state := c.connection.GetState()
    if state == connectivity.Idle {
        c.connection.Connect()
    }
    return state
}
Enter fullscreen mode Exit fullscreen mode

Connect() is non-blocking and a no-op when the connection is not Idle, so it is safe to call from a probe. The probe still returns the pre-nudge state (so a just-nudged Idle connection reports NotReady that cycle), but the connection is now moving toward Connecting and then Ready on its own. The replica recovers without a human. The probe went from a pure observer to an observer-with-one-side-effect, and that single side effect is what closes the recovery loop.

Reproduce it yourself

The regression test does not need a cluster or a running operator. It exercises the connection-state logic directly with an in-process gRPC client, so you can run it from a clean checkout:

git clone --depth 20 https://github.com/kedacore/keda.git
cd keda
go test ./pkg/metricsservice/... -run 'TestGetConnectionState' -v
Enter fullscreen mode Exit fullscreen mode

Two tests cover the two behaviors above. One asserts that a freshly created client (gRPC does not dial until first use) is not Ready, so a probe built on it correctly reports NotReady for a replica that has not connected yet. The other asserts the review fix: calling the accessor on an Idle connection moves it off Idle within a couple of seconds without any RPC being issued, so a sidelined replica can recover on its own.

Expected: both subtests pass. (If you are on a network that cannot reach the Go module proxy, the clone's large dependency tree may need go clean -modcache or a working GOPROXY first.)

How to adopt the pattern

The KEDA-specific takeaway is small: upgrade to a release that includes #8226 and your Metrics Server will stop advertising readiness while blind to the operator.

The portable takeaway is the one worth internalizing. A readiness probe should check the dependency that the component cannot work without, not just that the process is up. For a forwarding service, that dependency is the downstream connection. And when you make a probe gate traffic on a connection's health, check whether removing traffic from an unhealthy replica also removes the thing that would let it recover. If recovery depends on activity, and your probe stops activity, you need the probe itself to drive recovery, carefully and without blocking. Check the real dependency, and make sure the healthy path back exists.

A probe that only ever reports the process is alive is answering a question nobody asked. The question Kubernetes is actually asking is "can you serve." Answer that one.


If you run forwarding or adapter services behind readiness probes, it is worth auditing what they actually check. The gap between "process is up" and "can serve" is where the quiet outages live.