Vlad ZThe one I think about most is the account where the infrastructure was correct and the cost was...
The one I think about most is the account where the infrastructure was correct and the cost was still wrong
ML startup. Eight engineers. Production inference API with real traffic
Their AWS bill was $22,000 a month. Mostly GPU compute for inference
I went in expecting to find over-provisioned instances. I found the opposite. The instance types were appropriate for the workload. Utilization was reasonable. Nothing obviously wasteful
What I found instead was the scheduling
Their inference traffic followed a completely predictable pattern. 90 percent of requests came between 9am and 11pm US Eastern time. The overnight period was nearly silent. A few health checks. Background jobs. No real user traffic
The GPU instances ran 24 hours a day
At full capacity. Full price. Through eight hours of near-zero utilization every single night
We implemented a scale-down schedule. At 11pm Eastern, scale the inference cluster down to a minimal warm state - enough to handle the background jobs and respond to the health checks. At 8am Eastern, scale back up before the traffic arrives
The change took two days to implement and test safely
Monthly savings: $4,800
Not from changing the architecture. Not from changing the model. Not from changing anything about how the product worked
From matching the cost to the actual demand pattern instead of running at peak capacity all the time because that felt safer
The infrastructure was fine
The schedule was wrong
Sometimes that's all it is