Scale VectorDuring a high-concurrency flash sale—such as ten limited-inventory items going live at 08:00...
During a high-concurrency flash sale—such as ten limited-inventory items going live at 08:00 AM—traffic does not arrive in a smooth, manageable curve. It arrives as an instantaneous impulse function: 100,000 requests hitting the ingress layer within the same millisecond window.
In a conventional architecture where an API gateway proxies requests directly to an application cluster backed by a relational database, this volume will not merely oversell inventory. It causes a cascading system collapse across network, gateway, and storage layers.
Why unbuffered architectures fail during sudden demand spikes:
TIME_WAIT state, surfacing cascading HTTP 504 Gateway Timeout errors.SELECT ... FOR UPDATE) causes queueing across connection pools, stalling unrelated services across the platform.A resilient architecture shifts concurrency management out of the relational engine:
[ 100,000 Users ]
│
▼ (1. Rate limit & In-Memory Check)
[ Edge / Ingress ] ──(Lua Atomic Decr)──► [ Redis Cluster ]
│ │
│ (Only 10 requests pass) │ (99,990 rejected instantly)
▼ ▼
[ Kafka Queue ] ────────────────────────► [ Background Workers ]
│
▼ (Optimistic Lock)
[ PostgreSQL / SQL Server ]
Before an HTTP payload reaches user space, the Linux kernel manages connection handshakes across two primary queues:
Client ──► [ SYN Queue ] ──► [ Accept Queue ] ──► Application accept()
▲
└── Saturated queue: Packets dropped silently
net.core.somaxconn to conservative values (such as 128 or 4096). Under an instantaneous 100,000-connection burst, the SYN backlog overflows immediately.ETIMEDOUT). The application logs zero errors because packets are dropped at the kernel boundary.Reverse proxies and load balancers (such as Envoy, Nginx, or an AWS ALB) maintain outbound connections toward backend instances.
32768–60999, providing roughly 28,232 ports).TIME_WAIT state for the duration of 2 * MSL (typically 60 seconds).HTTP 504 Gateway Timeout.For requests that reach the relational database, naive inventory deduction relies on row-level pessimistic locking:
-- Anti-pattern: Naive pessimistic row lock
SELECT stock FROM inventory WHERE id = 42 FOR UPDATE;
UPDATE inventory SET stock = stock - 1 WHERE id = 42;
Warning: When row 42 is locked exclusively, the first transaction acquires the lock while subsequent concurrent transactions queue behind it. A standard application pool of 100 connections is fully saturated in milliseconds. Once the pool is depleted, every other endpoint sharing the pool—including authentication, user profiles, and product catalogs—stalls entirely.
To protect the database from lock contention, the architecture must filter non-viable requests in memory at the edge.
Stock validation and reservation are executed atomically in memory using an embedded Lua script on Redis:
-- Atomic stock reservation in Redis
local stock = redis.call('get', KEYS[1])
if stock and tonumber(stock) > 0 then
redis.call('decr', KEYS[1])
return 1 -- Reservation acquired
else
return 0 -- Insufficient stock; reject
end
Successful reservations are not persisted via synchronous writes. Instead, they are published to a durable append-only log:
[ Checkout Event ] ──► [ Kafka Topic: orders ] ──► [ Consumer Pool ]
This decouples request ingress from disk I/O, allowing persistent database writes to execute at a steady, controlled rate irrespective of the traffic burst.
When background consumers persist the order into the primary relational store, they use Optimistic Concurrency Control (OCC) rather than exclusive locks:
-- Recommended: Optimistic concurrency control
UPDATE inventory
SET allocated = allocated + 1, version = version + 1
WHERE product_id = 42 AND version = 5;
If the version check fails due to a concurrent write, zero rows are updated. The worker backs off with jitter and retries. This eliminates database-level row lock waits and prevents thread pool starvation.
| Layer | Failure Mode | Mitigation Strategy |
|---|---|---|
| Linux Kernel | SYN backlog queue saturation | Tune sysctl -w net.core.somaxconn=65535 and net.ipv4.tcp_max_syn_backlog=65535
|
| Reverse Proxy | Ephemeral port exhaustion (TIME_WAIT) |
Enable HTTP keep-alives and sysctl -w net.ipv4.tcp_tw_reuse=1
|
| Application | Connection pool exhaustion | Shift initial inventory decrement to in-memory store (Redis) |
| Database | Lock queues and connection starvation | Replace pessimistic row locks with optimistic concurrency control (version) |
When architecting high-demand flash inventory systems, does your team rely on edge virtual waiting rooms (such as Cloudflare Waiting Room or AWS Virtual Waiting Room) to shape ingress traffic, or do you absorb spikes in-memory using distributed caching and message brokers?
Share your production experiences and trade-offs in the comments below.
The visual animated breakdown is available on YouTube:
Watch: What Actually Happens When 100,000 Users Click 'Buy' Concurrently
ScaleVector publishes deep-dive engineering breakdowns covering distributed systems, systems programming, and cloud infrastructure.