What Actually Happens When 100,000 Users Click 'Buy' at the Exact Same Millisecond

# architecture# backend# scalability# systemdesign
What Actually Happens When 100,000 Users Click 'Buy' at the Exact Same MillisecondScale Vector

During a high-concurrency flash sale—such as ten limited-inventory items going live at 08:00...

During a high-concurrency flash sale—such as ten limited-inventory items going live at 08:00 AM—traffic does not arrive in a smooth, manageable curve. It arrives as an instantaneous impulse function: 100,000 requests hitting the ingress layer within the same millisecond window.

In a conventional architecture where an API gateway proxies requests directly to an application cluster backed by a relational database, this volume will not merely oversell inventory. It causes a cascading system collapse across network, gateway, and storage layers.



Architecture at a Glance

Why unbuffered architectures fail during sudden demand spikes:

  • Network Layer (Linux Kernel): Incoming SYN packets saturate kernel socket queues, causing silent connection drops before the application runtime receives execution time.
  • Gateway Layer (Reverse Proxy): High connection churn exhausts ephemeral ports into TIME_WAIT state, surfacing cascading HTTP 504 Gateway Timeout errors.
  • Storage Layer (Database): Pessimistic locking (SELECT ... FOR UPDATE) causes queueing across connection pools, stalling unrelated services across the platform.

A resilient architecture shifts concurrency management out of the relational engine:

[ 100,000 Users ]
       │
       ▼ (1. Rate limit & In-Memory Check)
[ Edge / Ingress ] ──(Lua Atomic Decr)──► [ Redis Cluster ] 
       │                                         │
       │ (Only 10 requests pass)                 │ (99,990 rejected instantly)
       ▼                                         ▼
[ Kafka Queue ] ────────────────────────► [ Background Workers ]
                                                 │
                                                 ▼ (Optimistic Lock)
                                          [ PostgreSQL / SQL Server ]
Enter fullscreen mode Exit fullscreen mode

Failure Modes Across the Stack

1. Kernel Backlog Saturation (Network Layer)

Before an HTTP payload reaches user space, the Linux kernel manages connection handshakes across two primary queues:

Client ──► [ SYN Queue ] ──► [ Accept Queue ] ──► Application accept()
                 ▲
                 └── Saturated queue: Packets dropped silently
Enter fullscreen mode Exit fullscreen mode
  • The Failure Mechanism: Default kernel configurations often set net.core.somaxconn to conservative values (such as 128 or 4096). Under an instantaneous 100,000-connection burst, the SYN backlog overflows immediately.
  • Observed Behavior: Clients experience stalled TLS handshakes and connection timeouts (ETIMEDOUT). The application logs zero errors because packets are dropped at the kernel boundary.

2. Ephemeral Port Starvation (Gateway Layer)

Reverse proxies and load balancers (such as Envoy, Nginx, or an AWS ALB) maintain outbound connections toward backend instances.

  • Outbound connections allocate an ephemeral port from the local range (typically 32768–60999, providing roughly 28,232 ports).
  • When connections close rapidly without connection reuse, sockets transition to the TIME_WAIT state for the duration of 2 * MSL (typically 60 seconds).
  • Observed Behavior: The proxy exhausts its ephemeral port pool, unable to open new sockets to upstream backends. Callers receive HTTP 504 Gateway Timeout.

3. Database Connection Pool and Lock Exhaustion (Storage Layer)

For requests that reach the relational database, naive inventory deduction relies on row-level pessimistic locking:

-- Anti-pattern: Naive pessimistic row lock
SELECT stock FROM inventory WHERE id = 42 FOR UPDATE;
UPDATE inventory SET stock = stock - 1 WHERE id = 42;
Enter fullscreen mode Exit fullscreen mode

Warning: When row 42 is locked exclusively, the first transaction acquires the lock while subsequent concurrent transactions queue behind it. A standard application pool of 100 connections is fully saturated in milliseconds. Once the pool is depleted, every other endpoint sharing the pool—including authentication, user profiles, and product catalogs—stalls entirely.


Core Architectural Patterns for Flash Traffic

To protect the database from lock contention, the architecture must filter non-viable requests in memory at the edge.

1. Atomic In-Memory Reservation (Redis and Lua)

Stock validation and reservation are executed atomically in memory using an embedded Lua script on Redis:

-- Atomic stock reservation in Redis
local stock = redis.call('get', KEYS[1])
if stock and tonumber(stock) > 0 then
    redis.call('decr', KEYS[1])
    return 1 -- Reservation acquired
else
    return 0 -- Insufficient stock; reject
end
Enter fullscreen mode Exit fullscreen mode
  • Latency: Sub-millisecond execution (< 0.5 ms).
  • Throughput: Only the initial 10 requests obtain a successful return code. The remaining 99,990 requests are rejected immediately at the ingress boundary without touching persistent storage.

2. Asynchronous Buffering (Kafka / Event Queue)

Successful reservations are not persisted via synchronous writes. Instead, they are published to a durable append-only log:

[ Checkout Event ] ──► [ Kafka Topic: orders ] ──► [ Consumer Pool ]
Enter fullscreen mode Exit fullscreen mode

This decouples request ingress from disk I/O, allowing persistent database writes to execute at a steady, controlled rate irrespective of the traffic burst.

3. Optimistic Concurrency Control on Settlement

When background consumers persist the order into the primary relational store, they use Optimistic Concurrency Control (OCC) rather than exclusive locks:

-- Recommended: Optimistic concurrency control
UPDATE inventory 
SET allocated = allocated + 1, version = version + 1
WHERE product_id = 42 AND version = 5;
Enter fullscreen mode Exit fullscreen mode

If the version check fails due to a concurrent write, zero rows are updated. The worker backs off with jitter and retries. This eliminates database-level row lock waits and prevents thread pool starvation.


Production Tuning Reference

Layer Failure Mode Mitigation Strategy
Linux Kernel SYN backlog queue saturation Tune sysctl -w net.core.somaxconn=65535 and net.ipv4.tcp_max_syn_backlog=65535
Reverse Proxy Ephemeral port exhaustion (TIME_WAIT) Enable HTTP keep-alives and sysctl -w net.ipv4.tcp_tw_reuse=1
Application Connection pool exhaustion Shift initial inventory decrement to in-memory store (Redis)
Database Lock queues and connection starvation Replace pessimistic row locks with optimistic concurrency control (version)

Discussion

When architecting high-demand flash inventory systems, does your team rely on edge virtual waiting rooms (such as Cloudflare Waiting Room or AWS Virtual Waiting Room) to shape ingress traffic, or do you absorb spikes in-memory using distributed caching and message brokers?

Share your production experiences and trade-offs in the comments below.


Reference & Video Walkthrough

The visual animated breakdown is available on YouTube:
Watch: What Actually Happens When 100,000 Users Click 'Buy' Concurrently

ScaleVector publishes deep-dive engineering breakdowns covering distributed systems, systems programming, and cloud infrastructure.