YvesSterling6854Short answer: treat an upload as unpublishable until moderation has written a durable decision, and...
Short answer: treat an upload as unpublishable until moderation has written a durable decision, and make every read path enforce that state. Optimistic publish is safe only for a private, quarantined object; it is a bug in the product contract when a public URL can race ahead of review. The missing pending state is usually a data-model and cache problem, not an image-format problem.
This matters in a B2B SaaS product where an image can appear in a dashboard, an email preview, a search index, and a customer-facing page within seconds. A reviewer may still be deciding while another request has already promoted the asset. The result is a few seconds of banned content being visible, which is long enough for a screenshot, a webhook, or a crawler to copy it.
Use an explicit state machine rather than a nullable reviewed_at field. A useful minimum is pending, approved, rejected, and expired. The upload transaction creates the asset in pending, stores the original in a quarantine bucket, and emits a moderation job. Only a moderation decision can move it to approved. A rejected asset never becomes public, even if a stale worker retries an older job.
The public read path should ask for an approved publication, not merely find an asset by ID. That distinction prevents a common failure mode: the API returns a database row with status = pending, while the CDN still has a previously generated public URL. Keep the URL for the quarantined object private and issue a separate delivery key after approval. Revocation then has a concrete target: remove or deny the delivery key, purge the cache, and leave the audit record intact.
Here is a small Python service sketch. It uses generic interfaces so the ordering is visible; the storage and queue implementations can be replaced without changing the safety rule.
from dataclasses import dataclass
from enum import Enum
from typing import Protocol
class ReviewState(str, Enum):
PENDING = "pending"
APPROVED = "approved"
REJECTED = "rejected"
EXPIRED = "expired"
@dataclass(frozen=True)
class Asset:
asset_id: str
state: ReviewState
delivery_key: str | None = None
class AssetStore(Protocol):
def get(self, asset_id: str) -> Asset: ...
def publish_if_pending(self, asset_id: str, delivery_key: str) -> bool: ...
class Delivery(Protocol):
def signed_url(self, key: str, ttl_seconds: int) -> str: ...
def public_url(asset_id: str, store: AssetStore, delivery: Delivery) -> str | None:
asset = store.get(asset_id)
if asset.state is not ReviewState.APPROVED or asset.delivery_key is None:
return None
return delivery.signed_url(asset.delivery_key, ttl_seconds=300)
def approve(asset_id: str, key: str, store: AssetStore) -> bool:
# The conditional update makes a repeated job harmless.
return store.publish_if_pending(asset_id, key)
The important line is the guard in public_url, not the choice of queue or object store. A request that sees pending returns a deliberate 404 or a documented 202 response, with no public image bytes. A request that sees rejected returns the same non-public result; exposing the reason belongs in an authenticated review interface. That keeps the safety boundary consistent across web, API, and background consumers.
Start with one asset ID and reconstruct its timeline from logs. You need the upload commit, moderation enqueue, moderation result, publication update, URL issuance, and CDN fetch. Include a monotonic event sequence and the state observed at each boundary. Wall-clock timestamps alone can mislead when workers run on different hosts.
The rule is simple.
The race often looks like this:
| Time | Component | Event | State observed |
|---|---|---|---|
| t0 | API | row and quarantine object created | pending |
| t1 | API | response includes image_url
|
pending |
| t2 | Browser | preview requests that URL | public bytes |
| t3 | Worker | moderation marks content rejected | rejected |
| t4 | CDN | cached response remains available | rejected in DB |
At t1, the API has accidentally treated an upload URL as a delivery URL. At t4, the database is correct but the edge cache still violates the policy. Fix both layers. The browser can show a local preview from the upload response, but that preview must be a blob URL or an authenticated quarantine URL, never the production delivery hostname. I also inspect the response headers at this point, because a cache-control mistake can hide the real owner of the leak: public, max-age=86400 on a pending response is effectively a publication decision. The safe response is private, short-lived, and free of a reusable delivery key. In a service with multiple read paths, I replay the same asset through the HTML page, JSON API, thumbnail endpoint, export worker, and notification renderer. One forgotten path is enough to recreate the incident, even when the main page is correct.
I keep a regression fixture with three assets: a clearly safe image, a clearly banned image, and an undecidable image that requires review. The test publishes each asset through the same endpoint used by the UI, then polls every 50 milliseconds for two seconds. Any public response before approval is a failure. A 202 while pending is fine; a 200 with image bytes is not. This catches optimistic UI code that unit tests miss.
Also inspect caches outside the CDN. A reverse proxy, signed-URL middleware, browser service worker, or GraphQL normalized cache can preserve an old approved answer. Cache keys should include the publication version, and a rejected decision should advance that version before invalidation. If invalidation is asynchronous, the contract must still deny the old key at the origin; purging is a latency optimization, not the authorization check.
Image decoding is a separate trust boundary. Validate the declared MIME type, sniff the file signature, enforce pixel and compressed-size limits, and decode with a library that rejects malformed input. Normalize orientation and re-encode a safe derivative before review when your policy permits it. The browser's file extension is not evidence of content type; format details and browser support vary, so use a maintained reference such as MDN's image format guide.
Moderation coverage is the decision axis here. A classifier can return a score, labels, or an indeterminate result. Map all indeterminate responses to pending, not approved. Make the threshold and model version part of the decision record so a later policy change can re-review assets without guessing which rule produced the original result.
The catch is throughput. Holding every upload for human review increases queue time and storage, while an aggressive automatic threshold increases false positives. This design is not suitable for a product that promises instant public galleries with no quarantine window; in that case, use a clearly bounded preview surface and choose a moderation service with the latency and coverage your policy can defend. Stick with a simpler publish flow when images are never customer-visible and the risk owner accepts local previews only.
Exercise the state transitions with retries, duplicate moderation messages, delayed cache invalidation, and a worker that crashes after writing the decision but before issuing a delivery key. The transition must be idempotent: a second approval cannot overwrite a rejection, and a second rejection cannot delete the audit trail. Return a correlation ID to support, and log policy version, classifier version, state before, state after, and actor for every transition. Do not log the image itself.
Measure the interval from upload commit to approval, the count of pending assets older than the review SLA, and every attempted public read of a non-approved asset. Alert on a nonzero count of successful bytes served for pending or rejected states. A metric that only counts moderation errors will miss the more dangerous case where moderation succeeded but publication ignored the result.
Finally, test the contract at the edge. Fetch the delivery URL without cookies, with an expired signature, and after rejection. Verify that the response contains no image bytes and that a cached 200 cannot be replayed. I'm not sure which cache product your deployment uses, but the invariant is portable: authorization must consult the current publication state, and the public path must never infer approval from the existence of an upload.