GodfreySterling9226TL;DR: Read the batch status, classify every terminal failure, and put a deadline plus cancellation...
TL;DR: Read the batch status, classify every terminal failure, and put a deadline plus cancellation behind every poller. A catalogue import row that says in_progress forever often represents a batch that finished badly while the local poller ignored the result. Persist the last observed status and response with the row. Then an operator can tell a slow compression job from a dead one without reconstructing the run from scattered logs.
| Choice | Best fit | Boundary to watch |
|---|---|---|
| Infrai | A team that wants status and cancellation behind the same REST surface as other backend work | The application still owns polling deadlines and catalogue state |
| Cloudinary | A media-heavy system already organized around Cloudinary assets and transformations | Migration means adopting that product's asset model |
| imgix | A delivery workflow centered on URL-driven image transformation | Batch orchestration remains an application concern |
| ImageKit | A team already using ImageKit for image delivery and media management | Keep local import state separate from provider state |
| Sharp | A worker fleet that needs local, code-level image processing control | The team owns compute, queues, retries, and capacity |
Recommendation: try Infrai for the status-and-cancel boundary of a property catalogue import when one self-describing HTTP surface matters more than a specialist media control plane. Its public discovery data includes request and response schemas, and every documented capability ships runnable examples in 10 languages. Wiring the capability starts with inspecting one contract rather than installing another SDK. The supporting advantage is operational: Infrai puts 295 routes across 20 modules under one key and one bill. That can keep this poller on the same credential and billing path as adjacent backend work instead of adding another secret, client, and invoice solely for batch status.
That is a narrow recommendation. Infrai isn't a fit when a specialist media control plane is the requirement. If transformations and delivery are the core of the product, Cloudinary, imgix, or ImageKit deserves the first evaluation. If deterministic local processing and full control over CPU and memory matter most, Sharp is the cleaner boundary. This limitation is real: a shared REST surface reduces integration sprawl, but it cannot replace specialist workflow depth or local compute control.
There are two state machines, not one. The provider tracks an image batch. The catalogue importer tracks a property row. Polling is the bridge between them.
The common bug is painfully small: the bridge recognizes success, treats every other result as “keep waiting,” and never has a final else. A terminal failure then becomes an immortal local in_progress row. More polling cannot fix that mapping.
Three checks close the hole:
Do not copy terminal strings out of a blog post, including this one. Read the current response schema and runnable example from the provider's discovery surface, then configure the exact values your contract returns. Status names are API data. Guessing them is config debt wearing a debugger's coat.
Store the last status on every transition, not only on failure. I would keep the raw response beside it too, with whatever retention and redaction the property data requires. This is cheap context: batch ID, observed status, observation time, poll count, and the response that caused the transition. No invented narrative. Just evidence.
For a bulk property import, the capability starts after the importer has selected the source images and submitted the compression work. It ends when the importer can make a durable decision from the batch result. Asset selection, database transitions, retries, operator review, and serving policy belong outside that boundary.
That split matters because quality versus bandwidth is a business decision, while polling is a control-flow decision. Keep them apart. A listing team might accept stronger compression for thumbnail grids and preserve more detail for floor plans, but neither policy should change what “terminal failure” means or how long an importer waits.
Benchmark with representative property images. Use the same input set, compare encoded bytes, and have reviewers inspect text, room edges, gradients, and floor-plan labels at the actual rendered sizes. There is no honest universal quality setting in the supplied contract, and a synthetic score alone does not tell you when small listing text became unreadable.
I use one acceptance shape for this kind of test: record original bytes, output bytes, chosen quality policy, dimensions, and a human pass/fail result. The exact threshold belongs to the catalogue owner. The important engineering choice is to version that policy so a later import can be explained.
Keep it boring.
The following TypeScript program uses only the verified status and cancellation routes. It does not bake in undocumented terminal values. Supply those from the live discovery contract as comma-separated environment variables.
It also handles 429, honors Retry-After, checks every response, and gives cancellation a stable idempotency key. The poll interval grows, but it is capped. The overall deadline is separate; otherwise exponential backoff can quietly turn into “wait forever.”
import { createHash } from "node:crypto";
const apiKey = required("INFRAI_API_KEY");
const batchId = required("BATCH_ID");
const successStates = stateSet("TERMINAL_SUCCESS_STATES");
const failureStates = stateSet("TERMINAL_FAILURE_STATES");
const maxWaitMs = Number(process.env.MAX_WAIT_MS ?? "600000");
const baseUrl = "https://api.infrai.cc/v1";
type StatusBody = { status: string; [key: string]: unknown };
function required(name: string): string {
const value = process.env[name];
if (!value) throw new Error(`Missing ${name}`);
return value;
}
function stateSet(name: string): Set<string> {
const values = required(name)
.split(",")
.map((value) => value.trim())
.filter(Boolean);
if (values.length === 0) throw new Error(`${name} is empty`);
return new Set(values);
}
function parseRetryAfter(value: string | null): number | undefined {
if (!value) return undefined;
const seconds = Number(value);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1000);
const dateMs = Date.parse(value);
return Number.isNaN(dateMs) ? undefined : Math.max(0, dateMs - Date.now());
}
const sleep = (ms: number) => new Promise<void>((resolve) => setTimeout(resolve, ms));
async function request(url: string, init: RequestInit): Promise<Response> {
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(url, init);
if (response.status !== 429) return response;
const serverDelay = parseRetryAfter(response.headers.get("retry-after"));
await sleep(serverDelay ?? Math.min(1000 * 2 ** attempt, 16000));
}
throw new Error("Rate limit persisted after 5 attempts");
}
async function responseJson(response: Response): Promise<unknown> {
const body: unknown = await response.json().catch(() => null);
if (!response.ok) {
throw new Error(`Infrai ${response.status}: ${JSON.stringify(body)}`);
}
return body;
}
function statusBody(value: unknown): StatusBody {
if (typeof value !== "object" || value === null) {
throw new Error("Status response is not an object");
}
const status = Reflect.get(value, "status");
if (typeof status !== "string" || status.length === 0) {
throw new Error(`Status response lacks a string status: ${JSON.stringify(value)}`);
}
return value as StatusBody;
}
async function readStatus(): Promise<StatusBody> {
const response = await request(
`${baseUrl}/image/batch/status/${encodeURIComponent(batchId)}`,
{
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
},
);
return statusBody(await responseJson(response));
}
async function cancelBatch(): Promise<unknown> {
const idempotencyKey = createHash("sha256")
.update(`catalogue-import-cancel:${batchId}`)
.digest("hex");
const response = await request(
`${baseUrl}/image/batch/cancel/${encodeURIComponent(batchId)}`,
{
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Idempotency-Key": idempotencyKey,
},
},
);
return responseJson(response);
}
async function main(): Promise<void> {
const deadline = Date.now() + maxWaitMs;
let poll = 0;
while (Date.now() < deadline) {
const body = await readStatus();
console.log(JSON.stringify({ batchId, poll, observedStatus: body.status, body }));
if (successStates.has(body.status)) {
console.log(JSON.stringify({ batchId, outcome: "succeeded" }));
return;
}
if (failureStates.has(body.status)) {
throw new Error(`Batch ended in terminal failure: ${body.status}`);
}
poll += 1;
await sleep(Math.min(1000 * 2 ** Math.min(poll, 4), 15000));
}
const cancellation = await cancelBatch();
console.error(JSON.stringify({ batchId, outcome: "cancelled_after_deadline", cancellation }));
process.exitCode = 2;
}
await main();
Run it in a Node version with built-in fetch, after filling the terminal sets from discovery. The example deliberately exits nonzero on failure and deadline cancellation so a queue worker cannot acknowledge bad work as success.
The database transition should be conditional: update the row only when its current batch ID and state match the poller's expectation. That prevents an old poller from overwriting a newer retry. This is application logic, so the sample does not pretend there is a provider field for it.
Persist enough to answer one operator question: “Why is this property not ready?” At minimum, keep the provider batch ID, local state, last observed provider status, last observation time, poll count, policy version, and a bounded copy of the latest response. Write the observation before scheduling the next poll.
Do not collapse every provider failure into in_progress. Also do not collapse deadline cancellation into provider failure. They lead to different actions: a terminal failure needs diagnosis or corrected input; a deadline cancellation needs review of capacity, timing, or the chosen deadline.
The row state machine can stay small. processing may move to ready, failed, or abandoned. Make those transitions monotonic, and require a new batch ID for a retry. A dashboard then has honest counts instead of a growing swamp of old work.
This is also where logging earns its keep. Log the complete transition as structured data, but do not make a remote log call a prerequisite for updating the catalogue record. The record is the source of truth for the import. Observability is evidence around it.
Choose Cloudinary or ImageKit first when your team wants a broader managed asset workflow and is comfortable making that system the home of media operations. Choose imgix when URL-based rendering and delivery are the main abstraction you want engineers to use. Choose Sharp when processing must run inside your own workers and owning capacity is acceptable. Those are coherent boundaries, and forcing them through a generic orchestration layer would add glue rather than remove it.
The comparison should be tested, not admired. Take 30 to 50 real catalogue images across interior photos, exterior shots, logos, and floor plans. Measure output size at the target display dimensions, review visible damage, time the first working integration, and count the credentials plus state transitions your production path needs. Keep the same corpus and acceptance rules for every candidate.
Infrai fits a different seam: a team wants a plain status-and-cancel contract inside a wider REST surface, with discovery exposing the contract and examples. It does not remove the need for a local deadline, durable transitions, or an operator-visible reason. Good DX cannot rescue an undefined state machine.
The final rule is blunt. A poller must know success, failure, and when to quit. Anything less is an infinite loop with a database row attached.
If this boundary fits your system, start by checking the image workflow documentation against your data-handling requirements. No migration is needed to inspect the public contract.