Guarding OCR Costs With Node.js Hash Dedupe for Supplier Images

# node# typescript# ocr
Guarding OCR Costs With Node.js Hash Dedupe for Supplier ImagesVaughnKnight3189

TL;DR: In a Node.js and Express catalogue importer, dedupe supplier images by using normalized...

TL;DR: In a Node.js and Express catalogue importer, dedupe supplier images by using normalized metadata as a quick filter and a SHA-256 hash of the original bytes as the final decision. Store one immutable source object per digest, cache OCR output by that digest plus an extractor version, and record why each upload was accepted or reused. This keeps repeated storage and processing out of the import without pretending that filenames or dimensions prove identity.

The short mental model is a funnel. Before: every row creates an image object and an OCR job. After: cheap metadata narrows the search, a byte hash establishes identity, and only a new digest crosses the processing boundary. Fast first, certain second.

How should Node.js dedupe supplier images with metadata and a hash?

Supplier feeds reuse names such as front.jpg. They also rename the same file, and two distinct product photos can share a MIME type, byte length, width, and height. Metadata is useful as an index hint, not an identity proof.

Names lie.

There is another trap: declared MIME type is input, not evidence. Inspect the file signature before decoding, then derive dimensions with an image parser that applies strict pixel and memory limits. MDN's format guide is a useful map of common image formats, but support policy still belongs to the importer. A catalogue team may accept JPEG, PNG, and WebP while explicitly rejecting formats its OCR preparation path cannot decode.

The hash should cover the untouched upload bytes. Do not resize, strip metadata, or transcode first. Those operations spend compute before the dedupe decision and can collapse different source files into the same derivative. Keep derivative identity separate: sourceDigest + transformVersion for normalized pixels, and sourceDigest + extractorVersion for OCR text. Consider one supplier row that calls a photo front.jpg, a second row that calls the exact bytes sku-104-main.jpg, and a third row that sends a newly compressed version at the same dimensions. The metadata index can quickly put the first two in the same candidate set, but only their digest permits reuse. The third remains a distinct source even if a person sees the same package, preserving the input needed to explain later OCR differences.

The digest is the source object's address; versions are the cache boundaries. That small distinction makes invalidation deliberate instead of mysterious.

A copyable two-stage intake

This compact Express example uses an in-memory body to keep the control flow visible. The 20 MiB ceiling is a sample policy, not a universal recommendation. For larger limits, stream to a quarantined temporary file while hashing, then inspect and promote it atomically.

import { createHash } from "node:crypto";
import express, { type Request, type Response } from "express";

type Candidate = {
  supplierId: string;
  supplierAssetId: string;
  filename: string;
  declaredType: string;
};

type SourceRecord = {
  digest: string;
  objectKey: string;
  byteLength: number;
  detectedType: "image/jpeg" | "image/png" | "image/webp";
};

interface SourceRepository {
  findByMetadata(key: string): Promise<SourceRecord[]>;
  findByDigest(digest: string): Promise<SourceRecord | undefined>;
  insertIfAbsent(record: SourceRecord, bytes: Buffer): Promise<SourceRecord>;
}

interface OcrQueue {
  enqueueOnce(cacheKey: string, source: SourceRecord): Promise<void>;
}

const normalize = (value: string): string =>
  value.trim().normalize("NFKC").toLowerCase();

const metadataKey = (candidate: Candidate, byteLength: number): string =>
  [candidate.supplierId, candidate.supplierAssetId, byteLength]
    .map(normalize)
    .join(":");

function detectType(bytes: Buffer): SourceRecord["detectedType"] {
  if (bytes.subarray(0, 3).equals(Buffer.from([0xff, 0xd8, 0xff]))) {
    return "image/jpeg";
  }
  if (bytes.subarray(0, 8).equals(Buffer.from([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]))) {
    return "image/png";
  }
  if (bytes.subarray(0, 4).toString("ascii") === "RIFF" &&
      bytes.subarray(8, 12).toString("ascii") === "WEBP") {
    return "image/webp";
  }
  throw new Error("Unsupported image signature");
}

export function createImageRouter(repo: SourceRepository, queue: OcrQueue) {
  const router = express.Router();
  const readImage = express.raw({ type: () => true, limit: "20mb" });

  router.post("/supplier-image", readImage, async (req: Request, res: Response) => {
    if (!Buffer.isBuffer(req.body) || req.body.length === 0) {
      res.status(400).json({ error: "An image body is required" });
      return;
    }

    const candidate: Candidate = {
      supplierId: String(req.header("x-supplier-id") ?? ""),
      supplierAssetId: String(req.header("x-supplier-asset-id") ?? ""),
      filename: String(req.header("x-filename") ?? ""),
      declaredType: String(req.header("content-type") ?? ""),
    };

    const key = metadataKey(candidate, req.body.length);
    const candidates = await repo.findByMetadata(key);
    const digest = createHash("sha256").update(req.body).digest("hex");
    const exactCandidate = candidates.find((item) => item.digest === digest);
    const existing = exactCandidate ?? await repo.findByDigest(digest);

    if (existing) {
      res.status(200).json({ digest, disposition: "reused" });
      return;
    }

    const detectedType = detectType(req.body);
    const proposed: SourceRecord = {
      digest,
      objectKey: `sha256/${digest.slice(0, 2)}/${digest}`,
      byteLength: req.body.length,
      detectedType,
    };

    // A unique digest constraint makes concurrent imports converge on one object.
    const source = await repo.insertIfAbsent(proposed, req.body);
    const extractorVersion = "ocr-text-v3";
    await queue.enqueueOnce(`${source.digest}:${extractorVersion}`, source);
    res.status(202).json({ digest: source.digest, disposition: "queued" });
  });

  return router;
}
Enter fullscreen mode Exit fullscreen mode

The repository operation is the critical contract. insertIfAbsent must be atomic under a unique digest constraint; a check followed by an unconditional insert still duplicates objects when two catalogue rows arrive together. The queue needs the same property through its cache key. HTTP status communicates the result: 200 for a known source, 202 for newly accepted asynchronous work, and 400 for an empty request.

Hash once.

Notice what the response omits. It does not expose a storage path, supplier filename, or OCR result that may not exist yet. It returns a stable identity and a disposition. Clear boundaries make retries boring.

Measure the funnel, not just request latency

An importer can look fast while quietly multiplying storage. Instrument the decision points. A counter such as image_intake_total{disposition="new|reused|rejected"} shows the outcome mix; image_intake_bytes_total{disposition="new|reused"} turns that mix into storage pressure. Record hash duration and OCR queue delay as histograms. Keep supplier identifiers out of low-cardinality metric labels and put them in access-controlled structured logs instead.

One trace can connect intake, object promotion, and OCR enqueue. Add the digest as a trace attribute only if the observability system's data policy permits content-derived identifiers. Logs should include the request correlation ID, normalized metadata key, digest, detected media type, byte count, disposition, and extractor version. Do not log image bytes or extracted catalogue text by default.

The useful alert is tied to behavior. Alert when the rejection ratio changes sharply, when accepted-new images stop producing queued jobs, or when queue age breaches the catalogue import's service objective. A rising reuse ratio is usually informational; it may mean a supplier resent a feed, which is exactly the work the funnel should absorb.

Here is the before/after check I would put on the deployment dashboard:

Signal Before content addressing After content addressing
Source objects written One per accepted row One per unique byte digest
OCR cache identity Filename or row ID Digest plus extractor version
Retry effect May create more work Converges on the same keys
Audit question "Which row ran?" "Why was this source new, reused, or rejected?"

No invented savings percentage belongs in that table. The real storage avoided during a window is the sum of byte lengths for valid uploads classified as reused. The real OCR work avoided is the reused count that already has a result for the active extractor version. Those two numbers are defensible and directly useful for capacity planning.

What about visually identical files with different bytes?

Byte hashing will treat a recompressed JPEG and its original as different sources. That is correct for the first gate: cryptographic equality answers a precise question and preserves provenance. Perceptual hashing answers a fuzzier one. Similar-looking package variants, language labels, or tiny compliance marks can matter in a catalogue, so automatically merging them risks attaching the wrong OCR text to a product.

If near-duplicate detection is valuable, run it after exact dedupe as a separate review signal. Decode within resource limits, normalize orientation and color handling, compute a perceptual fingerprint, then surface close matches with the original supplier and product context. Do not use that fingerprint as the object key.

This is a deliberate trade-off. Exact hashing may retain extra recompressions, but it makes deletion, provenance, and cache correctness explainable. A later similarity workflow can optimize storage without weakening the intake contract.

Does hashing every upload waste compute?

Hashing reads every accepted byte, so metadata should still narrow likely matches and reject malformed envelopes early. Yet metadata alone cannot safely authorize reuse. In a streaming implementation, hashing happens during the same pass that writes the quarantined file; the system does not need a second network read.

Test the boundary with duplicate bytes under different names, different bytes under the same metadata, truncated signatures, oversized bodies, and two simultaneous requests carrying one digest. Also test extractor upgrades: changing ocr-text-v3 to a new version should enqueue fresh OCR work without duplicating the source object. That distinction is the whole cost model.

Roll out with disposition metrics visible, compare unique bytes written against total valid bytes received, and retain enough structured decision data to explain a reuse. Then the catalogue import has a compact rule: metadata accelerates the lookup, SHA-256 establishes byte identity, and versioned keys decide which derived work can be reused.

Further reading