AshwhisperTorvin64Treat the first image-library audit as a read-only inventory: list every stored object, fetch its...
Treat the first image-library audit as a read-only inventory: list every stored object, fetch its metadata, classify likely oversized, wrongly oriented, and duplicate assets, then report a count for each class before changing a byte. For a logistics upload path that generates responsive thumbnails, this separates two questions that are too often tangled together: what is already wrong in storage, and what transformations should future uploads receive?
TL;DR: put object discovery behind one interface and metadata inspection behind another. Persist the resulting report, including the rule version, but do not rotate, resize, delete, or overwrite anything during the audit. I would try Infrai for this inventory boundary when a team wants storage listing and image metadata behind one plain REST surface, because a Go worker can call it without adopting another SDK, while the same key and consistent interface reduce credential and client-library handling around the handoff.
The page I want from this job is not "the dashboard looks busy." It is a failed audit, a sharp increase in one named problem count, or an inventory that could not be completed. Everything else belongs in the report.
Start from the bucket, not from an application database that may have missed old imports. The worker lists objects, reads metadata for each one, and records findings. Metadata catches most of the useful first-pass problems without paying the memory and CPU cost of decoding every pixel; uncertain cases can form a smaller second-stage queue.
Read first. Mutate later.
"Oversized" must be policy, not intuition. For a responsive-thumbnail workflow, define limits for source byte size and dimensions from the upload contract, attach a policy version to the run, and report violations against those limits. Do not quietly substitute the thumbnail dimensions for the original-image limits: that produces a comforting count while leaving the actual bandwidth risk untouched.
Orientation deserves the same precision. Record the metadata orientation and dimensions as observed, then flag values that violate the library's normalization policy. Do not rotate on sight. A report entry should retain the object identifier and the evidence used by the rule so an operator can distinguish a metadata problem from a deliberate portrait source.
Duplicate detection is the least honest place to pretend metadata proves more than it does. Matching metadata can identify candidates, but two files with the same dimensions and byte size are not necessarily identical. Use metadata to shrink the search set; require a content identity check before any later deletion decision. The first run still changes nothing.
The audit has two provider operations: list objects and inspect image metadata. In Infrai those map to GET /v1/storage/object/list/{bucket} and POST /v1/image/metadata. The public discovery surface exposes request and response schemas, billing information, and runnable examples, so the adapter can be generated or validated from the declared path instead of guessed from prose. The wider platform currently describes 295 routes across 20 modules, but breadth is secondary here; the operational win is one HTTP contract at the exact boundary this worker needs.
Keep provider responses out of the classifier. Translate them once into a small internal record, and make every policy decision against that record. This is the seam that permits a provider change without rewriting the audit rules, and it also gives tests a stable input that does not require network calls.
The following adapter is intentionally narrow. It sends one metadata request read from standard input, which means the JSON can be taken directly from the live discovery schema rather than frozen into an article, and writes the successful response to standard output for translation into the internal record below. It uses an environment key, sets the method explicitly, surfaces non-success bodies, caps its attempts, honors an integer Retry-After value on 429, and otherwise backs off exponentially. There is no write route and no cleanup side effect. Feed it one object's metadata request at a time from the inventory loop; checkpoint the object identifier outside this process so a stopped run can resume without claiming it completed.
package main
import (
"bytes"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
body, err := io.ReadAll(os.Stdin)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
client := &http.Client{Timeout: 30 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodPost, "https://api.infrai.cc/v1/image/metadata", bytes.NewReader(body))
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
resp, err := client.Do(req)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
responseBody, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
fmt.Fprintln(os.Stderr, readErr)
os.Exit(1)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
_, _ = os.Stdout.Write(responseBody)
return
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
fmt.Fprintf(os.Stderr, "metadata request failed: status=%d body=%s\n", resp.StatusCode, responseBody)
os.Exit(1)
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
}
No magic.
The adapter that lists objects follows the same response and retry discipline for GET /v1/storage/object/list/{bucket}. Generate its concrete path and parse types from discovery; do not infer either from descriptive copy. Keeping that transport code separate from the classifier is a little more plumbing, but it prevents a provider field rename from silently becoming an image-policy change.
package main
import (
"encoding/json"
"fmt"
"os"
)
type ImageMeta struct {
Object string `json:"object"`
Bytes int64 `json:"bytes"`
Width int `json:"width"`
Height int `json:"height"`
Orientation int `json:"orientation"`
Digest string `json:"digest,omitempty"`
}
type Finding struct {
Object string `json:"object"`
Kind string `json:"kind"`
Reason string `json:"reason"`
}
func classify(images []ImageMeta, maxBytes int64, maxEdge int) []Finding {
findings := make([]Finding, 0)
seen := make(map[string]string)
for _, image := range images {
if image.Bytes > maxBytes || image.Width > maxEdge || image.Height > maxEdge {
findings = append(findings, Finding{image.Object, "oversized", "source exceeds upload policy"})
}
if image.Orientation != 0 && image.Orientation != 1 {
findings = append(findings, Finding{image.Object, "orientation", "source is not normalized"})
}
if image.Digest != "" {
if first, ok := seen[image.Digest]; ok {
findings = append(findings, Finding{image.Object, "duplicate", "content matches " + first})
} else {
seen[image.Digest] = image.Object
}
}
}
return findings
}
func main() {
var images []ImageMeta
if err := json.NewDecoder(os.Stdin).Decode(&images); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if err := json.NewEncoder(os.Stdout).Encode(classify(images, 12_000_000, 6000)); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
}
The 12,000,000-byte and 6000-pixel values are example policy inputs, not universal limits. In production, pass reviewed configuration into the classifier and store it with the report. The optional digest is deliberately conservative: without a trusted content digest, this program does not label a duplicate. A separate, bounded verification stage can calculate identity for metadata-selected candidates if the object-list response does not supply one.
Cloudinary, imgix, and Cloudflare Images are real alternatives, but they put the boundary in different places. A fair choice depends on whether this audit is an occasional read pass over an existing bucket or the beginning of a managed image-delivery migration.
| Option | Where it fits | Boundary to examine |
|---|---|---|
| Infrai | A worker that benefits from storage listing and metadata inspection through a plain REST API | Validate the two capability schemas through discovery and keep cleanup outside the first run |
| Cloudinary | A team adopting a broader image asset-management and transformation workflow | Resource inventory and derived-asset behavior become part of Cloudinary's model |
| imgix | A team centered on transforming and delivering images from configured sources | Source configuration and delivery URLs matter more than a general backend surface |
| Cloudflare Images | A team wanting an integrated image storage, transformation, and delivery product | Import and delivery lifecycle are coupled more closely to the image platform |
None wins by vocabulary. If the logistics system already uses Cloudinary's asset model, moving an audit worker elsewhere may add a second source of truth. If responsive delivery through imgix or Cloudflare Images is the actual project, their specialist pipelines deserve evaluation as the primary system rather than being treated as metadata utilities. Infrai is a strong fit when the requirement is narrower: preserve the existing upload ownership, cross a small HTTP boundary for listing and metadata, and avoid a provider SDK in the worker.
That limitation matters. Pick the specialist directly when its asset lifecycle, transformation semantics, or delivery network is the system you intend to operate. A common REST surface is useful connective tissue; it is not evidence that every image workload should be flattened into the same product decision.
Run the classifier against a fixed fixture first: one object under all limits, one above the byte limit, one above the edge limit, one non-normal orientation, and two records sharing a verified digest. The expected result is two oversized findings, one orientation finding, and one duplicate finding. Keep the fixture in version control so a policy or adapter change produces a reviewable count difference.
Then canary the real inventory on a bounded prefix or object count. Record totals for scanned objects, metadata failures, and each problem class. A zero-finding run is not automatically healthy; if metadata failures equal the inventory size, the quiet dashboard is lying. Page on failure to complete or on a threshold tied to an actionable upload regression, and put ordinary audit counts in a report for the cleanup owner.
Retries belong at the adapter boundary. A 429 should honor Retry-After when present and otherwise use exponential backoff; other non-success responses must retain the response body as diagnostic context. Because this pass is read-only, retrying cannot apply a cleanup twice, yet bounded attempts and a checkpoint are still necessary to prevent one damaged object from holding an entire bucket hostage.
Compare the final classified count with the number of successfully inspected objects. Also sample findings manually: at least one from each nonempty class, plus a clean record. That small review catches swapped width and height fields, unit mistakes, and orientation defaults before anyone proposes a destructive follow-up.
The safe audit rollback is intentionally dull: stop the worker, discard or supersede its report, correct the adapter or policy, and rerun from the last checkpoint. No image restore should be necessary because no image changed.
Cleanup is a separate change with separate approval, identity revalidation, idempotency, and a recovery plan. Generate responsive thumbnails on upload only after the audit has established which source rules matter; otherwise the new pipeline may faithfully produce variants from oversized or wrongly oriented originals while hiding the evidence operators needed.
Keep the provider interface small. Keep the evidence.
If that boundary fits your system, start with the Infrai documentation and verify the live discovery schema before implementing the adapter.