Hay EquiposTurn a list of image links or web pages into saved image files, each with a download link, format, width, height, file size and SHA256, while skipping duplicates and broken links.
You have a spreadsheet of product image links from a supplier feed, or a list of pages whose images you need for a catalog, a migration or a dataset. Right clicking and saving each one is out of the question, and a quick script soon runs into broken links, HTML error pages that pretend to be images, tiny icons and the same file under three different URLs.
This guide shows how to handle the whole list in one run with a small Apify Actor published by Hay Equipos called Bulk Image Downloader: Image URLs and Page Images to Files.
It works in two ways, and you can mix them in one run:
img (taking the largest size from srcset, plus lazy loading attributes such as data-src), picture sources, and the share image from og:image and twitter:image. Then it downloads them.Each file is checked by its first bytes, so a web page served at an image address is rejected. Supported formats are jpg, png, webp, gif, svg, avif, ico, bmp, tiff and heic. Every saved image gets a row like this (values are illustrative):
{
"imageUrl": "https://shop.example.com/media/products/blue-mug-large.jpg",
"pageUrl": "https://shop.example.com/mugs/",
"alt": "Blue ceramic mug",
"foundIn": "img",
"success": true,
"fileName": "blue-mug-large.jpg",
"storeKey": "9f2c41aa07be-blue-mug-large.jpg",
"fileUrl": "https://api.apify.com/v2/key-value-stores/<storeId>/records/9f2c41aa07be-blue-mug-large.jpg?signature=<signature>",
"format": "jpg",
"contentType": "image/jpeg",
"bytes": 184220,
"width": 1200,
"height": 1200,
"sha256": "9f2c41aa07be..."
}
fileUrl is a signed link that downloads the image directly, with no API token needed, so you can drop it into a spreadsheet, a CMS or another tool. Rows that failed or were skipped have success: false and a plain error, such as The server answered HTTP 404, Not an image (content type text/html) or Same file already saved in this run.
With curl:
curl -X POST "https://api.apify.com/v2/acts/pistachio_implementation~bulk-image-downloader/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"pageUrls": ["https://books.toscrape.com/"], "minWidth": 100, "maxImagesPerPage": 20}'
In Python, with the apify-client package, you can run it and then save every file locally through the signed links:
import os
import urllib.request
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("pistachio_implementation/bulk-image-downloader").call(
run_input={"imageUrls": ["https://www.w3.org/Icons/w3c_home.png"], "allowedTypes": ["jpg", "png", "webp"]}
)
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
if row.get("success") and row.get("fileUrl"):
urllib.request.urlretrieve(row["fileUrl"], row["storeKey"])
Pay per event: $0.002 per image saved, which is $2 per 1,000 images. There is no start fee and no platform usage charge on top. Broken links, files that are not images, files over the size limit, images below your minimum size, formats you excluded, duplicates and robots.txt skips are all free. You can cap the spend of any run with the maximum charge setting in Apify.
ApifyImageDownloader. Turn robots.txt checking off only for sites you own or may download from.Whether you may use a given image is between you and its owner. Download your own images, images you have a license for, or uses the law allows, and respect each source site's terms.
Try it on the Apify Store: https://apify.com/pistachio_implementation/bulk-image-downloader