How to download all images from a list of URLs or web pages in bulk

# images# ecommerce# python# automation
How to download all images from a list of URLs or web pages in bulkHay Equipos

Turn a list of image links or web pages into saved image files, each with a download link, format, width, height, file size and SHA256, while skipping duplicates and broken links.

You have a spreadsheet of product image links from a supplier feed, or a list of pages whose images you need for a catalog, a migration or a dataset. Right clicking and saving each one is out of the question, and a quick script soon runs into broken links, HTML error pages that pretend to be images, tiny icons and the same file under three different URLs.

This guide shows how to handle the whole list in one run with a small Apify Actor published by Hay Equipos called Bulk Image Downloader: Image URLs and Page Images to Files.

What the tool returns

It works in two ways, and you can mix them in one run:

  1. Image URLs. You already have direct links. The Actor downloads each one.
  2. Page URLs. You name web pages. The Actor reads each page's HTML and collects every img (taking the largest size from srcset, plus lazy loading attributes such as data-src), picture sources, and the share image from og:image and twitter:image. Then it downloads them.

Each file is checked by its first bytes, so a web page served at an image address is rejected. Supported formats are jpg, png, webp, gif, svg, avif, ico, bmp, tiff and heic. Every saved image gets a row like this (values are illustrative):

{
  "imageUrl": "https://shop.example.com/media/products/blue-mug-large.jpg",
  "pageUrl": "https://shop.example.com/mugs/",
  "alt": "Blue ceramic mug",
  "foundIn": "img",
  "success": true,
  "fileName": "blue-mug-large.jpg",
  "storeKey": "9f2c41aa07be-blue-mug-large.jpg",
  "fileUrl": "https://api.apify.com/v2/key-value-stores/<storeId>/records/9f2c41aa07be-blue-mug-large.jpg?signature=<signature>",
  "format": "jpg",
  "contentType": "image/jpeg",
  "bytes": 184220,
  "width": 1200,
  "height": 1200,
  "sha256": "9f2c41aa07be..."
}
Enter fullscreen mode Exit fullscreen mode

fileUrl is a signed link that downloads the image directly, with no API token needed, so you can drop it into a spreadsheet, a CMS or another tool. Rows that failed or were skipped have success: false and a plain error, such as The server answered HTTP 404, Not an image (content type text/html) or Same file already saved in this run.

Step by step in the Apify Console

  1. Open the Actor from its Apify Store page and sign in to Apify Console.
  2. In the Input tab, paste direct links into Image URLs and/or pages into Page URLs, one per line.
  3. Set Maximum images per page and Maximum images per run (up to 20,000).
  4. To skip icons and tracking pixels, set Minimum width and Minimum height in pixels.
  5. Optionally limit formats under Only these formats, and set Maximum file size (25 MB by default).
  6. Leave Skip identical files on, so the same bytes from different URLs are saved once.
  7. If you want to keep the files after the run's own storage expires, enter a name under Save into a named store.
  8. Click Start. Open the Output tab for the rows, or the run's Storage tab, Key value store, to browse and download the files.

How to call it from code

With curl:

curl -X POST "https://api.apify.com/v2/acts/pistachio_implementation~bulk-image-downloader/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"pageUrls": ["https://books.toscrape.com/"], "minWidth": 100, "maxImagesPerPage": 20}'
Enter fullscreen mode Exit fullscreen mode

In Python, with the apify-client package, you can run it and then save every file locally through the signed links:

import os
import urllib.request
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("pistachio_implementation/bulk-image-downloader").call(
    run_input={"imageUrls": ["https://www.w3.org/Icons/w3c_home.png"], "allowedTypes": ["jpg", "png", "webp"]}
)
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    if row.get("success") and row.get("fileUrl"):
        urllib.request.urlretrieve(row["fileUrl"], row["storeKey"])
Enter fullscreen mode Exit fullscreen mode

Pricing

Pay per event: $0.002 per image saved, which is $2 per 1,000 images. There is no start fee and no platform usage charge on top. Broken links, files that are not images, files over the size limit, images below your minimum size, formats you excluded, duplicates and robots.txt skips are all free. You can cap the spend of any run with the maximum charge setting in Apify.

Limits and what it does not do

  • Page mode reads the HTML the server sends. Images that a page adds later with JavaScript, CSS background images and images inside iframes are not collected.
  • No getting around blocks. Some sites refuse downloads from cloud servers or need a login. Those files come back as free error rows. The Actor does not try to get past blocks, logins or captchas.
  • It respects robots.txt by default and identifies itself honestly as ApifyImageDownloader. Turn robots.txt checking off only for sites you own or may download from.
  • Polite pacing. Requests to the same site are spaced out (about 4 per second for files, 1 per second for pages), while different sites are fetched in parallel.
  • Some sizes stay empty. A few formats (AVIF, HEIC, some SVGs) do not state their size in a way the Actor reads without decoding the whole file. The file is still saved.
  • No ZIP output. Files are stored one by one, each with its own link.
  • Up to 20,000 images per run and 100 MB per file. Files in a run's default store follow your Apify data retention, so use a named store to keep them.

Whether you may use a given image is between you and its owner. Download your own images, images you have a license for, or uses the law allows, and respect each source site's terms.

Try it on the Apify Store: https://apify.com/pistachio_implementation/bulk-image-downloader