Hay EquiposTurn a list of news and blog URLs into clean text or Markdown with title, author, publish date, site, tags and word count.
Copying an article out of a web page sounds easy until you try it at scale. You get menus, cookie banners, related links and ads mixed into the text, and the author and publish date are hidden in a different place on every site. If you feed that into an LLM, a RAG index or a media monitoring sheet, the noise comes with it.
The Article Extractor, published by Hay Equipos on the Apify Store, takes a list of article URLs and returns one clean row per article: the main text without the clutter, plus title, author, publish and update dates, site name, language, section, tags, lead image, word count and reading time. It uses Mozilla Readability (the engine behind Firefox Reader View) for the text, and reads the publisher's own metadata (Open Graph, article tags, JSON LD and time tags) for the facts, so dates and authors come from the source rather than guesses.
One row per URL (illustrative values, text shortened):
{
"url": "https://blog.example.com/2026/09/why-we-rewrote-our-api",
"success": true,
"statusCode": 200,
"title": "Why we rewrote our API",
"author": "Example Engineering Team",
"publishedAt": "2026-09-18T08:00:00.000Z",
"modifiedAt": "2026-09-19T10:15:00.000Z",
"siteName": "Example Blog",
"language": "en",
"section": "Engineering",
"tags": ["api", "architecture"],
"leadImageUrl": "https://blog.example.com/images/api-cover.png",
"wordCount": 1480,
"readingTimeMinutes": 6,
"isAccessibleForFree": null,
"text": "Two years ago our API had grown into...",
"markdown": "Two years ago our API had grown into..."
}
Failed rows have success: false and an error that says why: HTTP error, bot check, not HTML, too little text, or robots.txt.
curl:
curl -X POST \
"https://api.apify.com/v2/acts/pistachio_implementation~article-extractor/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls": ["https://en.wikipedia.org/wiki/Web_scraping"], "includeMarkdown": true, "maxTextLength": 20000}'
Python with apify-client, keeping only successful rows for an LLM pipeline:
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("pistachio_implementation/article-extractor").call(
run_input={
"urls": [
"https://blog.apify.com/what-is-web-scraping/",
"https://en.wikipedia.org/wiki/Web_scraping",
],
"includeText": True,
"includeMarkdown": True,
}
)
docs = []
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
if row["success"]:
docs.append({"title": row["title"], "date": row["publishedAt"], "body": row["markdown"]})
else:
print("skipped", row["url"], row["error"])
The same rows work for a content audit of your own site: filter for posts where author, publishedAt or leadImageUrl is empty.
Pay per event: $0.0015 per article extracted ($1.50 per 1,000 articles), charged only for rows with at least 100 words of main text. Apify's own start event of $0.00005 per run also applies. Failed URLs, pages that are not articles, bot checks and robots.txt skips are free. You can set a maximum charge per run and the actor stops when it is reached.
isAccessibleForFree tells you when the publisher marks an article as paid.You are responsible for having the right to use the content you extract. Please use it in line with each publisher's terms.
Try it on the Apify Store: https://apify.com/pistachio_implementation/article-extractor