How to extract clean article text, author and date from news URLs

# nlp# llm# api# python
How to extract clean article text, author and date from news URLsHay Equipos

Turn a list of news and blog URLs into clean text or Markdown with title, author, publish date, site, tags and word count.

Copying an article out of a web page sounds easy until you try it at scale. You get menus, cookie banners, related links and ads mixed into the text, and the author and publish date are hidden in a different place on every site. If you feed that into an LLM, a RAG index or a media monitoring sheet, the noise comes with it.

The Article Extractor, published by Hay Equipos on the Apify Store, takes a list of article URLs and returns one clean row per article: the main text without the clutter, plus title, author, publish and update dates, site name, language, section, tags, lead image, word count and reading time. It uses Mozilla Readability (the engine behind Firefox Reader View) for the text, and reads the publisher's own metadata (Open Graph, article tags, JSON LD and time tags) for the facts, so dates and authors come from the source rather than guesses.

What you get back

One row per URL (illustrative values, text shortened):

{
  "url": "https://blog.example.com/2026/09/why-we-rewrote-our-api",
  "success": true,
  "statusCode": 200,
  "title": "Why we rewrote our API",
  "author": "Example Engineering Team",
  "publishedAt": "2026-09-18T08:00:00.000Z",
  "modifiedAt": "2026-09-19T10:15:00.000Z",
  "siteName": "Example Blog",
  "language": "en",
  "section": "Engineering",
  "tags": ["api", "architecture"],
  "leadImageUrl": "https://blog.example.com/images/api-cover.png",
  "wordCount": 1480,
  "readingTimeMinutes": 6,
  "isAccessibleForFree": null,
  "text": "Two years ago our API had grown into...",
  "markdown": "Two years ago our API had grown into..."
}
Enter fullscreen mode Exit fullscreen mode

Failed rows have success: false and an error that says why: HTTP error, bot check, not HTML, too little text, or robots.txt.

Step by step in the Apify Console

  1. Open the actor on the Apify Store and start it in the Console.
  2. Paste your Article URLs, one per line. The actor reads only these pages and does not follow links.
  3. Choose the formats: Include plain text (on by default), Include Markdown and Include cleaned HTML.
  4. Optionally set Maximum text length to cut long articles (0 means no limit).
  5. Leave Respect robots.txt on.
  6. Set Parallel sites (default 5, up to 10). Requests to the same site are always spaced one second apart.
  7. Click Start and export the dataset as JSON, CSV or Excel.

Calling it from code

curl:

curl -X POST \
  "https://api.apify.com/v2/acts/pistachio_implementation~article-extractor/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://en.wikipedia.org/wiki/Web_scraping"], "includeMarkdown": true, "maxTextLength": 20000}'
Enter fullscreen mode Exit fullscreen mode

Python with apify-client, keeping only successful rows for an LLM pipeline:

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])

run = client.actor("pistachio_implementation/article-extractor").call(
    run_input={
        "urls": [
            "https://blog.apify.com/what-is-web-scraping/",
            "https://en.wikipedia.org/wiki/Web_scraping",
        ],
        "includeText": True,
        "includeMarkdown": True,
    }
)

docs = []
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    if row["success"]:
        docs.append({"title": row["title"], "date": row["publishedAt"], "body": row["markdown"]})
    else:
        print("skipped", row["url"], row["error"])
Enter fullscreen mode Exit fullscreen mode

The same rows work for a content audit of your own site: filter for posts where author, publishedAt or leadImageUrl is empty.

Pricing

Pay per event: $0.0015 per article extracted ($1.50 per 1,000 articles), charged only for rows with at least 100 words of main text. Apify's own start event of $0.00005 per run also applies. Failed URLs, pages that are not articles, bot checks and robots.txt skips are free. You can set a maximum charge per run and the actor stops when it is reached.

Limits and what it does not do

  • Plain HTTP only, no browser. Pages that build their text with JavaScript, or sit behind a bot check or a login, come back as a free error row.
  • Some publishers refuse automated readers or tools run on Apify, either in robots.txt or at their server. The actor respects that, returns a free error row and does not disguise itself, so some large news sites will not work.
  • Paywalled articles return only what anonymous readers see. isAccessibleForFree tells you when the publisher marks an article as paid.
  • Home, section and video pages are rejected as "not an article" when they have under 100 words of main text.
  • One request per second per site, so 1,000 URLs from one site take about 17 minutes. URLs spread across many sites run in parallel.
  • Up to 5,000 URLs per run, and no crawling. Collect links first with a sitemap or RSS tool, then pass them in.
  • If no author is published, the field is left empty rather than guessed.

You are responsible for having the right to use the content you extract. Please use it in line with each publisher's terms.

Try it on the Apify Store: https://apify.com/pistachio_implementation/article-extractor