Actor/apify

Website Content Crawler

Crawl websites and extract text content to feed AI models, LLM applications, vector databases, or RAG pipelines. The Actor supports rich formatting using Markdown, cleans the HTML, downloads files, and integrates well with 🦜🔗 LangChain, LlamaIndex, and the wider LLM ecosystem.

Publisher description, verbatim
Last 7 dayson this actor

No changes seen in the last 7 days.

The morning one happens, followers hear first.Follow this actor
Key figures30 readings
Pricefreefree
Users / 30d11,837↑ 2,004 over 30 readings
Users, all time160,320↑ 10,184 over 30 readings
Runs / 30d3.2m↑ 621,880 over 30 readings
Runs / user / mo269runs per monthly user
Store search presence
11top-10 keywords

best #1 for “content”

  • content#1
  • content crawler#1
  • content scraper#1
  • 8 more#2–#9
re-crawled daily
Recorded changes
  • 2026-09-10schemaInput schema changed—on record
  • 2026-09-08schemaInput schema changed—on record
  • 2026-09-07reviewsNew reviews+8 (233 total)+8233 total
showing 3 of 4 in the last 30 days · followers get all 4 by email

What the reviews say

Good at accurate, fast and easy extraction for many sites, but often unreliable on some URLs and CAPTCHAs, costly, and limited in page and site coverage.

65 of 65 reviews carry text
ThemeSentimentSaidWhat they mean
Does it run2 praised · 7 complained9works reliably in Cheerio modeserves as a fallback for n8n http“Did not work. The content on the site is in a JavaScript-rendered table. Even though this actor claims to render JavaScript (tested with both recommended options), it returned nothing. Instead, it just cycled through pages and ran up usage costs.” — 1/5
What it costs0 praised · 6 complained6consumes credits and runs up usage costsemail extraction requires payment“Why does this need to be a paid AI application? I figured it would scrape what it determines as "important" pages but this is simply scraping every page of a website, this could easily be done with a simple python script and not be paid.” — 1/5
Does it return everything2 praised · 4 complained6comprehensive extraction with multiple formatting optionsfails to extract emails and website links“Good not great, scarping behind a newpaper paywall i have access to. does some urls, not others, when i redo the ones missed it again does some and not others. basically from a list of 50 urls I've had to scrape 4 times to get the 50” — 3/5
Setup and docs6 praised · 1 complained7easy to use and speeds workflowfails on arbitrary pasted URLs“I am amazed, this is so cool. The scraping with Google Maps was so straightforward, but i notice i couldn't scrape for email without paying.” — 5/5
How fast1 praised · 1 complained2generally fast response timescan be very slow occasionally“Je l’aime parce qu’il est rapide” — 4/5
Caps and ceilings0 praised · 1 complained1hard limit on pages per website“Can we maybe not crawl all pages, create a hard limit on pages per website?” — 5/5
What it supports1 praised · 1 complained1covers most common websitesstruggles with tricky websites“Very good Crawler. Works better than others and is useful for most of the websites. Keen to see if this can be improved for other websites that are tricky to handle.” — 5/5
Is the data right1 praised · 0 complained1high crawl accuracy and quality“Very happy with the Crawl quality and ability to select different formattings” — 5/5

Does it run

2 ↑ · 7 ↓ of 9
works reliably in Cheerio modeserves as a fallback for n8n http“Did not work. The content on the site is in a JavaScript-rendered table. Even though this actor claims to render JavaScript (tested with both recommended options), it returned nothing. Instead, it just cycled through pages and ran up usage costs.” — 1/5

What it costs

0 ↑ · 6 ↓ of 6
consumes credits and runs up usage costsemail extraction requires payment“Why does this need to be a paid AI application? I figured it would scrape what it determines as "important" pages but this is simply scraping every page of a website, this could easily be done with a simple python script and not be paid.” — 1/5

Does it return everything

2 ↑ · 4 ↓ of 6
comprehensive extraction with multiple formatting optionsfails to extract emails and website links“Good not great, scarping behind a newpaper paywall i have access to. does some urls, not others, when i redo the ones missed it again does some and not others. basically from a list of 50 urls I've had to scrape 4 times to get the 50” — 3/5

Setup and docs

6 ↑ · 1 ↓ of 7
easy to use and speeds workflowfails on arbitrary pasted URLs“I am amazed, this is so cool. The scraping with Google Maps was so straightforward, but i notice i couldn't scrape for email without paying.” — 5/5
4 more themes · read all 65 →
Every count is a review we read. 41 of the 65 said only that they liked it, with nothing specific — those are excluded from the themes above. Quotes are one reviewer, unedited.Read all 65 →

History

30 readings since 2026-08-27
Users / 30d30 readings
08-27↑ 2,00409-25
Users, all time30 readings
08-27↑ 10,18409-25
Runs / 30d30 readings
08-27↑ 621,88009-25

Follow this actor, compare it against its rivals, and get every overnight change in one morning report.

Publish it yourself? Claim your publisher page and its full history is yours free, forever.

Open in the app Claim this page

Numbers on this page are measured only by an independent census of the Apify Store — not self-reported.