Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
<a href="https://trendshift.io/repositories/11716" target="_blank"><img src="https://trendshift.io/api/badge/repositories/11716" alt="unclecode%2Fcrawl4ai | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>
Latest: v0.9.4 (23 Sep 2026) · all releases →
<a href="https://crawl4ai.com/?ref=readme-banner"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/assets/cloud-launch-banner-dark.svg"> <img alt="Crawl4AI Cloud is live. Soft launch: your first $10 is on us until 31 December 2026, no card. Get your key." src="https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/assets/cloud-launch-banner-light.svg" width="960"> </picture> </a> </div>Crawl4AI turns any website into clean, LLM-ready Markdown for RAG, AI agents and data pipelines. Run the open-source web crawler and scraper yourself, free forever, or use it hosted with one key: scrape, search and extract through one API, with MCP for your agent.
pip install -U crawl4ai
crawl4ai-setup # installs the browser, onceimport asyncio
from crawl4ai import AsyncWebCrawler
async def main():
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(url="https://news.ycombinator.com")
print(result.markdown)
asyncio.run(main())Docker server, CLI and every option: Installation · docs.crawl4ai.com
curl -s https://api.crawl4ai.com/scrape \
-H "Authorization: Bearer $CRAWL4AI_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://news.ycombinator.com"}' | jq -r .markdownThe same key works for /search, /answer, /extract and many URLs at once (/scrape/batch, /scrape/jobs). Pay as you go: live prices.claude mcp add --transport http crawl4ai https://api.crawl4ai.com/mcp \
--header "Authorization: Bearer $CRAWL4AI_KEY"| 🐍 Library | 🐳 Your own server | ☁️ Crawl4AI Cloud | |
|---|---|---|---|
| Runs the browsers | you, in your Python process | you, in Docker on your machine | we do |
| JS-heavy pages and bot walls | your settings, your proxies | your settings, your proxies | handled for you, automatically |
| Web search | – | – | /search and /answer |
| Price | free, forever | free (your hosting) | pay as you go; your first $10 is on us |
I grew up on an Amstrad, thanks to my dad, and never stopped building. In grad school I specialized in NLP and built crawlers for research. That’s where I learned how much extraction matters.
In 2023, I needed web-to-Markdown. The “open source” option wanted an account, API token, and $16, and still under-delivered. I went turbo anger mode, built Crawl4AI in days, and it went viral. Now it’s the most-starred crawler on GitHub.
I made it open source for availability, anyone can use it without a gate. Now I’m building the platform for affordability, anyone can run serious crawls without breaking the bank. If that resonates, join in, send feedback, or just crawl something amazing.
That platform is live now: Crawl4AI Cloud.
</details> <details> <summary>Why developers pick Crawl4AI</summary>PruningContentFilterLXML, BM25ContentFilter (for a query) and LLMContentFilter.☁️ Same in the cloud: POST /scrape returns this Markdown, with no browser to run. Docs →
JsonCssExtractionStrategy, JsonXPathExtractionStrategy, RegexExtractionStrategy).generate_schema writes a reusable schema.LLMExtractionStrategy).CosineStrategy).☁️ Same in the cloud: POST /extract, with no LLM key of your own. Docs →
enable_stealth, and an undetected-browser adapter for sites that detect automation.resume_state) for long crawls.AdaptiveCrawler stops when it has learned enough to answer your query.AsyncUrlSeeder (sitemaps, Common Crawl) and DomainMapper; prefetch=True finds URLs 5 to 10 times faster.scan_full_page) for infinite scroll and lazy images.srcset, internal and external links, iframes, metadata.raw: and file://.arun_many with a memory-adaptive dispatcher.☁️ Same in the cloud: up to 50 URLs in one streamed call, or 10,000 in a background job. Docs →
</details> <details> <summary>🐳 <strong>Self-hosting (Docker)</strong></summary>CRAWL4AI_API_TOKEN./md, /html, /crawl, /crawl/stream, /screenshot, /pdf, /execute_js.☁️ Rather not run a server? The cloud is the same idea, hosted. Get a key →
</details> <details> <summary>☁️ <strong>What the cloud adds</strong></summary>GET /search, browser-free, ranked and cleaned. Docs →GET /answer gives a direct answer to a question (experimental). Docs →POST /extract. Docs →<a id="installation"></a>
pip install -U crawl4ai
crawl4ai-setup # installs and sets up the browser
crawl4ai-doctor # checks the installationIf the browser setup fails, install it by hand:
python -m playwright install --with-deps chromiumPre-release versions: pip install crawl4ai --pre
Development install, for contributors:
git clone https://github.com/unclecode/crawl4ai.git
cd crawl4ai
pip install -e ".[all]" # or: pip install -e . (the core only)</details>
<details>
<summary>🐳 <strong>Docker server</strong></summary>
The server needs a token. Without one it answers only inside its container.
export CRAWL4AI_API_TOKEN="$(openssl rand -hex 32)"
docker run -d -p 11235:11235 --name crawl4ai --shm-size=1g \
-e CRAWL4AI_API_TOKEN="$CRAWL4AI_API_TOKEN" \
unclecode/crawl4ai:latestTest it (allow about 10 seconds for the start):
curl -s http://localhost:11235/md \
-H "Authorization: Bearer $CRAWL4AI_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"url": "https://news.ycombinator.com"}' | jq -r .markdownThe dashboard is at http://localhost:11235/dashboard, the playground at http://localhost:11235/playground. LLM keys, MCP and every setting: Self-hosting guide.
# A page as Markdown
crwl https://news.ycombinator.com -o markdown
# Deep crawl, breadth first, at most 10 pages
crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
# Ask a question about a page (needs an LLM key: crwl config)
crwl https://www.example.com/products -q "Extract all product prices"</details>
More in docs/examples.
<details> <summary>📝 <strong>Clean and fit Markdown</strong></summary>import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
from crawl4ai.content_filter_strategy import PruningContentFilterLXML
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator
async def main():
run_config = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS,
markdown_generator=DefaultMarkdownGenerator(
content_filter=PruningContentFilterLXML(threshold=0.48, threshold_type="fixed", min_word_threshold=0)
),
)
async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
result = await crawler.arun(url="https://en.wikipedia.org/wiki/Web_crawler", config=run_config)
print(len(result.markdown.raw_markdown), "characters of raw Markdown")
print(len(result.markdown.fit_markdown), "characters after the filter")
asyncio.run(main())</details>
<details>
<summary>🖥️ <strong>A JavaScript page and structured data, without an LLM</strong></summary>
import asyncio, json
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, JsonCssExtractionStrategy
schema = {
"name": "Quotes",
"baseSelector": "div.quote",
"fields": [
{"name": "text", "selector": "span.text", "type": "text"},
{"name": "author", "selector": "small.author", "type": "text"},
{"name": "tags", "selector": "a.tag", "type": "list", "fields": [{"name": "tag", "type": "text"}]},
],
}
async def main():
run_config = CrawlerRunConfig(
extraction_strategy=JsonCssExtractionStrategy(schema),
scan_full_page=True, # scroll to the end, so the page loads every quote
scroll_delay=0.5,
cache_mode=CacheMode.BYPASS,
)
async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
result = await crawler.arun(url="https://quotes.toscrape.com/scroll", config=run_config)
quotes = json.loads(result.extracted_content)
print(f"Extracted {len(quotes)} quotes")
print(json.dumps(quotes[0], indent=2))
asyncio.run(main())</details>
<details>
<summary>📚 <strong>Structured data with an LLM</strong></summary>
import os, asyncio
from pydantic import BaseModel, Field
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, CacheMode, LLMConfig, LLMExtractionStrategy
class ModelFee(BaseModel):
model_name: str = Field(..., description="Name of the model.")
input_fee: str = Field(..., description="Fee for input tokens.")
output_fee: str = Field(..., description="Fee for output tokens.")
async def main():
run_config = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS,
extraction_strategy=LLMExtractionStrategy(
# any provider LiteLLM supports, e.g. "ollama/llama3.3" with api_token="no-token"
llm_config=LLMConfig(provider="openai/gpt-4o-mini", api_token=os.getenv("OPENAI_API_KEY")),
schema=ModelFee.model_json_schema(),
extraction_type="schema",
instruction="Extract every model name with its input and output token fee.",
),
)
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(url="https://openai.com/api/pricing/", config=run_config)
print(result.extracted_content)
asyncio.run(main())</details>
<details>
<summary>🤖 <strong>Your own browser with a saved profile</strong></summary>
import os, asyncio
from pathlib import Path
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
async def main():
user_data_dir = os.path.join(Path.home(), ".crawl4ai", "browser_profile")
os.makedirs(user_data_dir, exist_ok=True)
browser_config = BrowserConfig(headless=True, user_data_dir=user_data_dir, use_persistent_context=True)
run_config = CrawlerRunConfig(cache_mode=CacheMode.BYPASS, magic=True)
async with AsyncWebCrawler(config=browser_config) as crawler:
result = await crawler.arun(url="ADDRESS_OF_A_CHALLENGING_WEBSITE", config=run_config)
print(result.success, len(result.markdown))
asyncio.run(main())</details>
We welcome contributions from the open-source community. Check out our contribution guidelines for more information.
This project is licensed under the Apache License 2.0, attribution is recommended via the badges below. See the Apache 2.0 License file for details.
When using Crawl4AI, you must include one of the following attribution methods:
<details> <summary>📈 <strong>1. Badge Attribution (Recommended)</strong></summary> Add one of these badges to your README, documentation, or website:| Theme | Badge |
|---|---|
| Disco Theme (Animated) | <a href="https://github.com/unclecode/crawl4ai"><img src="./docs/assets/powered-by-disco.svg" alt="Powered by Crawl4AI" width="200"/></a> |
| Night Theme (Dark with Neon) | <a href="https://github.com/unclecode/crawl4ai"><img src="./docs/assets/powered-by-night.svg" alt="Powered by Crawl4AI" width="200"/></a> |
| Dark Theme (Classic) | <a href="https://github.com/unclecode/crawl4ai"><img src="./docs/assets/powered-by-dark.svg" alt="Powered by Crawl4AI" width="200"/></a> |
| Light Theme (Classic) | <a href="https://github.com/unclecode/crawl4ai"><img src="./docs/assets/powered-by-light.svg" alt="Powered by Crawl4AI" width="200"/></a> |
HTML code for adding the badges:
<!-- Disco Theme (Animated) -->
<a href="https://github.com/unclecode/crawl4ai">
<img src="https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/assets/powered-by-disco.svg" alt="Powered by Crawl4AI" width="200"/>
</a>
<!-- Night Theme (Dark with Neon) -->
<a href="https://github.com/unclecode/crawl4ai">
<img src="https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/assets/powered-by-night.svg" alt="Powered by Crawl4AI" width="200"/>
</a>
<!-- Dark Theme (Classic) -->
<a href="https://github.com/unclecode/crawl4ai">
<img src="https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/assets/powered-by-dark.svg" alt="Powered by Crawl4AI" width="200"/>
</a>
<!-- Light Theme (Classic) -->
<a href="https://github.com/unclecode/crawl4ai">
<img src="https://raw.githubusercontent.com/unclecode/crawl4ai/main/docs/assets/powered-by-light.svg" alt="Powered by Crawl4AI" width="200"/>
</a>
<!-- Simple Shield Badge -->
<a href="https://github.com/unclecode/crawl4ai">
<img src="https://img.shields.io/badge/Powered%20by-Crawl4AI-blue?style=flat-square" alt="Powered by Crawl4AI"/>
</a></details>
<details>
<summary>📖 <strong>2. Text Attribution</strong></summary>
Add this line to your documentation:
```
This project uses Crawl4AI (https://github.com/unclecode/crawl4ai) for web data extraction.
```
</details>
If you use Crawl4AI in your research or project, please cite:
@software{crawl4ai2024,
author = {UncleCode},
title = {Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper},
year = {2024},
publisher = {GitHub},
journal = {GitHub Repository},
howpublished = {\url{https://github.com/unclecode/crawl4ai}},
commit = {Please use the commit hash you're working with}
}Text citation format:
UncleCode. (2024). Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper [Computer software].
GitHub. https://github.com/unclecode/crawl4aiOur mission is to unlock the value of personal and enterprise data by turning digital footprints into structured, useful assets. Crawl4AI gives individuals and organizations open-source tools to extract and structure data, and a fair way to benefit from it. Full mission statement →
These companies provide core infrastructure and technology that power Crawl4AI’s capabilities — from web access and proxy networks to AI tooling and data pipelines.
| Company | About |
|---|---|
| <a href="https://www.joinmassive.com/" target="_blank"><picture><source media="(prefers-color-scheme: dark)" srcset="docs/assets/sponsors/massive_light.svg"><source media="(prefers-color-scheme: light)" srcset="docs/assets/sponsors/massive.svg"><img alt="Massive" src="docs/assets/sponsors/massive.svg" height="40"/></picture></a> | Massive is a web access API backed by millions of volunteer devices in 195+ countries. AI agents, models, and data pipelines use it to reach any site on the internet, reliably, in real time, and at scale. |
Our enterprise sponsors support Crawl4AI and help scale it to power production-grade data pipelines.
| Company | About | Sponsorship Tier |
|---|---|---|
| <a href="https://kipo.ai" target="_blank"><img src="https://docs.crawl4ai.com/uploads/sponsors/20251013045751_2d54f57f117c651e.png" alt="DataSync" height="40"/></a> | Helps engineers and buyers find, compare, and source electronic & industrial parts in seconds, with specs, pricing, lead times & alternatives. | 🥇 Gold |
| <a href="https://www.kidocode.com/" target="_blank"><img src="https://docs.crawl4ai.com/uploads/sponsors/20251013045045_bb8dace3f0440d65.svg" alt="Kidocode" height="40"/></a> | Kidocode is a hybrid technology and entrepreneurship school for kids aged 5–18, offering both online and on-campus education. | 🥇 Gold |
| <a href="https://www.alephnull.sg/" target="_blank"><picture><source media="(prefers-color-scheme: dark)" srcset="docs/assets/sponsors/aleph_null_light.svg"><source media="(prefers-color-scheme: light)" srcset="docs/assets/sponsors/aleph_null.svg"><img alt="Aleph null" src="docs/assets/sponsors/aleph_null.svg" height="40"/></picture></a> | Singapore-based Aleph Null is Asia’s leading edtech hub, dedicated to student-centric, AI-driven education—empowering learners with the tools to thrive in a fast-changing world. | 🥇 Gold |
Interested in partnering with Crawl4AI?
Whether you’re a proxy provider, AI infrastructure company, cloud platform, or an organization looking to support the Crawl4AI ecosystem, we’d love to hear from you.
📩 Contact: hello@crawl4ai.com
A heartfelt thanks to our individual supporters! Every contribution helps us keep our opensource mission alive and thriving!
<p align="left"> <a href="https://github.com/hafezparast"><img src="https://avatars.githubusercontent.com/u/14273305?s=60&v=4" style="border-radius:50%;" width="64px;"/></a> <a href="https://github.com/ntohidi"><img src="https://avatars.githubusercontent.com/u/17140097?s=60&v=4" style="border-radius:50%;"width="64px;"/></a> <a href="https://github.com/Sjoeborg"><img src="https://avatars.githubusercontent.com/u/17451310?s=60&v=4" style="border-radius:50%;"width="64px;"/></a> <a href="https://github.com/romek-rozen"><img src="https://avatars.githubusercontent.com/u/30595969?s=60&v=4" style="border-radius:50%;"width="64px;"/></a> <a href="https://github.com/Kourosh-Kiyani"><img src="https://avatars.githubusercontent.com/u/34105600?s=60&v=4" style="border-radius:50%;"width="64px;"/></a> <a href="https://github.com/Etherdrake"><img src="https://avatars.githubusercontent.com/u/67021215?s=60&v=4" style="border-radius:50%;"width="64px;"/></a> <a href="https://github.com/shaman247"><img src="https://avatars.githubusercontent.com/u/211010067?s=60&v=4" style="border-radius:50%;"width="64px;"/></a> <a href="https://github.com/work-flow-manager"><img src="https://avatars.githubusercontent.com/u/217665461?s=60&v=4" style="border-radius:50%;"width="64px;"/></a> </p>Want to join them? Sponsor Crawl4AI →
Discord · X @unclecode · GitHub @unclecode · hello@crawl4ai.com
Happy crawling! 🕸️🚀
unclecode/crawl4ai
May 9, 2024
September 28, 2026
Python