crawl4ai MCP Server: Extract LLM-Ready Markdown & Structured Data from Any Site
crawl4ai MCP Server: Extract LLM-Ready Markdown & Structured Data from Any Site
The crawl4ai MCP Server, built in Python, provides a focused solution for developers needing to feed web content into LLMs and AI agents. It specializes in converting arbitrary web pages into two primary formats: clean, LLM-ready Markdown and structured JSON data, addressing the common challenge of context acquisition for AI.
From Web Page to Clean Markdown
crawl4ai's core strength lies in its ability to transform web pages into Markdown optimized for LLM consumption. This isn't just a raw dump; it includes several strategies to ensure the output is genuinely useful:
- Clean Markdown Generation: The server produces Markdown with proper headings, lists, tables, and code blocks, structured for easy LLM parsing.
- Content Filtering: To prevent LLMs from ingesting irrelevant boilerplate,
crawl4aiincludes filters likePruningContentFilterLXMLto remove menus and footers. For query-specific content,BM25ContentFiltercan be applied, andLLMContentFilteroffers an AI-driven approach to relevance. - Citations: Page links are automatically converted into a numbered reference list within the Markdown, providing traceability for AI agents.
- Customization: Developers can plug in their own custom Markdown generation strategies if the defaults don't meet specific requirements.
For those who prefer not to manage local browser instances, crawl4ai offers a cloud service where a POST /scrape endpoint returns this processed Markdown directly.
Structured Data Extraction for AI
Beyond Markdown, crawl4ai excels at extracting structured data, a critical capability for agents needing specific information. It offers multiple extraction methodologies:
- Schema-Based Extraction: For predictable structures,
JsonCssExtractionStrategy,JsonXPathExtractionStrategy, andRegexExtractionStrategyenable fast, LLM-free data extraction using defined CSS selectors, XPath expressions, or regular expressions. - Schema Generation: A
generate_schemautility helps describe desired data once, then writes a reusable schema for consistent extraction. - LLM-Powered Extraction: For more complex or dynamic data,
LLMExtractionStrategyallows integration with any LLM provider (open-source or hosted) to extract data into a typed JSON schema. This offloads the schema definition and parsing to the LLM itself. - Content Chunking: Long pages can be broken down using topic, regex, or sentence chunking, making it easier to manage context windows for LLMs.
- Query-Based Chunk Retrieval: The
CosineStrategycan be used to find relevant chunks that match a specific query using cosine similarity.
Similar to Markdown generation, the cloud version provides a POST /extract endpoint, abstracting away the need for local LLM keys or browser control.
Running crawl4ai Locally
Developers can interact with crawl4ai directly from the command line for quick queries. For instance, to ask a question about a page and extract information, an LLM key is required, configured via crwl config:
crwl https://www.example.com/products -q "Extract all product prices"For more programmatic control, crawl4ai exposes an asynchronous Python API. An example of generating clean, filtered Markdown demonstrates its usage:
import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
from crawl4ai.content_filter_strategy import PruningContentFilterLXML
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator
async def main():
run_config = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS,
# ... additional configuration for crawler
)
# ... further code to instantiate and run the crawlerThis snippet illustrates how developers can import components like AsyncWebCrawler, PruningContentFilterLXML, and DefaultMarkdownGenerator to build custom scraping and processing workflows. The cache_mode setting, for example, allows bypassing the cache for fresh content.