Crawl websites locally, extract Markdown and inspect SEO issues through a desktop app, CLI or MCP connection, with plan-dependent exports.
crawler.sh
Explore features, practical uses and pricing below.
crawler.sh is a local website crawler with a desktop interface, command-line tool and MCP connection. It collects pages, extracts readable content and inspects technical SEO issues. Developers can use the extracted material as input for a retrieval system or an AI agent, while site owners can review links, metadata and redirects.
The product overview separates these local tools from an upcoming cloud API. The crawler supplies source material and diagnostics; it does not itself establish that an AI system trained or grounded on that material will answer correctly.
The CLI guide documents controls for depth, page limits, concurrency and delay. Crawls stay within the configured site scope. JavaScript rendering can be detected automatically or selected explicitly, which matters when a page's useful content appears only after scripts run.
Stored crawl results can be inspected and exported without starting the same collection again. Markdown extraction provides readable page content, while SEO checks identify conditions such as missing metadata, duplicate titles, redirects and broken outgoing links. A flagged condition is a review lead rather than a complete judgment about a page's quality.
The desktop interface provides a live crawl feed, content viewer and issue panels. The CLI suits scripted work and repeatable commands. The MCP server exposes local crawling and result tools to supported AI clients, giving an assistant a way to inspect collected material through an authorized connection.
Begin with a documentation section whose content you are entitled to use. Set a narrow start URL, a modest page limit and a depth that reaches the relevant guides without collecting the whole site. Crawl a sample and inspect the resulting Markdown before choosing a larger scope.
Check whether headings, code explanations and important qualifications survived extraction. If script-heavy pages are nearly empty, compare the rendering settings and inspect any reported access failures. The FAQ explains that rendering alone does not bypass bot protection.
Review the SEO and redirect results for duplicate or obsolete pages, then export the permitted content format for the next processing stage. Keep URLs with the extracted text so later answers can be checked against their source. Refresh the collection when the documentation changes rather than treating a local archive as continuously current.
The FAQ distinguishes free crawling from Pro access, with page capacity affected by sign-in and subscription status. Full Markdown archive export is a premium feature. Confirm the required export and crawl size before assuming the free download covers a complete content-ingestion job.
The download page lists macOS and Linux routes. It describes native Windows desktop support as in progress and points Windows users to a WSL2 CLI route. Check the supported environment before selecting the interface.
Local execution uses the machine's resources and still depends on reachable source pages. Robots rules, server responses, rendering behavior and account limits can reduce what a crawl captures. Review missing pages and extraction quality before passing the results into a larger analysis or AI workflow.