Cockroach Crawler is an open-source web crawler for AI agents. It maps permitted public sites, renders JavaScript when requested, and returns source-linked Markdown, JSON, or JSONL while enforcing explicit origin, redirect, robots, request, byte, depth, and time limits.
Static and JavaScript crawling
Sitemap and robots-aware discovery
Markdown, JSON, JSONL, CSS, XPath and PDF extraction
MCP, CLI, Node API, Docker API and Cloudflare Worker profiles
Source provenance, hashes, budgets and fail-closed boundaries
Map documentation sites for coding agents
Build cited retrieval datasets
Extract structured records from permitted pages
Render JavaScript-heavy public pages
Run bounded MCP and automation workflows

I built Cockroach Crawler because agents need useful web evidence without turning every crawl into an unbounded browser or proxy job. It maps permitted public pages, keeps source provenance, and makes origin, robots, redirect, request, byte, depth, and time limits explicit. I would love feedback from people building retrieval and research workflows.

I built Cockroach Crawler because agents need useful web evidence without turning every crawl into an unbounded browser or proxy job. It maps permitted public pages, keeps source provenance, and makes origin, robots, redirect, request, byte, depth, and time limits explicit. I would love feedback from people building retrieval and research workflows.
Find your next favorite product or submit your own. Made by @FalakDigital.
Copyright ©2025. All Rights Reserved