Launch
Cockroach Crawler
Visit
Example Image

Cockroach Crawler

AI web crawler for governed agents

Visit

Cockroach Crawler is an open-source web crawler for AI agents. It maps permitted public sites, renders JavaScript when requested, and returns source-linked Markdown, JSON, or JSONL while enforcing explicit origin, redirect, robots, request, byte, depth, and time limits.

Example Image
Example Image
Example Image
Example Image

Features

Static and JavaScript crawling

Sitemap and robots-aware discovery

Markdown, JSON, JSONL, CSS, XPath and PDF extraction

MCP, CLI, Node API, Docker API and Cloudflare Worker profiles

Source provenance, hashes, budgets and fail-closed boundaries

Use Cases

Map documentation sites for coding agents

Build cited retrieval datasets

Extract structured records from permitted pages

Render JavaScript-heavy public pages

Run bounded MCP and automation workflows

Comments

I built Cockroach Crawler because agents need useful web evidence without turning every crawl into an unbounded browser or proxy job. It maps permitted public pages, keeps source provenance, and makes origin, robots, redirect, request, byte, depth, and time limits explicit. I would love feedback from people building retrieval and research workflows.

Premium Products

Comments

I built Cockroach Crawler because agents need useful web evidence without turning every crawl into an unbounded browser or proxy job. It maps permitted public pages, keeps source provenance, and makes origin, robots, redirect, request, byte, depth, and time limits explicit. I would love feedback from people building retrieval and research workflows.

Premium Products