crawlee
Tidy Interface for Reproducible Web Crawling
A tidy, pipe-friendly toolkit for reproducible web crawling and structured data collection, inspired by the architecture of the 'Crawlee' library. Provides a unified crawler with a deduplicating, resumable request queue, content-type aware handlers, structured storage backends and rich console logging via 'cli'. Supports crawling HTML pages, sitemaps, RSS and Atom feeds and PDF documents, with optional headless-browser rendering and helpers for retrieval-augmented generation.
Versions across snapshots
| Version | Repository | File | Size |
|---|---|---|---|
0.1.0 |
rolling linux/jammy R-4.5 | crawlee_0.1.0.tar.gz |
340.9 KiB |
0.1.0 |
rolling linux/noble R-4.5 | crawlee_0.1.0.tar.gz |
340.7 KiB |
0.1.0 |
rolling source/ R- | crawlee_0.1.0.tar.gz |
97.6 KiB |
0.1.0 |
latest linux/jammy R-4.5 | crawlee_0.1.0.tar.gz |
340.9 KiB |
0.1.0 |
latest linux/noble R-4.5 | crawlee_0.1.0.tar.gz |
340.7 KiB |
0.1.0 |
latest source/ R- | crawlee_0.1.0.tar.gz |
97.6 KiB |
0.1.0 |
2026-04-23 source/ R- | crawlee_0.1.0.tar.gz |
0 B |