Replies: 1 comment
|
The pattern's solid, and it's worth splitting into two flavours, because they fail differently once you wire them into Crawlee. 1. Ship the whole render out (what anybrowse does here): hand the URL to a residential-browser service, get HTML/markdown back. Simple, but you leave Crawlee's session pool, request queue and proxy handling behind for that URL, and you pay per full page render. Fine for a handful of stubborn URLs, expensive if half your queue sits behind Cloudflare. 2. Solve just the token, inject it, stay in Crawlee. For Turnstile you don't actually need the page rendered somewhere else. The thing blocking you is one value: One check that saves time before you reach for either: log the real status. A live Turnstile challenge is a 403 with the widget present. A For the token route I keep a small Crawlee wrapper that does the inject-and-continue step: https://github.com/CircuitSavage/crawlee-turnstile (it calls Peak's solve API under the hood; disclosure: I work on Peak). Peak is ~$0.90/1k successful solves, down to $0.35 at volume, and you pay per success, so the ~1k free solves are enough to check whether the token route even works on your target before you commit. If it turns out the block is IP-based rather than a real challenge, save the money and fix the proxies first. |
Uh oh!
There was an error while loading. Please reload this page.
Context
Crawlee is excellent for large-scale crawling but Cloudflare-protected sites are still a challenge -- even with Playwright, datacenter IPs get blocked quickly.
Idea
anybrowse uses residential Chrome (real IP, real browser) and could work as a fallback request handler for URLs that fail Cloudflare challenges.
Conceptually:
Not suggesting a native integration necessarily -- just sharing the pattern in case it's useful for Cloudflare use cases.
Docs: https://anybrowse.dev/docs
All reactions