Context
Crawlee is excellent for large-scale crawling but Cloudflare-protected sites are still a challenge -- even with Playwright, datacenter IPs get blocked quickly.
Idea
anybrowse uses residential Chrome (real IP, real browser) and could work as a fallback request handler for URLs that fail Cloudflare challenges.
Conceptually:
const router = createPlaywrightRouter();
router.addDefaultHandler(async ({ request, page }) => {
// Try standard Playwright first
const content = await page.content();
if (content.includes('cf-browser-verification') || content.length < 500) {
// Fallback to anybrowse for Cloudflare sites
const r = await fetch('https://anybrowse.dev/scrape', {
method: 'POST',
headers: {'Content-Type': 'application/json'},
body: JSON.stringify({url: request.url})
});
const data = await r.json();
// Process data.markdown
}
});
Not suggesting a native integration necessarily -- just sharing the pattern in case it's useful for Cloudflare use cases.
Docs: https://anybrowse.dev/docs
Context
Crawlee is excellent for large-scale crawling but Cloudflare-protected sites are still a challenge -- even with Playwright, datacenter IPs get blocked quickly.
Idea
anybrowse uses residential Chrome (real IP, real browser) and could work as a fallback request handler for URLs that fail Cloudflare challenges.
Conceptually:
Not suggesting a native integration necessarily -- just sharing the pattern in case it's useful for Cloudflare use cases.
Docs: https://anybrowse.dev/docs