Skip to content

Integration idea: anybrowse as fallback for Cloudflare-protected sites #3506

Description

@kc23go

Context

Crawlee is excellent for large-scale crawling but Cloudflare-protected sites are still a challenge -- even with Playwright, datacenter IPs get blocked quickly.

Idea

anybrowse uses residential Chrome (real IP, real browser) and could work as a fallback request handler for URLs that fail Cloudflare challenges.

Conceptually:

const router = createPlaywrightRouter();

router.addDefaultHandler(async ({ request, page }) => {
    // Try standard Playwright first
    const content = await page.content();
    if (content.includes('cf-browser-verification') || content.length < 500) {
        // Fallback to anybrowse for Cloudflare sites
        const r = await fetch('https://anybrowse.dev/scrape', {
            method: 'POST',
            headers: {'Content-Type': 'application/json'},
            body: JSON.stringify({url: request.url})
        });
        const data = await r.json();
        // Process data.markdown
    }
});

Not suggesting a native integration necessarily -- just sharing the pattern in case it's useful for Cloudflare use cases.

Docs: https://anybrowse.dev/docs

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    t-toolingIssues with this label are in the ownership of the tooling team.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions