Guides

How to scrape Cloudflare-protected sites

A request can reach a website and still return a protection page instead of the content you expected. Treat access and extraction as separate checks: first obtain the page, then confirm it contains the fields your application needs. StealthASF supports requests through Cloudflare protection, but a particular page can still fail or change. This guide uses example.com only as a placeholder. It does not claim that example.com uses Cloudflare or needs protected mode. Replace it with a page you are authorised to access.

1. Choose the request mode

The API field engine selects the request mode. Use http when the initial response contains the data, browser for JavaScript rendering, and stealth for protected pages. For a known Cloudflare block, start with stealth so your test explicitly requests protected access. Alternatively, auto starts with HTTP and can step up after a detected block. Auto mode is not a general detector of missing JavaScript content: an unblocked HTML shell may still require an explicit browser request.

2. Send a small curl request

Create an API key in your dashboard after verifying your email. Set STEALTHASF_API_KEY in your local environment, then send the request below. Keep the key out of source control. Start with one URL and text extraction; adding actions before a basic request works makes failures harder to isolate.

curl --max-time 600 "https://stealthasf.com/v1/scrape" \
  -H "x-api-key: $STEALTHASF_API_KEY" \
  -H "content-type: application/json" \
  --data-raw '{"url":"https://example.com/","engine":"stealth","extract":"text"}'

The 600-second client timeout leaves room for a slower request; it is not a response-time guarantee. The endpoint accepts JSON and authenticates with the x-api-key header. Keep the API endpoint unchanged when replacing the placeholder target URL.

3. Separate API status from page status

Inspect both the HTTP status of the API response and the returned status, which describes the target page. A successful API response does not establish that you received the intended article, product or table. Check the returned text for an expected heading or identifier. A detected protection block returns HTTP 422 and is not charged. Authentication, credit and request errors need different fixes; repeating the same call does not fix an invalid key or an exhausted balance.

4. Move the request into Python

This example uses Python’s standard library. It stops on an HTTP error instead of treating an error object as extracted content. In your application, also catch connection failures and timeouts, record the job ID when available, and validate the shape of the returned data before writing it to your database.

import json
import os
from urllib.error import HTTPError
from urllib.request import Request, urlopen

payload = {
    "url": "https://example.com/",
    "engine": "stealth",
    "extract": "text"
}
request = Request(
    "https://stealthasf.com/v1/scrape",
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "x-api-key": os.environ["STEALTHASF_API_KEY"],
        "content-type": "application/json",
    },
    method="POST",
)
try:
    with urlopen(request, timeout=600) as response:
        result = json.load(response)
except HTTPError as error:
    detail = error.read().decode("utf-8")
    raise SystemExit(f"API error {error.code}: {detail}")

print("Target status:", result["status"])
print("Credits:", result["credits_charged"])
print(result.get("data"))

5. Budget for the complete request

A protected request starts at 25 credits. Browser tiers include the first 1 MB of transfer, then add 10 credits per additional MB, rounded up to a whole credit. MB means 1,048,576 bytes. A 3 MB protected request without a solved CAPTCHA costs 45 credits. A solved CAPTCHA adds 25 credits when applicable. The 200-credit trial therefore covers at most eight protected requests at the base rate. Blocked requests are never charged, but a client timeout alone does not prove that processing stopped.

6. Validate before increasing volume

Save the returned engine, credits_charged and job ID alongside your evaluation results. Test several page shapes, including missing pages and pages with no records, so your parser distinguishes empty content from a failed fetch. If the same URL keeps failing, pause and contact support with the job ID. Repeating it indefinitely gives you little diagnostic information even when the block itself costs no credits.

Before scheduling repeat runs, decide which missing fields should stop an import. A price without its currency, or an item without its identifier, may be unusable even though extraction returned text. Keep a small sample of expected values for comparison and alert on a sudden drop in complete records. This gives you a concrete signal to investigate when a layout changes without producing an API error.

Once a small sample works, increase concurrency gradually within your plan limit. Use bounded retries for temporary errors and respect Retry-After when a rate-limit response supplies it. Compare total credits against complete records collected, not just the number of HTTP 200 responses. See the API reference for error codes and the pricing page for plan allowances. Keep your initial request as a repeatable check when either your target page or extraction requirements change.