How to scrape DataDome-protected sites
When a page is protected by DataDome, a plain request may return a challenge instead of useful content. Start by checking what was returned, rather than assuming an empty extraction means your selector is wrong. StealthASF supports DataDome-protected requests and returns usage information so you can evaluate both results and costs. The examples below use example.com as a placeholder, not as a claim about its protection. Substitute a URL you are authorised to access and choose one specific field to validate first.
1. Request protected access explicitly
Use engine: "stealth" for the initial test of a known protected page. Use http for ordinary pages whose data is already present in the response, and browser for JavaScript content without a protection block. The auto mode can escalate from HTTP after a detected block, but an explicit setting makes an evaluation easier to reproduce. Keep the URL and extraction setting fixed while changing modes, so you can tell which change affected the result.
2. Test one URL with curl
Verify your account email, create an API key in the dashboard, and put it in the STEALTHASF_API_KEY environment variable. This request asks for text extraction. It deliberately omits country targeting and page actions so that the first result answers a simple question: can this request retrieve the expected page content?
curl --max-time 600 "https://stealthasf.com/v1/scrape" \
-H "x-api-key: $STEALTHASF_API_KEY" \
-H "content-type: application/json" \
--data-raw '{"url":"https://example.com/","engine":"stealth","extract":"text"}'Inspect the API response before launching a loop. Check the target status, returned text and credits_charged. A page title by itself may be insufficient: look for an expected record identifier or another field specific to the content you requested.
3. Handle failures separately in Python
The Python example uses only the standard library and reads the key from your environment. An API HTTP error stops the script. Add application-specific validation after the request succeeds, then save the data only if that validation passes. This prevents an error response or an unrelated page from silently becoming a record in your dataset.
import json
import os
from urllib.error import HTTPError
from urllib.request import Request, urlopen
payload = {
"url": "https://example.com/",
"engine": "stealth",
"extract": "text"
}
request = Request(
"https://stealthasf.com/v1/scrape",
data=json.dumps(payload).encode("utf-8"),
headers={
"x-api-key": os.environ["STEALTHASF_API_KEY"],
"content-type": "application/json",
},
method="POST",
)
try:
with urlopen(request, timeout=600) as response:
result = json.load(response)
except HTTPError as error:
detail = error.read().decode("utf-8")
raise SystemExit(f"API error {error.code}: {detail}")
print("Target status:", result["status"])
print("Credits:", result["credits_charged"])
print(result.get("data"))4. Add location only when needed
If the data depends on country, add a country field with a two-letter country code, such as "country": "US". Country targeting requires Pro or Scale; Free and Hobby return HTTP 403 with plan_feature for that option. Keep the country constant across your evaluation. Changing location and request mode together can produce different content and make it difficult to distinguish a successful extraction from a regional variation.
5. Account for bandwidth and challenges
Protected requests start at 25 credits. The first 1 MB of transfer is included; additional transfer on browser tiers costs 10 credits per MB, rounded up to a whole credit. A solved CAPTCHA adds another 25 credits when applicable. For example, a protected request with 2 MB transferred costs 35 credits without a solved CAPTCHA, or 60 with one. Here, MB means 1,048,576 bytes. Read actual usage from the response rather than assuming every successful page costs the base amount.
A detected block returns HTTP 422 and is never charged. That does not make repeated blocked attempts useful. If multiple attempts return the same failure, stop the test and send support the job ID and a description of the expected content. Do not include your API key. For an HTTP 429 response, reduce concurrency and follow the retry delay when supplied. Credit exhaustion and authentication errors require changes to the account or request before retrying.
6. Make the evaluation reproducible
Separate access checks from extraction changes. If the response contains the expected page but your table is empty, inspect the extraction preset or selector before switching request modes again. If the response is a challenge, work on access first. Keep one known request as your baseline and change only one parameter between attempts. This avoids spending credits on a series of tests whose differences you cannot explain.
Record the target URL, requested mode, country, timestamp, job ID and credits used for each test. Compare a few page types and include empty results as well as populated ones. Define success as usable records with the fields you require, then calculate credits per usable record. This captures extraction problems that a simple count of successful responses would miss. Use the API reference to interpret errors and the plan comparison to set a budget before increasing volume.