Guides

How to scrape JavaScript-heavy sites

Some pages send useful HTML immediately. Others send a shell and fill it with data after JavaScript runs. If a plain request returns navigation and a loading message but none of the records you need, try rendering before rewriting your parser. StealthASF provides separate request modes so you can choose explicitly. Every example here targets example.com as a placeholder; it does not describe how example.com is built. Replace the URL and any selector with values from a page you are authorised to access.

1. Decide whether rendering is necessary

Start with engine: "http" and inspect the returned HTML. If it already includes the needed fields, rendering may add cost without adding information. Choose engine: "browser" when the content requires JavaScript. Use stealth when protection is also blocking access. Do not rely on auto to recognise a JavaScript shell: it escalates after a detected block, while an ordinary empty shell can be a successful HTTP response. Rendering and protected access solve different problems.

2. Render and extract text with curl

After verifying your account email, create an API key and set STEALTHASF_API_KEY in your environment. This request uses browser mode and asks for extracted text. Run it once, inspect the output and identify a phrase or record that appears only after the content loads. Keep this check small enough to inspect by hand.

curl --max-time 600 "https://stealthasf.com/v1/scrape" \
  -H "x-api-key: $STEALTHASF_API_KEY" \
  -H "content-type: application/json" \
  --data-raw '{"url":"https://example.com/","engine":"browser","extract":"text"}'

The API returns the target’s status separately from its own HTTP response status. Also inspect the data and credits_charged fields. A successful fetch can still return a loading message, an empty list or a page that needs an additional interaction.

3. Wait for a useful element in Python

When the needed content appears after initial rendering, add a wait_for step. The example uses css:main as a placeholder selector. For a real page, select an element that indicates the requested data is ready, such as a populated result row. Waiting for a container that exists before its data arrives may finish too early. A specific condition makes the wait easier to reason about than a large fixed delay.

import json
import os
from urllib.error import HTTPError
from urllib.request import Request, urlopen

payload = {
    "url": "https://example.com/",
    "engine": "browser",
    "extract": "text",
    "steps": [
        {
            "wait_for": "css:main"
        }
    ]
}
request = Request(
    "https://stealthasf.com/v1/scrape",
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "x-api-key": os.environ["STEALTHASF_API_KEY"],
        "content-type": "application/json",
    },
    method="POST",
)
try:
    with urlopen(request, timeout=600) as response:
        result = json.load(response)
except HTTPError as error:
    detail = error.read().decode("utf-8")
    raise SystemExit(f"API error {error.code}: {detail}")

print("Target status:", result["status"])
print("Credits:", result["credits_charged"])
print(result.get("data"))

This Python code uses the standard library. Handle connection errors and timeouts in your application as well as HTTP errors. If the chosen selector never appears, inspect the page and the returned error before lengthening waits. The page may have changed or may legitimately contain no records.

4. Choose an extraction strategy

Text extraction is a useful first check, but structured datasets usually need clearer fields. The API also supports extraction presets such as links, headings and tables. Page actions can collect data after a click or other interaction. Introduce one action at a time and verify its effect. If the response includes discovered_api, inspect those details as another source of information about how the page obtains data; discovery is optional and is not a promise that every page exposes a usable endpoint.

5. Measure the rendering cost

Browser mode starts at 5 credits, compared with 1 for a plain request and 25 for protected mode. Browser tiers include the first 1 MB of transfer and add 10 credits per extra MB, rounded up to a whole credit. MB means 1,048,576 bytes. A browser request transferring 2 MB without a solved CAPTCHA therefore costs 15 credits. A solved CAPTCHA adds 25 credits when applicable. Blocked requests are never charged. Larger pages and longer interactions can consume more transfer, so use the returned credit total for planning.

6. Turn a working example into a job

For a paginated view, define how you will recognise the last page before adding a loop. For a load-more button, track identifiers already collected and stop when another click adds no new records. Rendering alone does not mean every record has been loaded. Check the number of unique records against a small manual sample before treating the result as a complete export.

Keep a small set of representative URLs and expected fields as a repeatable check. Include pages with no results so you do not mistake a valid empty state for a failure. Record job IDs, target statuses and credits, and give retries a limit. Increase concurrency only after the data is correct; more simultaneous requests will not repair a selector. If regional content matters, country targeting is available on Pro and Scale. Read the API documentation for action syntax and pricing for the complete usage rules.