Site guides

How to scrape Walmart

Walmart runs PerimeterX (HUMAN) bot protection. Automated requests often land on a "Robot or human?" page with a "Press & Hold" button. StealthASF Ultra mode returns the real page: in our production test on 8 October 2026, the Walmart electronics category came back with HTTP 200 and its real title. Here is how to go from a category page to product URLs and product fields.

1. Use Ultra mode

Set engine to ultra, our strongest mode for the hardest protected sites and the one that passed Walmart in our tests. Ultra renders the page and supports extraction. It does not run browser steps such as clicks, which you do not need here: Walmart category pages, search results and product pages all have stable URLs, and the next page of a category is a URL too.

2. Request a category page with curl

Verify your email, create an API key in the dashboard and set STEALTHASF_API_KEY. This request fetches the electronics category and uses links extraction, so data lists every link on the page with its text, including the product links.

curl --max-time 600 "https://stealthasf.com/v1/scrape" \
  -H "x-api-key: $STEALTHASF_API_KEY" \
  -H "content-type: application/json" \
  --data-raw '{"url":"https://www.walmart.com/browse/electronics/3944","engine":"ultra","extract":"links"}'

Swap in any category or search URL from your browser. Keep the 600-second client timeout from the example, since a rendered protected page takes longer than a plain request.

3. What comes back

The response has the target's status, the rendered html, the engine used and credits_charged. With links extraction, data is a table of text and url. Walmart product pages have /ip/ in the path, followed by the product name and item ID, so filtering on /ip/ gives you the products on the page. If the text shows the "Robot or human?" page, you were blocked. Detected blocks return HTTP 422 and are not charged.

4. Collect products in Python

The Python example uses only the standard library and exits on an API error, so error bodies never reach your data.

import json
import os
from urllib.error import HTTPError
from urllib.request import Request, urlopen

payload = {
    "url": "https://www.walmart.com/browse/electronics/3944",
    "engine": "ultra",
    "extract": "links"
}
request = Request(
    "https://stealthasf.com/v1/scrape",
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "x-api-key": os.environ["STEALTHASF_API_KEY"],
        "content-type": "application/json",
    },
    method="POST",
)
try:
    with urlopen(request, timeout=600) as response:
        result = json.load(response)
except HTTPError as error:
    detail = error.read().decode("utf-8")
    raise SystemExit(f"API error {error.code}: {detail}")

print("Target status:", result["status"])
print("Credits:", result["credits_charged"])
print(result.get("data"))

Then keep the product links and request each one. The helper below filters the links and reads JSON-LD from a product page's HTML. Retail product pages commonly embed schema.org Product data this way, which gives you the name, brand, price and rating as clean JSON when it is there.

import re

# result is the listing-page response from the example above.
detail_urls = sorted({url for _text, url in result["data"]["rows"] if "/ip/" in url})
print(len(detail_urls), "detail pages")

def json_ld(html):
    """Every JSON-LD block on the page that parses as JSON."""
    blocks = re.findall(r"<script[^>]*application/ld\+json[^>]*>(.*?)</script>", html, re.S | re.I)
    found = []
    for block in blocks:
        try:
            found.append(json.loads(block))
        except json.JSONDecodeError:
            pass
    return found

# Send each detail URL with the same payload, then read its fields:
# for item in json_ld(detail["html"]): print(item.get("@type"), item.get("name"))

5. Fields and cost

Typical Walmart fields are product name, item ID (the number at the end of the URL), brand, current price, rating, review count, seller and availability. Take the item ID from the URL so every record has a stable key, even if the title changes. Validate that each record has a price before writing it. A product without a price may be out of stock, so store that state too.

An Ultra request costs 50 credits and includes the first 1 MB of transfer. Each additional MB adds 10 credits, rounded up to a whole credit, and a solved CAPTCHA adds 25. A category page plus 40 product pages at the base rate costs 2,050 credits. The Hobby plan's 70,000 credits cover 1,400 Ultra requests and Pro covers 5,000. Use credits_charged for the exact amount per page. Blocked requests are never charged.

6. Run it on a schedule

For price tracking, keep the list of product URLs and request only those on each run instead of re-crawling categories. Store the URL, item ID, engine, target status, job ID and credits with each price. Increase concurrency step by step within your plan limit, follow Retry-After on a 429 and cap retries. If a URL keeps returning 422, send support the job ID. The PerimeterX guide covers the protection in more depth and the API reference lists every field.