Site guides

How to scrape Home Depot

Home Depot is protected by Akamai Bot Manager. A plain request typically gets a 403 or an error page instead of products. StealthASF gets the real page in more than one mode: in our production tests on 8 October 2026, the Home Depot tools category came back with HTTP 200 and its real title in both stealth mode and Ultra mode. That means you can use auto mode and pay the protected rate of 25 credits instead of the Ultra rate of 50.

1. Use auto mode

Set engine to auto. Auto sends a plain request first, and when it detects a block it steps up to protected mode (stealth). It also remembers sites that needed protection, so once Home Depot has needed it, later requests start in protected mode straight away. You pay only for the mode that produced the page. If you prefer to skip the plain attempt from the first request, set engine to stealth.

2. Request a category page with curl

Verify your email, create an API key in the dashboard and set STEALTHASF_API_KEY. This request fetches the tools category with links extraction, so data lists every link on the page with its text, product links included.

curl --max-time 600 "https://stealthasf.com/v1/scrape" \
  -H "x-api-key: $STEALTHASF_API_KEY" \
  -H "content-type: application/json" \
  --data-raw '{"url":"https://www.homedepot.com/b/Tools/N-5yc1vZc1xy","engine":"auto","extract":"links"}'

Replace the URL with any Home Depot category or search page from your browser. Keep the 600-second client timeout, which leaves room for the step up to protected mode.

3. What comes back

The response contains the target's status, the page html, the engine that produced it and credits_charged. The engine field matters here: it shows whether the page came from the plain request or from protected mode, and it decides the charge. A protected request can escalate to Ultra when a site needs it. With links extraction, data is a table of text and url. Home Depot product pages have /p/ in the path, ending in the product ID, so filtering on /p/ gives you the products. A detected block returns HTTP 422 and is never charged.

4. Collect products in Python

The Python example uses only the standard library and stops on an API error, so a failed request never becomes a record.

import json
import os
from urllib.error import HTTPError
from urllib.request import Request, urlopen

payload = {
    "url": "https://www.homedepot.com/b/Tools/N-5yc1vZc1xy",
    "engine": "auto",
    "extract": "links"
}
request = Request(
    "https://stealthasf.com/v1/scrape",
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "x-api-key": os.environ["STEALTHASF_API_KEY"],
        "content-type": "application/json",
    },
    method="POST",
)
try:
    with urlopen(request, timeout=600) as response:
        result = json.load(response)
except HTTPError as error:
    detail = error.read().decode("utf-8")
    raise SystemExit(f"API error {error.code}: {detail}")

print("Target status:", result["status"])
print("Credits:", result["credits_charged"])
print(result.get("data"))

Then keep the product links and request each one with the same payload. The helper reads JSON-LD from a product page's HTML. Retail product pages commonly embed schema.org Product data, which gives you name, brand, price and rating as JSON when it is present.

import re

# result is the listing-page response from the example above.
detail_urls = sorted({url for _text, url in result["data"]["rows"] if "/p/" in url})
print(len(detail_urls), "detail pages")

def json_ld(html):
    """Every JSON-LD block on the page that parses as JSON."""
    blocks = re.findall(r"<script[^>]*application/ld\+json[^>]*>(.*?)</script>", html, re.S | re.I)
    found = []
    for block in blocks:
        try:
            found.append(json.loads(block))
        except json.JSONDecodeError:
            pass
    return found

# Send each detail URL with the same payload, then read its fields:
# for item in json_ld(detail["html"]): print(item.get("@type"), item.get("name"))

5. Fields and cost

Typical Home Depot fields are product name, brand, model number, the product ID at the end of the URL, price, rating, review count and availability. Use the product ID as the record key. Prices and stock can depend on the selected store, so record which location your data reflects and keep it consistent between runs.

A protected request costs 25 credits and includes the first 1 MB of transfer. Each additional MB adds 10 credits, rounded up to a whole credit, and a solved CAPTCHA adds 25 credits. A category page and 40 product pages at the protected base rate cost 1,025 credits. The Hobby plan's 70,000 credits cover 2,800 protected requests, Pro covers 10,000 and Scale 40,000. Check credits_charged on each response for the exact amount. Blocked requests are never charged.

6. Run it on a schedule

For price and stock tracking, build the product list once and request only those product pages on each run. Store the URL, product ID, engine, target status, job ID and credits with every record. Raise concurrency gradually within your plan limit, follow Retry-After on a 429 and cap retries. If a URL keeps returning 422, send support the job ID. The Akamai guide covers this protection in more depth and the API reference lists every field.