How to scrape Zillow
Zillow is protected by PerimeterX (HUMAN). A plain request usually gets a "Press & Hold" challenge instead of listings. StealthASF Ultra mode returns the real page: in our production test on 8 October 2026, the Zillow for-sale search page came back with HTTP 200 and its real title. This guide takes you from that search page to individual home listings and the fields on them.
1. Use Ultra mode
Set engine to ultra. It is our strongest mode, built for the hardest protected sites, and it is the mode that passed Zillow in our tests. Ultra renders the page and supports extraction, but it does not run browser steps such as clicks. That suits Zillow well, because search results, map areas and home details all have their own URLs you can request directly.
2. Request a search page with curl
Create an API key in your dashboard after verifying your email and set it as STEALTHASF_API_KEY. The request below fetches the for-sale search page and asks for links extraction, which returns every link on the page with its text. That gives you the listing URLs without writing a parser.
curl --max-time 600 "https://stealthasf.com/v1/scrape" \
-H "x-api-key: $STEALTHASF_API_KEY" \
-H "content-type: application/json" \
--data-raw '{"url":"https://www.zillow.com/homes/for_sale/","engine":"ultra","extract":"links"}'Replace the URL with the search you need, for example a city or ZIP code search copied from your browser's address bar. Keep the 600-second client timeout, since rendering a protected page takes longer than a plain fetch.
3. What comes back
The JSON response contains the target's status, the full rendered html, the engine that produced it and the credits_charged. With links extraction, data is a table with two columns, text and url. Home detail pages on Zillow have /homedetails/ in their path, so those rows are your listings. If the text of the page shows a challenge instead of homes, the request was blocked. A detected block returns HTTP 422 and is never charged.
4. Collect listings in Python
This version uses only the Python standard library and stops on an API error, so a failed request never becomes a record.
import json
import os
from urllib.error import HTTPError
from urllib.request import Request, urlopen
payload = {
"url": "https://www.zillow.com/homes/for_sale/",
"engine": "ultra",
"extract": "links"
}
request = Request(
"https://stealthasf.com/v1/scrape",
data=json.dumps(payload).encode("utf-8"),
headers={
"x-api-key": os.environ["STEALTHASF_API_KEY"],
"content-type": "application/json",
},
method="POST",
)
try:
with urlopen(request, timeout=600) as response:
result = json.load(response)
except HTTPError as error:
detail = error.read().decode("utf-8")
raise SystemExit(f"API error {error.code}: {detail}")
print("Target status:", result["status"])
print("Credits:", result["credits_charged"])
print(result.get("data"))Next, keep the home detail links and request each one with the same payload. The helper below also reads JSON-LD, the structured data block many sites embed for search engines, from a page's HTML. Where a detail page carries one, you get fields like the address without touching CSS selectors.
import re
# result is the listing-page response from the example above.
detail_urls = sorted({url for _text, url in result["data"]["rows"] if "/homedetails/" in url})
print(len(detail_urls), "detail pages")
def json_ld(html):
"""Every JSON-LD block on the page that parses as JSON."""
blocks = re.findall(r"<script[^>]*application/ld\+json[^>]*>(.*?)</script>", html, re.S | re.I)
found = []
for block in blocks:
try:
found.append(json.loads(block))
except json.JSONDecodeError:
pass
return found
# Send each detail URL with the same payload, then read its fields:
# for item in json_ld(detail["html"]): print(item.get("@type"), item.get("name"))5. Fields and cost
Most Zillow projects want the price, address, beds, baths, living area and listing status. Read them from JSON-LD where present and from the rendered html otherwise. Check every record for a price and an address before you save it. A listing without a price may be off market rather than broken, so keep that state instead of discarding the row.
An Ultra request costs 50 credits including the first 1 MB of transfer, plus 10 credits per additional MB, rounded up to a whole credit. A solved CAPTCHA adds 25 credits. One search page and 40 home pages at the base rate cost 2,050 credits, so the Pro plan's 250,000 credits cover about 120 such batches. Your real figure is in credits_charged on each response. Blocked requests cost nothing.
6. Run it on a schedule
Deduplicate home URLs before the detail pass, because the same home can appear in several searches. Store the URL, engine, target status, job ID and credits next to each record. Raise concurrency gradually within your plan limit, follow Retry-After on a 429 response and cap retries. If a URL keeps returning 422, send support the job ID. For more on this protection, read the PerimeterX guide, and see the API reference for every field.