How to scrape Tripadvisor
Tripadvisor is protected by DataDome. When it flags a request, you get a CAPTCHA page from DataDome instead of hotels, restaurants or reviews. StealthASF Ultra mode returns the real page: in our production test on 8 October 2026, tripadvisor.com came back with HTTP 200 and its real title. This guide covers the request, the links you want to keep and the fields on a review page.
1. Use Ultra mode
Set engine to ultra. It is our strongest mode, built for the hardest protected sites, and the mode that passed Tripadvisor in our tests. Ultra renders the page and supports extraction. It does not run browser steps such as clicks, and you do not need them here: every hotel, restaurant and attraction on Tripadvisor has its own page URL, and so does each city listing.
2. Send a first request with curl
Verify your email, create an API key in the dashboard and set STEALTHASF_API_KEY. The request uses links extraction, so data lists every link on the page with its text. Start with the home page to confirm access, then point the same request at the city or category page you need, such as the hotels page for a destination copied from your browser.
curl --max-time 600 "https://stealthasf.com/v1/scrape" \
-H "x-api-key: $STEALTHASF_API_KEY" \
-H "content-type: application/json" \
--data-raw '{"url":"https://www.tripadvisor.com/","engine":"ultra","extract":"links"}'Keep the 600-second client timeout from the example. Rendering a protected page takes longer than a plain fetch.
3. What comes back
The response contains the target's status, the rendered html, the engine used and credits_charged. With links extraction, data is a table of text and url. Tripadvisor detail pages have _Review- in their path: Hotel_Review-, Restaurant_Review- and Attraction_Review-. Those rows are the places you want. If the page text is a CAPTCHA prompt instead of listings, the request was blocked. A detected block returns HTTP 422 and is never charged.
4. Collect places in Python
The Python example uses only the standard library. It stops on an API error so a failed request never becomes a record.
import json
import os
from urllib.error import HTTPError
from urllib.request import Request, urlopen
payload = {
"url": "https://www.tripadvisor.com/",
"engine": "ultra",
"extract": "links"
}
request = Request(
"https://stealthasf.com/v1/scrape",
data=json.dumps(payload).encode("utf-8"),
headers={
"x-api-key": os.environ["STEALTHASF_API_KEY"],
"content-type": "application/json",
},
method="POST",
)
try:
with urlopen(request, timeout=600) as response:
result = json.load(response)
except HTTPError as error:
detail = error.read().decode("utf-8")
raise SystemExit(f"API error {error.code}: {detail}")
print("Target status:", result["status"])
print("Credits:", result["credits_charged"])
print(result.get("data"))Next, keep the review-page links and request each one. The helper reads JSON-LD from a page's HTML. Pages that embed schema.org data for hotels, restaurants or attractions give you the name, address and rating as JSON without any selectors.
import re
# result is the listing-page response from the example above.
detail_urls = sorted({url for _text, url in result["data"]["rows"] if "_Review-" in url})
print(len(detail_urls), "detail pages")
def json_ld(html):
"""Every JSON-LD block on the page that parses as JSON."""
blocks = re.findall(r"<script[^>]*application/ld\+json[^>]*>(.*?)</script>", html, re.S | re.I)
found = []
for block in blocks:
try:
found.append(json.loads(block))
except json.JSONDecodeError:
pass
return found
# Send each detail URL with the same payload, then read its fields:
# for item in json_ld(detail["html"]): print(item.get("@type"), item.get("name"))5. Fields and cost
Typical Tripadvisor fields are the place name, address, overall rating, number of reviews, price range, category and the ID in the URL (the number after -d). Use that ID as the record key, so the same hotel is one row across runs. Check every record for a name and a rating before saving it.
An Ultra request costs 50 credits including the first 1 MB of transfer. Each additional MB adds 10 credits, rounded up to a whole credit, and a solved CAPTCHA adds 25 credits. A city listing page and 30 place pages at the base rate cost 1,550 credits. The Hobby plan's 70,000 credits cover 1,400 Ultra requests, Pro covers 5,000 and Scale 20,000. Read credits_charged for the exact cost of each page. Blocked requests are never charged.
6. Run it on a schedule
Deduplicate place URLs before the detail pass, since popular places appear on several listing pages. Store the URL, place ID, engine, target status, job ID and credits with each record. Ratings and review counts move slowly, so a weekly run is usually enough for monitoring. Raise concurrency gradually within your plan limit, follow Retry-After on a 429 and cap retries. If a URL keeps returning 422, send support the job ID. The DataDome guide covers this protection in more depth, and the API reference lists every field.