How to scrape Glassdoor
Glassdoor is protected by Cloudflare. Automated requests often stop at a "Just a moment..." check or a 403 instead of job listings. StealthASF Ultra mode returns the real page: in our production test on 8 October 2026, Glassdoor's jobs page came back with HTTP 200 and its real title. This guide covers the first request, the job links to keep and the fields on a job page.
1. Use Ultra mode
Set engine to ultra. It is our strongest mode, built for the hardest protected sites, and the mode that passed Glassdoor in our tests. Ultra renders the page and supports extraction, but it does not run browser steps such as typing into a search box. You do not need to: a Glassdoor job search for a title and location has its own URL, which you can copy from your browser and send directly.
2. Send a first request with curl
Create an API key in your dashboard after verifying your email and set STEALTHASF_API_KEY. The request below fetches the jobs page with links extraction, so data lists every link with its text. Run it once to confirm access, then replace the URL with the search results page for the jobs you track.
curl --max-time 600 "https://stealthasf.com/v1/scrape" \
-H "x-api-key: $STEALTHASF_API_KEY" \
-H "content-type: application/json" \
--data-raw '{"url":"https://www.glassdoor.com/Job/index.htm","engine":"ultra","extract":"links"}'Keep the 600-second client timeout, since a rendered protected page takes longer than a plain request.
3. What comes back
The JSON response contains the target's status, the rendered html, the engine used and credits_charged. With links extraction, data is a table of text and url. Individual job pages on Glassdoor have /job-listing/ in their path, so those rows are the jobs on a results page. If the page text is a Cloudflare check instead of jobs, the request was blocked. Detected blocks return HTTP 422 and cost nothing.
4. Collect jobs in Python
This example uses only the Python standard library and stops on an API error, so a failed request never becomes a job record.
import json
import os
from urllib.error import HTTPError
from urllib.request import Request, urlopen
payload = {
"url": "https://www.glassdoor.com/Job/index.htm",
"engine": "ultra",
"extract": "links"
}
request = Request(
"https://stealthasf.com/v1/scrape",
data=json.dumps(payload).encode("utf-8"),
headers={
"x-api-key": os.environ["STEALTHASF_API_KEY"],
"content-type": "application/json",
},
method="POST",
)
try:
with urlopen(request, timeout=600) as response:
result = json.load(response)
except HTTPError as error:
detail = error.read().decode("utf-8")
raise SystemExit(f"API error {error.code}: {detail}")
print("Target status:", result["status"])
print("Credits:", result["credits_charged"])
print(result.get("data"))Then keep the job links and request each job page with the same payload. The helper reads JSON-LD from the page's HTML. Job pages that publish schema.org JobPosting data give you the title, company, location, posting date and salary range as JSON.
import re
# result is the listing-page response from the example above.
detail_urls = sorted({url for _text, url in result["data"]["rows"] if "/job-listing/" in url})
print(len(detail_urls), "detail pages")
def json_ld(html):
"""Every JSON-LD block on the page that parses as JSON."""
blocks = re.findall(r"<script[^>]*application/ld\+json[^>]*>(.*?)</script>", html, re.S | re.I)
found = []
for block in blocks:
try:
found.append(json.loads(block))
except json.JSONDecodeError:
pass
return found
# Send each detail URL with the same payload, then read its fields:
# for item in json_ld(detail["html"]): print(item.get("@type"), item.get("name"))5. Fields and cost
Typical Glassdoor job fields are job title, company, location, salary estimate, posting date, employment type and the description. Use the job URL as the record key and store the date you first saw each job, so you can tell new postings from old ones. Check every record for a title and a company before saving it.
An Ultra request costs 50 credits including the first 1 MB of transfer. Each additional MB adds 10 credits, rounded up to a whole credit, and a solved CAPTCHA adds 25 credits. One results page and 30 job pages at the base rate cost 1,550 credits. The Pro plan's 250,000 credits cover 5,000 Ultra requests and Scale covers 20,000. The exact cost of each page is in credits_charged. Blocked requests are never charged.
6. Run it on a schedule
Run your saved searches daily, request only job URLs you have not seen before, and you spend credits on new postings instead of repeats. Store the URL, engine, target status, job ID and credits with each record. Raise concurrency gradually within your plan limit, follow Retry-After on a 429 and cap retries. If a URL keeps returning 422, send support the job ID. If results depend on location, add a country field on Pro or Scale. The Cloudflare guide covers this protection in more depth and the API reference lists every field.