ActorStack.dev

How to scrape Wellfound startup jobs with salary and equity

A walkthrough of extracting startup job listings: how Wellfound's URL-path filtering constrains what you can ask for, which caps actually bound a run, and what arrives in each row.

By Oswaldo Carabano9 min read

Short answer

Wellfound job listings are extracted by role slug, optionally crossed with location slugs, because Wellfound filters by URL path rather than by query string. Each row carries 58 fields including the full job description — present on 100% of 2,229 measured jobs — with salary split into min, max, currency and period, equity parsed into a percentage range, geocoded coordinates on 93.3% of enriched rows, and the applicant tracking system behind the posting. Cap a run with `maxJobs` rather than with pages per role, because measured pagination ceilings run from 37 pages to 130 depending on the role.

Key points

  • Wellfound filters by URL path, not query string, so `remoteOnly` is ignored when `locations` is set — the combined form returns a 404 rather than an empty page.
  • One listing request returns about 28 unique jobs, and a role yields about 1,950 unique jobs on average.
  • The full job description ships in the listing row at no extra charge, so a run's cost is the row count times the row price with nothing hidden behind an enrichment step.
  • `maxJobs` is the cap that actually bounds a run. `maxPagesPerRole` is a safety net, because a fixed page count means a whole corpus for one role and a third of it for another.
  • A residential proxy is effectively required: at the same concurrency a datacentre IP returned 24 of 36 requests as HTTP 429 where residential returned 48 of 48 as HTTP 200.
On this page8 sections

Wellfound is what AngelList Talent became: the job side of AngelList, where startups post vacancies with a pitch, a size band and — often — a salary range and an equity range. Getting that out is mostly a matter of understanding how the site addresses its own listings.

Wellfound filters by path, and that constrains the input

Filters live in the URL path rather than in a query string. A role is a path segment, a location is a path segment, and the combinations that exist are the ones the site has routes for. That is not a detail: it is why one filter combination cannot be requested at all.

Step 1 — choose roles or categories

Wellfound publishes 518 role slugs across 13 categories. Name them individually in roles, or name a whole category in roleCategories and let the Actor expand it. Engineering alone holds 227 roles, so a category run is a long one.

A role yields about 1,950 unique jobs on average, and one listing request returns about 28 of them.

Step 2 — add locations, or remote, but not both

There are 306 location slugs, and each role is crossed with each location — so three roles and four locations is twelve sweeps, not seven. Leave locations empty for no location filter, which is also the only way remoteOnly takes effect.

Step 3 — cap the run with the right knob

maxJobs is the cap that matters, because delivered rows are what you pay for. maxPagesPerRole looks like the equivalent and is not: measured ceilings run from 37 pages for office-manager to 130 for engineer, so one value gives you a whole corpus in one role and a third of it in another — the measurement.

maxConcurrency defaults to 6 and caps at 16, which is the highest level measured without a block — where the rate-limit ceiling is.

Step 4 — run it

input.json
{
  "roles": ["software-engineer", "product-manager"],
  "locations": ["san-francisco", "new-york"],
  "maxJobs": 500,
  "enrichFromJobPage": true,
  "emitCompanies": true,
  "maxCacheAgeDays": 1
}

enrichFromJobPage adds structured salary, coordinates, benefits and the company's geocoded headquarters at one extra request per job. The full description arrives without it.

Reading the output

58 fields per job. The description is at 100% of 2,229 measured jobs, location names at 90.5%, the raw compensation string at 79.5%, parsed salary at 86.7% of enriched rows and coordinates at 93.3% of enriched rows. The full field reference has all of them.

Two fields are published as ranges rather than averages, because their category spread is wider than a mean can describe: ats_source at 21–83% and years_experience_min at 12.8–75% — why the average is wrong.

What a run costs

$3.00 per 1,000 jobs, $1.50 per 1,000 deduplicated companies, $0.01 per 1,000 starts. Five hundred jobs with companies is about $1.60. The description and the detail-page enrichment are in the job price and are never billed separately, so a run's cost is the row count times the row price with nothing hidden.

Four mistakes that waste a run

Setting remote and a location together. The site has no route for it, so the remote flag is dropped. Run them as separate jobs if you need both.

Capping on pages instead of rows. The same page cap means different coverage in every role, and nothing in the output says which you got.

Running from a datacentre IP. 24 of 36 requests came back as HTTP 429 where residential returned 48 of 48 as 200.

Treating `scraped_at` as freshness. Cloudflare's cache sat in front of job detail pages at up to 23.4 hours, and data_age_hours counts it — why the two differ.

Frequently asked questions

Can I filter for remote jobs in a specific city?
No, and the reason is Wellfound's own routing: it filters by URL path rather than query string, and the combined remote-plus-location form returns a 404. `remoteOnly` is therefore ignored whenever `locations` is set, rather than silently producing an empty run.
How many jobs will one role return?
About 1,950 unique jobs on average, with one listing request returning about 28 of them. There is no site-wide pagination limit — each role paginates until its own corpus runs out and then redirects.
Do I need to enrich from the job page?
Only if you need structured salary, coordinates, benefits or the company's geocoded headquarters. The full description is already in the listing row without it. Enrichment costs one extra request per job and is on by default.
What concurrency is safe?
16 is the highest value measured without a block, and 24 triggers a 429 with roughly a five-minute cooldown. The default is 6. Raising it past the ceiling makes a run slower rather than faster, because rate-limited requests still cost time.

Sources

Every URL below was requested and returned a page on the date shown.

  1. Operator claimchecked 9 Sept 2026
    Wellfound Jobs Scraper — Actor README and input schemaActorStack / Apify Store
  2. Site declarationchecked 9 Sept 2026
    wellfound.com/robots.txtWellfound
  3. Operator claimchecked 9 Sept 2026
    Wellfound Terms of ServiceWellfound
A laptop screen showing a plain text-mode terminal with a command prompt.
WellfoundExplainer

The Wellfound API question

There is no documented public endpoint for startup job listings. What robots.txt permits, what Cloudflare sits in front of, and which surface is actually readable.

6 min
A white measuring tape curving across a dark background, showing the numbers 15 to 45.
WellfoundMeasured

The ATS spread

`ats_source` names the applicant tracking system behind a posting. Its fill rate varies by 62 percentage points across role categories, which is why the average is not published.

6 min
Racks of network equipment in a dimly lit server room, lit blue by their indicators.
WellfoundMeasured

Cache age through a CDN

A row can be fetched seconds ago and still be a day old. Cloudflare's cache sat in front of Wellfound job pages at up to 23.4 hours, and `data_age_hours` counts it.

6 min