Wellfound (AngelList) Startup Jobs & Companies
The description is in the row, not behind a second charge.
Startup job listings from Wellfound, formerly AngelList Talent, with the full job description in every row rather than behind a separate enrichment charge. 58 fields: salary split into min, max, currency and period, equity as a parsed percentage range, geocoded locations, and the applicant tracking system each company hires through.
oswaldocarabano/wellfound-jobs-scraper
{
"roles": ["software-engineer", "product-manager"],
"locations": ["san-francisco", "new-york"],
"maxJobs": 500,
"enrichFromJobPage": true,
"emitCompanies": true,
"maxCacheAgeDays": 1
}- Version
- v0.1.13
- Memory
- 512 MB
- Browser
- none
- Proxy
- Residential, strongly recommended
Short answer
The Wellfound Jobs Actor extracts startup job listings from wellfound.com with 58 fields per row, including the full job description in markdown and HTML — present on 100% of 2,229 measured jobs across 25 roles and all 13 categories. Salary is split into min, max, currency and period, filled on 86.7% of enriched rows; equity is parsed into a percentage range; latitude and longitude are on 93.3% of enriched rows; and `ats_source` names the applicant tracking system behind the posting on 21% to 83% of rows depending on the role category. Pricing is $3.00 per 1,000 jobs with descriptions and enrichment included, and $1.50 per 1,000 deduplicated companies.
Key points
- The full job description is in every row, in markdown and HTML, and it is never a separate charge. It was present on 100% of 2,229 jobs measured across 25 roles and all 13 categories.
- Salary is parsed into `salary_min`, `salary_max`, `salary_currency` and `salary_period` on 86.7% of enriched rows, and equity into `equity_min_pct` and `equity_max_pct` — not left as a display string.
- `ats_source` names the applicant tracking system behind the posting — Greenhouse, Ashby, Lever or Workable — on 83% of rows in HR and recruiting and 21% in Marketing. The 62-point spread is published as a range because the average would mislead.
- `yearsExperienceMax` is deliberately not shipped: it was filled on 2 of 2,229 rows, so the column would have been nulls.
- Companies are deduplicated across the run: a company posting 56 jobs is charged once, and the company is attached inline to every job so no second run is needed.
- `data_age_hours` includes Cloudflare's own cache age, measured at up to 23.4 hours on job detail pages — so it reports how old the data is, not how recently the Actor asked for it.
What it does
Startup job listings from Wellfound, formerly AngelList Talent, with the full job description in every row rather than behind a separate enrichment charge. 58 fields: salary split into min, max, currency and period, equity as a parsed percentage range, geocoded locations, and the applicant tracking system each company hires through.
{
"job_id": "3184920",
"title": "Senior Backend Engineer",
"description": "## About the role\nWe are looking for…",
"job_type": "full-time",
"is_remote": true,
"remote_type": "remote-only",
"posted_at": "2026-08-19",
"compensation_raw": "$140k – $180k • 0.1% – 0.4%",
"salary_min": 140000,
"salary_max": 180000,
"salary_currency": "USD",
"salary_period": "year",
"equity_offered": true,
"equity_min_pct": 0.1,
"equity_max_pct": 0.4,
"location_names": ["San Francisco"],
"location_locality": "San Francisco",
"location_region": "California",
"location_country": "US",
"latitude": 37.7749,
"longitude": -122.4194,
"years_experience_min": 5,
"ats_source": "greenhouse",
"company_name": "Example Labs",
"company_size_min": 11,
"company_size_max": 50,
"company_website": "https://example.com",
"from_cache": false,
"data_age_hours": 0
}Why this one
The description is not an upsell
The pattern in this category is to sell a cheap listing row and then charge again to fetch the page that has the text on it, which makes a run's real cost unknowable until it finishes. Here the description ships in the job price, in markdown and HTML, at a measured 100% fill rate — so multiplying the price by the number of jobs gives you the actual bill.
A fill rate range where the average lies
`ats_source` is filled on 83% of HR and recruiting postings and 21% of Marketing ones. Publishing the mean would tell a marketing recruiter to expect something the data will not give them. The range is what gets published, by category, because a 62-point spread is not noise around an average — it is the finding.
A column removed rather than shipped empty
Wellfound exposes a maximum-years-of-experience value and it is filled on 0.1% of rows: 2 of 2,229. A column that is null 99.9% of the time is not data, it is the appearance of data, and it was cut. The minimum, at 31.6% overall and 12.8% to 75% by category, is shipped with its range.
Freshness that counts somebody else's cache
Every row reports `from_cache` and `data_age_hours`, and that figure includes Cloudflare's cache age in front of Wellfound — measured at up to 23.4 hours on job detail pages. A freshness field that only counted the Actor's own cache would say a row was minutes old when the underlying page was a day old. Set `maxCacheAgeDays` to 0 to force a fresh fetch.
No people, and the reason is on the page
Company profile pages carry named founders and employees, and Wellfound puts them behind a Cloudflare challenge. This Actor does not scrape them, so there are no funding rounds, investors, perks or team members in the output. Job descriptions are delivered exactly as the company wrote them and nothing is extracted or derived from them.
Use cases
- Track startup hiring demand by role and location, with salary bands you can aggregate rather than parse.
- Build a lead list of companies actively hiring, deduplicated, with size band, industries, website and geocoded headquarters.
- Segment prospects by the applicant tracking system they hire through, where `ats_source` is well filled for the category.
- Compare advertised compensation across cities using the geocoded location rather than the free-text one.
- Measure remote-work supply using `is_remote`, `remote_type` and the accepted applicant countries.
Input
Every field has a default, and the defaults are deliberately small so a first run is cheap enough to inspect before you commit to a sweep. This table mirrors the Actor's own input schema field for field.
| Field | Default | What it does |
|---|---|---|
rolesstring[] | [] | Job rolesRole slugs such as `software-engineer` or `product-manager`. Wellfound publishes 518 role slugs across 13 categories. Leave empty to use `roleCategories` instead. |
roleCategoriesstring[] | [] | Role categoriesScrape every role in a category rather than naming roles one by one. Engineering alone holds 227 roles, so expect long runs. |
locationsstring[] | [] | LocationsLocation slugs such as `san-francisco` or `london`, of which Wellfound publishes 306. Each role is crossed with each location. |
remoteOnlyboolean | false | Remote jobs onlyIgnored when `locations` is set: Wellfound returns a 404 for the combined form rather than an empty result. |
maxJobsinteger | 500 | Maximum jobsA hard cap on delivered rows; there is no unlimited mode. One listing request returns about 28 unique jobs, and a role yields about 1,950 on average. |
maxPagesPerRoleinteger | 40 | Maximum pages per roleA safety cap only. Measured ceilings run from 37 pages for office-manager to 130 for engineer, so one fixed value delivers a role's whole corpus or a third of it depending on the role. Cap with `maxJobs` instead. |
enrichFromJobPageboolean | true | Enrich from the job detail pageAdds structured salary, latitude and longitude, benefits, the company website and its geocoded headquarters, at one extra request per job. The full description is already included without it. |
emitCompaniesboolean | false | Also emit a companies datasetWrites a deduplicated companies dataset. A company with 56 jobs appears once and is charged once. |
maxCacheAgeDaysinteger | 1 | Maximum cache age (days)Serves cached rows up to this age. 0 forces a fresh fetch. Every row reports `from_cache` and `data_age_hours`, and that figure includes Cloudflare's own cache age. |
maxConcurrencyinteger | 6 | Maximum concurrencypersonal data16 is the highest value measured without a block; 24 triggers a 429 with a roughly five-minute cooldown. Raising it does not make a run faster once you are rate-limited. |
proxyConfigurationobject | RESIDENTIAL | ProxyResidential is strongly recommended and is the default. Measured: 24 of 36 requests returned 429 from a datacentre IP where 48 of 48 returned 200 from a residential one. |
Output and fill rates
A field being in the schema is not the same as it having a value. The percentages below were counted on real runs; the sample sizes are in Measurements. Anything not listed here is not promised.
| Field | Filled | Meaning |
|---|---|---|
descriptionstring | 100% | The full job description in markdown, with the HTML alongside. Included in the job price, never charged separately. |
titlestring | 100% | With `job_type`, `is_remote`, `remote_type`, `posted_at` and `direct_apply`. |
location_namesstring[] | 90.5% | The locations as Wellfound lists them. Minimum observed across categories was 83.7%. |
compensation_rawstring | 79.5% | The compensation string as displayed. Minimum observed across categories was 60.3%. |
salary_minnumber | 86.7% | With `salary_max`, `salary_currency` and `salary_period`. The rate is of enriched rows, so it depends on `enrichFromJobPage`. |
equity_offeredboolean | not measured | With `equity_min_pct` and `equity_max_pct`, parsed into a percentage range rather than left as a display string. |
latitudenumber | 93.3% | With `longitude`. Of enriched rows. |
location_localitystring | not measured | With `location_region` and `location_country`, split out of the display name. |
applicant_location_requirementsstring[] | not measured | Countries the company will accept applicants from. |
years_experience_mininteger | 31.6% | 12.8% to 75% depending on role category. The maximum is not shipped: it was filled on 2 of 2,229 rows. |
ats_sourcestring | not measured | Greenhouse, Ashby, Lever or Workable. 83% in HR and recruiting, 21% in Marketing — the range is the finding, not noise around an average. |
company_namestring | 100% | With `company_size_min`/`_max`, `company_pitch`, `company_badges`, `company_website`, `company_industries` and the geocoded headquarters. Attached inline, so no second run is needed. |
from_cacheboolean | 100% | With `data_age_hours`, which includes Cloudflare's cache age in front of Wellfound — up to 23.4 hours measured on job detail pages. |
scraped_atdatetime | 100% | With `source_surface`, which records which listing surface the row came from. |
Every key is always present. A field that exists but is empty comes back as explicit null, so a parser never has to guess.
Datasets
Different record types go to different datasets, so the main table never carries columns that are blank on most rows.
defaultOne row per job, 58 fields, deduplicated by job id across the run, with the company attached inline.billedcompaniesOne row per hiring company, deduplicated across the run. Optional, off by default.billedError rowsRequests that failed, with the reason.never billed
Pricing
Pay per delivered result. Charges are applied as each row is produced rather than in a lump at the end, so an aborted run bills only for what it actually gave you.
| Event | Price | Notes |
|---|---|---|
actor-startActor start | $0.00001 | $0.01 per 1,000 starts. A run that finds nothing costs you nothing. |
jobJob | $0.003 | One startup job with 58 fields, deduplicated by job id across the run. The full description and the detail-page enrichment are included at this price, never billed separately. Error rows are never charged. |
companyCompany | $0.0015 | One hiring company, deduplicated across the run: a company posting 56 jobs is charged once. Optional, off by default. |
Measurements
Each figure is shown with the method that produced it. A benchmark without a method is a marketing claim wearing a number's clothes.
Fill-rate sample
2,229 unique jobs
Across 25 roles and all 13 categories. Every rate on this page comes from that sample, and the ones that vary by category are published as ranges rather than means.
Description present
100% of 2,229 jobs
In markdown and HTML, in the listing row itself. This is why the description is not a separate charge.
ATS source by category
83% HR and recruiting, 21% Marketing
A 62-percentage-point spread across categories. The mean would tell a marketing recruiter to expect four times what the data gives them.
Cloudflare cache age
up to 23.4 hours on job detail pages
Measured in front of Wellfound and included in `data_age_hours`, so the field reports the age of the data rather than the age of the request.
Residential versus datacentre
48 of 48 HTTP 200 against 24 of 36 HTTP 429
Same concurrency, same targets, two proxy types. This is why residential is the default rather than a suggestion.
Rate-limit ceiling
16 concurrent safe, 24 blocked
24 triggers a 429 with roughly a five-minute cooldown, so the input caps at 16.
Corpus per role
~1,950 unique jobs per role, ~28 per request
Pagination ceilings measured per role: engineer 130 pages, software-engineer 91, office-manager 37. There is no site-wide limit.
What it will not do
Stated plainly so you can judge fit before spending anything.
- No company profile data: no funding rounds, investors, perks or team. Wellfound protects `/company/*` behind a Cloudflare challenge, and that page also exposes named founders and employees.
- `ats_source` ranges from 21% to 83% by role category, so its usefulness depends entirely on which category you are pulling.
- `years_experience_min` is filled on 31.6% of rows overall, between 12.8% and 75% depending on category.
- `remoteOnly` is ignored when `locations` is set, because Wellfound returns a 404 for the combined form. It filters by URL path, not by query string.
- There is no site-wide pagination limit: each role paginates until its corpus runs out, and measured ceilings vary from 37 pages for office-manager to 130 for engineer. Cap a run with `maxJobs` rather than with pages per role.
- Wellfound rate-limits per IP. 16 concurrent requests is the highest value measured without a block and 24 triggers a 429 with roughly a five-minute cooldown, so raising concurrency past that makes a run slower, not faster.
- A residential proxy is effectively required: the same concurrency that returned 24 of 36 requests as HTTP 429 from a datacentre IP returned 48 of 48 as HTTP 200 from a residential one.
Privacy
- No personal data is extracted. Job descriptions are delivered exactly as the company wrote them, and nothing is derived from them.
- Company profile pages are not scraped, and one reason is that they expose named founders and employees.
- The output describes companies and their published vacancies, which are business facts published for candidates to find.
- Removal requests: privacy@actorstack.dev
See also the data removal process.
Frequently asked questions
Is the full job description included?
Is Wellfound the same as AngelList?
Do I get salary and equity as numbers?
What is `ats_source` and how often is it there?
Does it return company profiles, funding or investors?
Do I need a residential proxy?
How do I cap what a run costs?
Guides for this Actor
- Scrape Wellfound jobsA walkthrough of extracting startup job listings: how Wellfound's URL-path filtering constrains what you can ask for, which caps actually bound a run, and what arrives in each row.
- The Wellfound API questionThere is no documented public endpoint for startup job listings. What robots.txt permits, what Cloudflare sits in front of, and which surface is actually readable.
- The ATS spread`ats_source` names the applicant tracking system behind a posting. Its fill rate varies by 62 percentage points across role categories, which is why the average is not published.
- The column not shippedWellfound exposes a maximum years-of-experience value. It was filled on 2 of 2,229 jobs. Shipping it would have added a field that looks like data and is not.
- Salary and equity parsingCompensation on Wellfound is a display string. What it takes to split it into minimum, maximum, currency, period and an equity range — and why the original string still ships.
- Cache age through a CDNA row can be fetched seconds ago and still be a day old. Cloudflare's cache sat in front of Wellfound job pages at up to 23.4 hours, and `data_age_hours` counts it.
- Rate limits and proxiesWellfound rate-limits per IP, and the difference between a datacentre address and a residential one was 24 failures out of 36 against zero out of 48.
- Pagination ceilingsEach role paginates until its own corpus runs out. Measured ceilings ranged from 37 pages to 130, which is why a fixed page cap is the wrong way to bound a run.
- Field referenceAll 58 fields grouped by what they describe, with the rate each was filled on across 2,229 jobs — and the three that are published as ranges rather than averages.
- Why no people dataFounders, employees, funding rounds and investors are all on Wellfound and none of them are in the output. What the line is, and why it sits where it does.