ActorStack.dev

Your freshness field is lying if it does not count Cloudflare

A row can be fetched seconds ago and still be a day old. Cloudflare's cache sat in front of Wellfound job pages at up to 23.4 hours, and `data_age_hours` counts it.

By Oswaldo Carabano6 min read

Short answer

A scraper that reports how long ago it made a request is not reporting how old the data is, because a CDN in front of the target can serve a response that was cached hours earlier. Measured on Wellfound job detail pages, Cloudflare's own cache age reached 23.4 hours. This Actor's `data_age_hours` includes that figure rather than only its own cache, so a row fetched a minute ago can correctly report an age of 23 hours. Setting `maxCacheAgeDays` to 0 forces a fresh fetch from the Actor's side, which still cannot bypass the CDN in front of the site.

Key points

  • Cloudflare cache age in front of Wellfound job detail pages was measured at up to 23.4 hours.
  • A freshness field that counts only the scraper's own cache reports the age of the request, not the age of the data.
  • `data_age_hours` adds the upstream cache age, so it can exceed the time since the fetch — which is the correct behaviour, not a bug.
  • `maxCacheAgeDays: 0` forces the Actor to refetch, and cannot force the CDN to revalidate.
  • Anything time-sensitive — a posting that closed, a salary that changed — should be read against `data_age_hours`, not against `scraped_at`.
On this page5 sections

“How fresh is this data?” has two answers, and most tools give you the wrong one without saying so.

There are two caches, and only one is yours

A scraper caches to avoid refetching. A CDN in front of the target caches to avoid serving the origin. A request that misses the first can still be answered from the second, which means a response can be hours old at the moment it arrives.

What was measured

On Wellfound job detail pages, Cloudflare's own cache age was measured at up to 23.4 hours. Not a theoretical ceiling — an observed value on the pages this Actor reads for enrichment.

What `data_age_hours` therefore means

It is the age of the data, not the age of the request. It includes the upstream cache age, so a row fetched sixty seconds ago can legitimately report 23 hours. That looks wrong the first time you see it and is the only honest value: the alternative is a field that says one minute about a page that has not changed since yesterday.

What forcing a fresh fetch can and cannot do

maxCacheAgeDays: 0 makes the Actor refetch rather than serving its own cached row. It cannot make a CDN revalidate against its origin. So a forced-fresh run still reports real ages, and those ages are still sometimes large — which is the information, not the failure.

The consequence for a pipeline

Set a freshness threshold on data_age_hours rather than on run time, and accept that a sub-hour threshold may be unreachable for pages behind an aggressive cache. Every row in this catalogue carries the same three provenance fields for this reason — why every row carries its own age.

Frequently asked questions

Why is `data_age_hours` larger than the time since the run?
Because it includes the cache age of the response Cloudflare served, measured at up to 23.4 hours on Wellfound job detail pages. A row fetched a minute ago can legitimately be 23 hours old, and reporting it as one minute old would be the actual error.
Can I force completely fresh data?
You can force the Actor to refetch by setting `maxCacheAgeDays` to 0. You cannot force a CDN sitting in front of the site to revalidate, so a fresh request can still return a cached response — which is exactly why the age is reported rather than assumed.
Does this affect the listing pages too?
The 23.4-hour figure was measured on job detail pages, which is where enrichment reads from. Every row carries its own `data_age_hours` regardless of surface, so the number is per-row rather than a global claim.

Sources

Every URL below was requested and returned a page on the date shown.

  1. Platform docschecked 9 Sept 2026
    Cache-Control — Cloudflare cache conceptsCloudflare
  2. Operator claimchecked 9 Sept 2026
    Wellfound Jobs Scraper — Actor README and input schemaActorStack / Apify Store
A white measuring tape curving across a dark background, showing the numbers 15 to 45.
WellfoundMeasured

The ATS spread

`ats_source` names the applicant tracking system behind a posting. Its fill rate varies by 62 percentage points across role categories, which is why the average is not published.

6 min
A laptop showing lines of code on a wooden desk in a dimly lit room.
WellfoundMeasured

Pagination ceilings

Each role paginates until its own corpus runs out. Measured ceilings ranged from 37 pages to 130, which is why a fixed page cap is the wrong way to bound a run.

5 min
Racks of network equipment in a dimly lit server room, lit blue by their indicators.
WellfoundMeasured

Rate limits and proxies

Wellfound rate-limits per IP, and the difference between a datacentre address and a residential one was 24 failures out of 36 against zero out of 48.

5 min