“How fresh is this data?” has two answers, and most tools give you the wrong one without saying so.
There are two caches, and only one is yours
A scraper caches to avoid refetching. A CDN in front of the target caches to avoid serving the origin. A request that misses the first can still be answered from the second, which means a response can be hours old at the moment it arrives.
What was measured
On Wellfound job detail pages, Cloudflare's own cache age was measured at up to 23.4 hours. Not a theoretical ceiling — an observed value on the pages this Actor reads for enrichment.
What `data_age_hours` therefore means
It is the age of the data, not the age of the request. It includes the upstream cache age, so a row fetched sixty seconds ago can legitimately report 23 hours. That looks wrong the first time you see it and is the only honest value: the alternative is a field that says one minute about a page that has not changed since yesterday.
What forcing a fresh fetch can and cannot do
maxCacheAgeDays: 0 makes the Actor refetch rather than serving its own cached row. It cannot make a CDN revalidate against its origin. So a forced-fresh run still reports real ages, and those ages are still sometimes large — which is the information, not the failure.
The consequence for a pipeline
Set a freshness threshold on data_age_hours rather than on run time, and accept that a sub-hour threshold may be unreachable for pages behind an aggressive cache. Every row in this catalogue carries the same three provenance fields for this reason — why every row carries its own age.


