130 guides in 13 clusters
Guides
Reference material for the data these Actors return: what each field contains, how often it is actually populated, where the target site's own rules set the limits, and which obligations attach to which data. Written for somebody about to build a pipeline, not for a keyword.
Everything, most recently updated first

Pricing a held dataset
The same answer costs $1 per thousand from a file and $20 per thousand from a WHOIS query, and the gap is a cost structure rather than a discount. What that means for how an Actor should be priced.

Naming for the event
Two Actors in this catalogue are named for what they actually measure rather than for the keyword that would sell them, and both decisions cost search volume on purpose.

DNSSEC adoption by TLD
Measured across the whole zone, 5.0% of domains publish a DS record. Per registry the rate runs from 0% to 100%, which is what makes the global figure unusable on its own.

What dns_provider can tell you
The nameserver identifies who runs the DNS for 69.1% of domains. It does not identify what a site is built with, and the measurement that settled that question is worth seeing.

Detecting domain hijacking
Whoever controls the nameservers controls the mail, the site and the certificates. That change happens in the registry zone, which is the one layer most monitoring never looks at.

Zone exit is not expiry
Leaving the zone runs roughly 35 days ahead of a domain being released, and most exits are never releases at all. The `minDaysAbsent` filter is what turns the signal into something usable.

Zone entry is not registration
Zone entry is the earliest public signal that a domain went live, and roughly 20–30% of those events are re-appearances rather than new registrations. The number is an outside estimate and is labelled as one.

Passive DNS versus zone files
Two sources that look interchangeable for infrastructure questions and answer differently shaped questions. Which one to reach for depends on whether you need history, resolution or completeness.

Reverse nameserver lookup
Pivoting from a nameserver to the domains delegated to it is useful on infrastructure that belongs to one organisation and useless on a large provider's. The difference is the whole technique.

dnstwist versus zone search
Permutation engines find the squats their rules predicted. Searching the registry zone finds what is actually registered, including the spellings no generator would produce — and each approach misses something the other catches.

Typosquat detection
How to run a brand sweep across 1,075 gTLDs, how to read `exact`, `typo` and `contains` differently, and why an empty result is the outcome worth paying for.

Probably free, not available
Absence from a zone file usually means unregistered, and sometimes does not. The gap was measured against ICANN's own monthly registry reports rather than estimated, and it differs by TLD.

Bulk availability without WHOIS
Screening a naming shortlist against zone files instead of querying WHOIS per name: what it costs, what the verdicts mean, and where the method stops being enough.

gTLDs and ccTLDs
Country-code TLDs are run outside ICANN's contracts, so their zone files are not available through the zone data service at any price. The gap matters most in exactly the namespaces a startup cares about.

ICANN CZDS access
Access to gTLD zone files is granted one TLD at a time, by each registry operator, under an agreement that shapes what a product built on the data is allowed to look like.

What a zone file contains
A zone file is a delegation record: which domains exist in a TLD and where each one points its nameservers. It holds no registrant, no registrar, no dates and none of the domain's own records — and knowing that is what makes the six zone-file Actors readable.

Listing versus detail
Which fields only exist on product pages, what turning on `scrapeDetail` does to a run's page loads and its bill, and how to get the enriched fields for only the rows that need them.

Stock and sales ranges
The site shows `+25 vendidos` and the official API answers `RANGO_1_50`. Any dataset with an exact sales integer in it invented that integer.

The anti-bot wall
Mercado Libre's anti-bot response is a well-formed page with a 200 status and zero listings. Six different responses have to be told apart, and only one of them is data.

item_id and link shapes
Mercado Libre results mix `articulo.…/MLA-…`, `/p/MLA…` and `/up/MLAU…` links. Deriving the item id from the URL looks fine until the third shape appears, which can be nearly half a page.

Pagination and robots.txt
Measured in September 2026: the plain paginated path returns the anti-bot wall and the `_NoIndex_True` form returns results — and robots.txt disallows the second by name. There is currently no polite pagination available.

Venezuela's dual price
Venezuelan product pages show a dollar price and a bolívar price at once. Deriving the implied rate from the page's structured data rather than the rendered price is the difference between a stable number and one that drifts.

Mixed-currency pages
In Uruguay, Paraguay, the Dominican Republic, Nicaragua, Guatemala and Panama the local currency and US dollars appear in the same list of results — which makes a per-country currency column quietly wrong.

The 17 marketplaces
A reference table of every Mercado Libre site id, its country, its currency, the depth of a single query and whether it has catalog pages, mixed currencies and installments.

API alternative
The official Mercado Libre API requires OAuth and returns 403 to an anonymous request. A comparison of what the API gives an authorised caller, what scraping gives anyone, and which fields exist in only one of the two.

How to scrape Mercado Libre
A working method for extracting listings and product pages from any of the 17 Mercado Libre marketplaces, including the two decisions — pagination and detail pages — that decide what a run costs.

The refused countries
Aggregating name, phone and coordinates for hundreds of thousands of small businesses produces a personal-data file. Where that leads under the GDPR, UK GDPR and the LGPD.

Field reference
Every field the Actor returns, which are always present, which depend entirely on the market, and the four that are deliberately absent.

Reading the run summary
Four fields in the run summary decide whether your data covers the area you asked for. What each one means and which combinations should stop a pipeline.

API versus scraping
Google's own API is the right tool for looking up a place. It is a different proposition when the question is every business in a region, and the difference is structural rather than about price alone.

Joining on place_id
Two identifiers ship on every row and they do different jobs: one deduplicates within and across runs, the other joins to anything you already have from Google.

The email question
Tools advertising emails with Google Maps data are getting them somewhere else. Zero email strings across three verticals, in the search response and on the place page.

Coverage is the market
Phone, hours, rating and website vary enormously by city and by vertical. Two independent samples, 16,725 businesses, and why an average would be the wrong thing to publish.

The census pass
The first pass extracts no businesses at all. It tells you which parts of an area are dense, which belong to a neighbouring country and which are empty — before you spend anything.

The 200-result cap
The cap is per query, not per page, so the answer is geographic rather than paginated. What saturation means, why the threshold is 190, and how deep the splitting goes.

Scrape Google Maps
A walkthrough of the two-pass approach: classify the area first so you know what it costs, then extract only the tiles that need it.

Why odds are off
The same results page carries match scores and bookmaker odds, and the two sit in different positions. Why one is on by default and the other is not.

Field reference
What each of the four entity types returns, which fields are derived rather than read, and the one column that is recounted when the source publishes a flag instead.

Player profiles
Profiles with as many past seasons as you ask for, one request per season each — and a deliberate reduction for players under 18.

Rankings
Singles, doubles and the season race, for the current week or a past one — with the constraint that a historical week has to be a date the source actually published.

Backfilling history
History goes back to 1995 with the same input. What a season actually costs, why the price has two tiers, and the settings that decide how long it takes.

Truncated doubles names
The visible team label in doubles is cut short — `Roger-Vas` rather than the two full surnames. The complete names are in the markup, and that is 17.4% of all rows.

The moving day boundary
TennisExplorer cuts its day using a timezone cookie, so two runs of the same date can legitimately disagree. Pinning every request to UTC is what makes a dataset reproducible.

Deriving match status
An unfinished match just shows a scoreline that never closes. Working out which ones those are, without being told whether a match is best-of-3 or best-of-5, across 2,364 matches.

The tennis API question
The tours publish rankings and draws for readers, not as APIs. Commercial feeds are licensed per use. What is left is reading a results site, and what that does and does not entitle you to.

Scrape tennis results
A walkthrough of extracting tennis data: one entity type per run, the day boundary that has to be pinned, and the two enrichments that cost an extra request each.

Why no people data
Founders, employees, funding rounds and investors are all on Wellfound and none of them are in the output. What the line is, and why it sits where it does.

Field reference
All 58 fields grouped by what they describe, with the rate each was filled on across 2,229 jobs — and the three that are published as ranges rather than averages.

Pagination ceilings
Each role paginates until its own corpus runs out. Measured ceilings ranged from 37 pages to 130, which is why a fixed page cap is the wrong way to bound a run.

Rate limits and proxies
Wellfound rate-limits per IP, and the difference between a datacentre address and a residential one was 24 failures out of 36 against zero out of 48.

Cache age through a CDN
A row can be fetched seconds ago and still be a day old. Cloudflare's cache sat in front of Wellfound job pages at up to 23.4 hours, and `data_age_hours` counts it.

Salary and equity parsing
Compensation on Wellfound is a display string. What it takes to split it into minimum, maximum, currency, period and an equity range — and why the original string still ships.

The column not shipped
Wellfound exposes a maximum years-of-experience value. It was filled on 2 of 2,229 jobs. Shipping it would have added a field that looks like data and is not.

The ATS spread
`ats_source` names the applicant tracking system behind a posting. Its fill rate varies by 62 percentage points across role categories, which is why the average is not published.

The Wellfound API question
There is no documented public endpoint for startup job listings. What robots.txt permits, what Cloudflare sits in front of, and which surface is actually readable.

Scrape Wellfound jobs
A walkthrough of extracting startup job listings: how Wellfound's URL-path filtering constrains what you can ask for, which caps actually bound a run, and what arrives in each row.

Data protection
A public post can contain an email, a phone number and a handle, and returning it in full is what makes it useful and what makes it personal data. Where that leaves the person running the run.

Channel id versus handle
A channel that renames itself breaks every dataset keyed on its handle. The internal id survives the rename, and it is on every row.

Channel discovery
Discover mode searches a curated catalogue of verified public channels by niche. What that catalogue is, what it is not, and why searching for a technology beats searching for an English phrase.

Field reference
Every field the Actor returns, with the rate it was actually filled on across 893 posts in 9 channels — including the two that vary enough by channel type that an average would mislead.

The attachment finding
Documents, voice notes, audio, stickers, locations and round videos never appeared once in the anonymous preview. Fields that would always be false were cut rather than shipped.

Stemming false positives
Telegram's search stems words rather than matching substrings, so it returns morphological relatives of your term. Filtering them out afterwards is free, and it has to happen before billing.

Exact subscriber counts
Every abbreviated figure Telegram displays is a rounded string. One extra 4 KB request returns the exact integer, and both numbers ship so the rounding stays auditable.

Search versus crawl
Reading a channel's history to find 20 posts about one term took 25 requests and 712 KB. Asking Telegram to filter took 1 request and 24 KB. Same 20 results.

The Telegram API question
The Bot API cannot read a channel it is not in. MTProto needs a phone number and an api_id. The public web preview needs neither, and it is what a channel publishes to the open web.

Scrape Telegram channels
A walkthrough of reading public Telegram channels anonymously: the four modes, when search beats reading a history, and what the anonymous preview does and does not render.

Is it legal?
Federal notices are published because the law requires agencies to publicise them, and U.S. Government works carry no copyright. That settles more than usual — and not everything.

Opportunity monitoring
A working monitoring pipeline: what to filter, how often to run, which deadline field to schedule against, and where the attachment rate changes your plan.

Exclusions and debarment
63% of the federal exclusion list is individuals and only about a fifth of those carry a UEI. A name-matching lookup would manufacture false positives about real people.

NAICS vs PSC
Two classification systems on the same notice, describing different things. One is better filled, and the difference decides which one your pipeline should key on.

Set-aside codes
SBA, 8(a), WOSB, SDVOSB, HUBZone — the codes that reserve a contract for a category of small business, present on 46.1% of notices.

SAM.gov data fields
62 fields, snake_case, null where a value is genuinely absent and never a missing key — with coverage measured across five days spread over ten months.

Deadline precision
SAM.gov stamps midnight UTC on deadlines that are date-only. Passing that through as an instant would present a precision that does not exist, so it is null instead.

Solicitation documents
Measured on a stratified sample covering 96.7% of the active index: Veterans Affairs posts documents on nearly every solicitation, the Department of Defense on about a quarter — and DoD is 63% of the index.

The API key question
Several SAM.gov scrapers require you to register for an api.sam.gov key. That moves the quota problem, the approval wait and the renewal onto you — for data that is public either way.

Scrape SAM.gov
A walkthrough of extracting U.S. federal contract notices: which filters actually narrow the index, how to get the solicitation documents, and why there are three deadline fields.

Skill slugs
`react` returns zero projects with HTTP 200. `react-js` returns hundreds. A wrong slug is the worst kind of failure, so it gets checked before the run rather than after.

Proposal velocity
Measured across 401 projects: competition on a freelance listing has a half-life measured in hours, which makes a stale row a wrong row.

Market analysis
What listing data supports — demand by skill, budget distributions, competition levels — and the four claims it cannot carry.

Workana data fields
55 fields measured on 718 projects across 12 subcategories and 8 countries — reported as how often a field carries useful information, not how often the key exists.

The rating that is not a rating
Workana returns a rating for every project, so "100% coverage" would be true. It would also throw away four fifths of the market if you believed it.

Parsing budgets
"USD 1,000" and "USD 1.000" are both one thousand. "Less than USD 50" has no lower bound. Both facts break naive budget parsing in ways that survive review.

The coverage ceiling
About 350 projects per query is the limit for every tool. The interesting part is that Workana cannot tell you what you are missing either — so the Actor writes a coverage report instead of a percentage.

Filters that do nothing
`budget_min`, `is_hourly`, `duration`, `payment_verified` and nine others return HTTP 200 and change nothing. One of them makes the result set bigger.

Workana API
No public API, a robots.txt that disallows the internal one, and a set of URL parameters that accept anything. What a compliant integration actually has to work with.

Scrape Workana
A walkthrough of extracting LATAM freelance demand: the five filters that work, the coverage ceiling nobody can exceed, and why the language setting changes what a budget means.

Korean search terms
Naver's search is Korean-first. The choice of term decides coverage more than any other input, and the region list decides whether the results are where you think they are.

Reputation aggregates
Average rating, a ten-band star distribution, reviewer counts and a theme analysis with counts — all computed by Naver, and all cheaper than deriving them from raw reviews.

Naver vs Google Maps
In most countries this is not a question. In Korea it is the whole question, because the platform Koreans use is not the one most tooling targets.

Reviews without identities
Every Naver review carries a nickname, a stable account id and a link to the reviewer's entire history. None of it is returned, and a test fails the build if it ever is.

Virtual phone numbers
Naver puts a relay number in front of many businesses' real lines. Between 37% and 87% of them, depending on vertical, publish only that — and a dataset that does not flag it is a call list into a redirect.

Naver Place fields
Measured on 3,295 businesses across 8 verticals and 8 regions — including which fields simply do not exist for clinics, academies and pharmacies.

The 300 ceiling
Every Naver query stops at 300 results regardless of what it claims to have matched. Coverage is a function of how you partition the map, not of how high you set a limit.

The padding finding
Naver silently pads its search results with businesses from neighbouring areas. Measured across ten queries — and it is the difference between a local dataset and a regional blur.

Naver Place API
Naver publishes developer APIs, but not the one people are looking for. What exists, what it covers, and why the public pages are the practical route.

Scrape Naver Place
Google Maps coverage in Korea is thin; Naver Place is where the data is. A walkthrough of querying it, and of the padding you have to remove before the results mean anything.

Agency lead lists
Agency profiles come with public business phones and a listing count, which together are a better qualification signal than either alone.

New developments
New-build projects have their own fields, their own pricing logic and their own reason to be tracked separately from resale stock.

Geolocated property data
Most property datasets give you a neighbourhood. Coordinates on the listing itself let you do the analysis that neighbourhood averages cannot.

Agency vs private seller
Zonaprop labels private owners separately from agencies, and that label decides whether a phone number is a business contact or an individual's personal data.

Zonaprop data fields
A field-by-field reference for the property, development and agency rows — including an explicit note about which fields carry a measured fill rate and which do not.

USD and ARS prices
Argentine listings are priced in dollars or pesos, and expenses almost always in pesos. Mixing them without care produces analysis that is wrong by two orders of magnitude.

The proxy requirement
Zonaprop sits behind Cloudflare and challenges the country its listings are in. The proxy requirement is a measurement, not a preference — and the concurrency default is part of the same finding.

The coverage ceiling
The gap between what a portal says it has and what it will hand over is where most real-estate datasets quietly go wrong. Here is the ceiling, where it comes from, and the only way past it.

Zonaprop API
Zonaprop has no public listing API. What it does have is a detailed robots.txt and a Cloudflare layer, which together define what a compliant integration can read.

Scrape Zonaprop
A walkthrough of extracting Argentine property data: why the proxy country matters, what the 5-page robots.txt cap means for coverage, and how to read prices quoted in two currencies.

Measured fill rates
A field list is a promise. A fill rate is a measurement. This is what happened when one Actor's numbers were re-measured at n=741 instead of n=30.

Cache metadata
Shared caching makes runs fast and cheap and quietly destroys trust — unless the row itself reports where it came from and when.

Pay-per-event pricing
Charging per delivered row changes what you are allowed to ship. If an error row costs the user money, every bug becomes a billing dispute.

robots.txt as a spec
A site's robots.txt is the only machine-readable statement it makes about crawling. Treating it as the specification — and downloading it on every run — is why a scraper survives.

Comparing data sources
Coches.net, Milanuncios, Wallapop and AutoScout24 all list Spanish cars, and they differ in the one thing that matters for data work: how much structure the site publishes.

Spanish market analysis
What you can and cannot conclude from marketplace listings: asking prices are not transaction prices, inventory is not demand, and the badge distribution tells you more than the fuel type does.

GDPR and classifieds
Scraping a marketplace means scraping individuals. Article 14 attaches an obligation to that which no scraper can discharge at scale — and a setting does not remove it, it transfers it.

Coches.net data fields
A reference for the 55 fields in a coches.net listing row, each with the percentage of 550 real listings that had it filled, plus the two fields deliberately left out.

Car dealer lead lists
Dealer profiles come with business phone, address, postcode and province at 100% coverage on a 60-dealer sample. Private sellers' numbers are never included, and that is not a configuration gap.

DGT badge data
Spain's environmental badge decides where a car can drive and what it is worth. What each code means, how often it appears in listing data, and why it stays untranslated.

Cars below market price
A concrete screen: which fields to filter on, which segments carry the valuation data, and how to avoid the two biases that make a below-market list look better than it is.

Price rank explained
Coches.net computes its own valuation for each model and publishes it next to the asking price. Two fields, filled on 77.6% of listings, that replace a pricing model you would otherwise have to build.

Coches.net API
Coches.net publishes no public listing API. What it does publish is a server-rendered payload and a detailed robots.txt — which together define exactly what a compliant integration can read.

Scrape coches.net
A walkthrough of extracting used-car listings from Spain's largest vehicle marketplace: how discovery works by facet rather than pagination, what the 55 fields contain, and what a thousand listings cost.

Is it legal?
The interesting question is not "is scraping legal" but which specific data creates which specific obligation. Public business contact details and neighbor posts are not the same thing.

Trade area vs city
Asking Nextdoor for 200 dentists in one city returns the surrounding four towns too. We counted where the switch happens: results 1-75 were 96-100% local, everything past 76 was 0%.

Nextdoor city data
The city dataset is the least known part of Nextdoor's public surface and the most useful for market sizing: residents, income, age, homeownership, subjective scores and full category coverage.

Local lead generation
Which Nextdoor fields are usable for outreach, what the fill rates mean for a target list of a given size, and how to avoid building a list that is 40% dead ends.

Scraping without cookies
Most Nextdoor scrapers ask for your session cookies. They work on setup day and fail silently later. The trade-off of refusing them, stated in both directions.

Nextdoor city slugs
Nextdoor identifies a city as `city-name--state`. The rule sounds trivial and breaks on saints, hyphens and states that share a city name — here is the exact format and how to confirm a slug before spending a run on it.

Nextdoor data fields
A field-by-field reference for the 34 attributes in a public Nextdoor business record, each with the percentage of 741 real businesses that had it filled in.

Recommendations vs reviews
Measured across 891 recommendations, 50% were neighbors asking for a provider and only 19% were actual reviews. Why the feed mixes three things, and how to separate them before they poison a sentiment model.

Nextdoor API: what exists
Nextdoor has no public data API for business listings or recommendations. Here is what its actual interfaces cover, why the gap exists, and how to get structured data without one.

Scrape Nextdoor business listings
A complete walkthrough: how Nextdoor exposes business data to signed-out visitors, how to select cities and categories, what a first run costs, and how to read the 34 fields you get back.