Typosquatting Detection — Find Lookalike Domains, 1,075 TLDs
A match means the domain really exists today.
Typosquats, homograph tricks, brand-plus-keyword domains and lookalikes in TLDs nobody thought to check, in one run across 1,075 gTLDs. Because it searches the registries' own zone files rather than generating permutations and testing them, it finds squats whose spelling no permutation engine would have guessed.
oswaldocarabano/brand-domain-sweep
{
"brand": ["stripe"],
"matchTypes": ["exact", "typo"],
"maxEditDistance": 1,
"maxResults": 100
}- Version
- v0.1.8
- Memory
- 4096 MB
- Browser
- none
- Proxy
- None — the zone data is held, not fetched
Short answer
The Typosquatting Detection Actor finds every delegated domain across 1,075 gTLDs that matches, imitates or contains a given brand, reading the registries' own zone files so a match means the domain really exists today rather than that a scraper saw it last month. Matches are classified as `exact`, `typo` with an edit distance, or `contains`, and each row carries the domain's nameservers, its DNS provider and whether it sits on a parking or domain-sale service. Pricing is $0.004 per matching domain, about $4 per 1,000, against roughly $30 per 1,000 for a dnstwist-based detector — and an empty result, which is the good news, costs only the $0.00001 start fee. A match is not an infringement: the Actor reports what exists and the judgement is the reader's.
Key points
- Searching the zone files finds squats whose spelling no permutation engine would have guessed, which is the structural difference from dnstwist-style tools that generate variants of a brand and then test whether each one resolves.
- A match proves the domain is delegated in the registry's zone file today, rather than that it was registered once or that a crawler observed it at some point in the past.
- Pricing is $0.004 per matching domain against roughly $0.03 for a dnstwist-based detector in the same Store, making a sweep about seven times cheaper per result found.
- An empty result is the good news and costs only the $0.00001 start fee, because nothing is charged when nothing matches.
- Narrowing the sweep to named TLDs makes a run substantially faster and cheaper, because the service skips every other TLD instead of reading through all 1,075.
- Edit distance 2 surfaces names that share letters with a brand by coincidence, which is why brands under five characters should start at distance 1.
- A match is not an infringement: `stripe-tools.com` may be a legitimate integration partner and `strlpe.com` may be a parked typo nobody will ever use.
What it does
Typosquats, homograph tricks, brand-plus-keyword domains and lookalikes in TLDs nobody thought to check, in one run across 1,075 gTLDs. Because it searches the registries' own zone files rather than generating permutations and testing them, it finds squats whose spelling no permutation engine would have guessed.
{
"domain": "strlpe.com",
"tld": "com",
"match_type": "typo",
"edit_distance": 1,
"nameservers": ["ns1.sedoparking.com", "ns2.sedoparking.com"],
"dns_provider": "Sedo",
"parked_for_sale": true
}Why this one
Search, not permutation
dnstwist and the tools built on it generate variants of a brand — character swaps, homoglyphs, added hyphens — and then test which ones resolve. That finds what the generator thought of. Reading the zone files finds what is actually there, including the squat spelled in a way no permutation rule would produce, and the brand-plus-keyword domain that is not a typo at all.
Existence, with a date on it
A domain in the zone file is delegated now. Passive-DNS and historical datasets answer a different question — whether something was ever seen — and a brand protection report full of domains that lapsed months ago wastes the time of whoever has to triage it. The zone snapshot is refreshed daily and a match means today.
Three match types that mean different things
`exact` is certain: the label is the brand. `typo` and `contains` are search rather than judgement, and the README says so. Edit distance is returned on every typo row so the reader can sort by how close a match actually is, instead of receiving one undifferentiated list where a one-character swap sits next to a coincidence.
A column that is true instead of one that looks impressive
`dns_provider` is filled on 69.1% of rows and names who runs the DNS — a hosting fact. It is not a technology fingerprint, and the Actor does not guess one: website platforms are configured with A and CNAME records, which a zone file does not contain. Shopify, for instance, is visible on 185 domains out of 64 million, which is why that column was not shipped.
Triage built into the row
`parked_for_sale` fires on 2.0% of domains and is exact when it does. A match sitting on Afternic or Sedo is a squatter waiting to sell; a match on a small unknown nameserver running a fake shop is a different problem with a different response. Separating those two in the output is what makes a long result list actionable.
Use cases
- Run a typosquatting sweep before a product launch, across every gTLD rather than the handful a brand already owns.
- Find phishing lookalikes imitating a bank or a payment brand, including the ones that contain the brand inside a longer name.
- Monitor a brand and its product names on a weekly schedule, and treat an empty result as a clean week.
- Build the evidence list for a UDRP or registrar complaint, with the delegation and the sale service recorded per domain.
- Check which exact-match brand domains exist across TLDs the brand never registered defensively.
Input
Every field has a default, and the defaults are deliberately small so a first run is cheap enough to inspect before you commit to a sweep. This table mirrors the Actor's own input schema field for field.
| Field | Default | What it does |
|---|---|---|
brandstring[] | [] | Brand or keywordThe word or words to look for — letters, digits and hyphens. This is what anchors the search; there is no way to ask for every domain. |
tldstring[] | [] | Limit to these TLDsOptional. Restricting the sweep to named TLDs makes a run substantially faster and cheaper, because the service can skip every other TLD instead of scanning them. Empty sweeps all 1,075. |
matchTypesstring[] | ["exact", "typo", "contains"] | Match types`exact` when the label is the brand, `typo` when it is one or two characters away, `contains` when the brand appears inside a longer name. |
maxEditDistanceinteger | 2 | Max edit distancepersonal dataHow far a typo may sit from the brand. 1 is conservative, 2 is thorough and noisier, and short brands should start at 1. |
maxResultsinteger | 100 | Max resultspersonal dataHard cap on rows returned, kept low on purpose so a first run cannot burn a free credit. The service never returns more than 50,000 rows in one run. |
Output and fill rates
A field being in the schema is not the same as it having a value. The percentages below were counted on real runs; the sample sizes are in Measurements. Anything not listed here is not promised.
| Field | Filled | Meaning |
|---|---|---|
domainstring | 100% | The matching domain, lowercase and in A-label form, so an internationalised squat arrives comparable rather than rendered. |
tldstring | 100% | The top-level domain, one of the 1,075 covered gTLDs. |
match_typestring | 100% | `exact`, `typo` or `contains`. Only `exact` is a certainty; the other two are search. |
edit_distanceinteger | not measured | How many characters the name sits from the brand, on `typo` rows, so a result list can be sorted by closeness. |
nameserversstring[] | 100% | Every nameserver the domain delegates to, as recorded in the zone. |
dns_providerstring | 69.1% | Who runs the DNS, inferred from the nameserver. A hosting fact rather than a technology one. |
parked_for_saleboolean | 100% | True when the nameserver belongs to a parking or domain-sale service. Fires on 2.0% of domains and is exact when it does. |
Every key is always present. A field that exists but is empty comes back as explicit null, so a parser never has to guess.
Datasets
Different record types go to different datasets, so the main table never carries columns that are blank on most rows.
defaultOne row per matching domain, with its match type, edit distance and delegation.billed
Pricing
Pay per delivered result. Charges are applied as each row is produced rather than in a lump at the end, so an aborted run bills only for what it actually gave you.
| Event | Price | Notes |
|---|---|---|
actor-startActor start | $0.00001 | Effectively free. A run that finds nothing costs you nothing, which is what makes a clean week cheap to confirm. |
match-foundMatching domain | $0.004 | One domain that matches, imitates or contains the brand, with its nameservers, DNS provider and sale-service flag — about $4 per 1,000, and $0.40 for a default run of 100 rows. Rows are only charged when they match the filters. |
Measurements
Each figure is shown with the method that produced it. A benchmark without a method is a marketing claim wearing a number's clothes.
Zone coverage
255,631,856 domains across 1,075 gTLDs
Counted from the zone files themselves rather than quoted from a registry summary, and refreshed daily.
Cost against a dnstwist-based detector
$4 against about $30 per 1,000 results
$0.004 per matching domain here against the $0.03 per registered domain found charged by a dnstwist-based typosquatting detector in the Apify Store.
`dns_provider` fill rate
69.1%
Measured across every domain in the zone. The remaining 30.9% delegate to nameservers that identify no known operator.
`parked_for_sale` rate
2.0% of domains
Measured across the whole zone by matching nameservers against known parking and domain-sale services. Exact when it fires.
Website platform visibility
Shopify on 185 domains out of 64 million
Measured while deciding whether to ship a platform column. Platforms are configured with A and CNAME records, which a zone file does not contain, so the column was not shipped.
Coverage concentration
.com is 166 million of 255 million
Counted per TLD from the zone files. 481 of the 1,075 covered TLDs hold fewer than a thousand domains each, so breadth catches the obscure case rather than doubling the volume.
What it will not do
Stated plainly so you can judge fit before spending anything.
- A match is not an infringement, and the Actor makes no legal judgement: deciding what to act on is the reader's job and their lawyer's.
- `typo` and `contains` are search rather than judgement. Edit distance 2 surfaces coincidental matches, especially for brands under five characters.
- A domain that exists is not necessarily live: the zone file proves delegation, not that a website answers.
- Country-code TLDs are not covered — `.io`, `.ai`, `.co`, `.me`, `.tv` and `.cc` are run outside ICANN's contracts and are unavailable at any price.
- Homograph squats registered in a ccTLD are therefore invisible here, which matters because ccTLDs are a common choice for them.
- A zone file contains no registrant, no email, no registrar and no registration or expiry date, so the output cannot tell you who is behind a squat.
- Zone data is refreshed once a day, so a squat registered this morning may not appear until tomorrow's snapshot.
- The service never returns more than 50,000 rows in one run.
Privacy
- Zone files contain delegation records, not people: there is no registrant, no email address and no registrar in the source.
- The data is obtained through ICANN's Centralized Zone Data Service under agreement with the Registry Operators, and ICANN does not endorse, sponsor or review this Actor.
- No WHOIS or RDAP query is made at any point, because querying registries at scale is prohibited by the agreement that provides this data.
- Every run is anchored on a brand term the operator supplies, and no input can return a substantial portion of a zone.
- Whoever runs the Actor is responsible for what they do with a match, which the output deliberately does not characterise as infringing.
- Removal requests: privacy@actorstack.dev
See also the data removal process.
Frequently asked questions
How is this different from dnstwist?
Does a match mean somebody is infringing my brand?
Should I use edit distance 1 or 2?
Can I limit the sweep to the TLDs I care about?
Does an empty result cost anything?
Will it find squats in .io or .ai?
Does a match mean the fake site is live?
Can I find out who registered a squatted domain?
Guides for this Actor
- Typosquat detectionHow to run a brand sweep across 1,075 gTLDs, how to read `exact`, `typo` and `contains` differently, and why an empty result is the outcome worth paying for.
- dnstwist versus zone searchPermutation engines find the squats their rules predicted. Searching the registry zone finds what is actually registered, including the spellings no generator would produce — and each approach misses something the other catches.