TennisExplorer Results, Schedule, Players & Rankings
A match status the source never labels, and a day that does not move.
ATP, WTA, Challenger and ITF tennis from TennisExplorer: finished results with set-by-set scores, upcoming schedules, player profiles with career history, and ranking tables back to 1995. Two things it adds that the source does not have — a derived match status, and a day boundary pinned to UTC so two runs of the same date return the same matches.
oswaldocarabano/tennisexplorer-scraper
{
"entityType": "results",
"dateFrom": "2026-08-01",
"dateTo": "2026-08-07",
"tours": ["atp-single", "wta-single"],
"includeMatchDetail": false,
"maxItems": 5000
}- Version
- v0.1.13
- Memory
- 4096 MB
- Browser
- none
- Proxy
- Residential pool, built in
Short answer
The TennisExplorer Actor extracts ATP, WTA, Challenger and ITF match results, upcoming schedules, player profiles and ranking tables, with history back to 1995. It derives a `status` field — `completed`, `retired` or `walkover` — that the source never labels anywhere, validated against 2,364 matches across nine tournaments including Davis Cup, and affecting about 4.6% of matches. Every request is pinned to UTC because the site's own day boundary follows a timezone cookie: the same date returned 227 matches in one timezone and 269 in another. Pricing is $0.0015 per result for the first 10,000 rows of a run and $0.0003 after that, so a full season of about 148,000 matches costs roughly $56.
Key points
- A derived `status` field — `completed`, `retired` or `walkover` — that TennisExplorer does not label anywhere. Validated against 2,364 matches across nine tournaments with different formats, including Davis Cup, which mixes best-of-3 and best-of-5 on one page. About 4.6% of matches are affected.
- Every request is pinned to UTC. The site decides which matches belong to which day from a timezone cookie, and the same date returned 227 matches in one timezone against 269 in another during testing.
- Full names in doubles. The visible team name is truncated to something like `Roger-Vas`; the Actor reads the full name from the underlying attribute, which is 17.4% of all rows.
- History back to 1995 with the same input. A full season is about 148,000 matches — measured, not estimated — and one day of results costs one request.
- Two price tiers so backfills are affordable: $0.0015 per result for the first 10,000 rows of a run and $0.0003 after that. A week of results is about $4 and a season about $56.
- Betting odds are off by default, because odds are licensed from bookmakers while match scores are facts. Failed requests are never charged and go to the ERRORS record instead.
What it does
ATP, WTA, Challenger and ITF tennis from TennisExplorer: finished results with set-by-set scores, upcoming schedules, player profiles with career history, and ranking tables back to 1995. Two things it adds that the source does not have — a derived match status, and a day boundary pinned to UTC so two runs of the same date return the same matches.
{
"entity_type": "results",
"match_id": "2026-08-03-atp-single-1184",
"start_time_utc": "2026-08-03T14:30:00Z",
"tour": "atp-single",
"tournament": "Toronto",
"home_player": "Fritz Taylor",
"away_player": "Rublev Andrey",
"sets": "6-4, 3-2",
"home_sets": 1,
"away_sets": 0,
"status": "retired",
"odds_home": null,
"odds_away": null
}Why this one
The status the source refuses to state
TennisExplorer does not label retirements or walkovers anywhere. An unfinished match simply shows a scoreline that never closes, and every consumer of that data has to guess. This Actor derives `status` from the scoreline itself, and the derivation works for best-of-3 and best-of-5 without being told which is which — which is the hard part, and why it was validated on Davis Cup, where both formats appear on the same page.
`retired` means one thing, and it is stated
`retired` means exactly that the match started and did not finish: a retirement, an injury, a default, or an abandoned dead rubber. The source does not distinguish between those, so neither does the Actor. Splitting them would be inventing a distinction the data does not carry, and a consumer who assumed injury from a `retired` flag would be wrong a good share of the time.
A day boundary that does not move
The same date returned 227 matches in one timezone and 269 in another, because the site cuts its day using a cookie. Any dataset built without pinning that is not reproducible: two runs of the same date disagree, and neither is wrong. Every request here is pinned to UTC and every time field is named `start_time_utc`, so the boundary is visible in the field name rather than assumed.
Sets recounted where the source publishes a flag
For a match that did not finish, the site's result column is a won/lost flag rather than a set count — it reads `1-0` no matter how many sets were actually played. `home_sets` and `away_sets` are recounted from the scoreline in that case, so they are never the flag, and `sets` always carries the raw scoreline for anyone who wants to check the recount.
One entity type per run
Results, schedules, players and rankings have genuinely different fields, so the Actor returns one type per run and the file you download has one clean shape. Mixing them would give every row the union of four schemas and leave the consumer to work out which columns apply.
Use cases
- Build a match-results dataset with a usable status field, for modelling or for reporting.
- Backfill a full season, or the whole 1995 onward archive, at the bulk row price.
- Track upcoming schedules and in-progress matches on a fixed UTC day boundary.
- Pull ATP or WTA ranking tables, singles, doubles or race, for a current or historical week.
- Collect player profiles with as many past seasons of match history as you need.
Input
Every field has a default, and the defaults are deliberately small so a first run is cheap enough to inspect before you commit to a sweep. This table mirrors the Actor's own input schema field for field.
| Field | Default | What it does |
|---|---|---|
entityTypestring | "results" | What to scrape`results` for finished matches, `schedule` for upcoming and in-progress ones, `players` for profiles, `rankings` for ATP or WTA tables. One type per run, so the dataset has one clean shape. |
dateFromstring | "yesterday" | From date (UTC)`YYYY-MM-DD`, or a relative date: `today`, `yesterday`, `-7d`, `+3d`. Days are cut at 00:00 UTC always, which is what makes two runs of the same date agree. History goes back to 1995. |
dateTostring | "today" | To date (UTC)Inclusive. One request covers one whole day and typically returns 150 to 600 matches. |
toursstring[] | all four | ToursATP singles, ATP doubles, WTA singles, WTA doubles. Leaving all four selected is one request per day, which is the cheapest option. |
includeOddsboolean | false | Include betting oddspersonal dataOff by design. Odds are licensed from bookmakers and are the part of the source with the weakest reuse position; match scores are facts and odds are not. Turn it on only with your own basis for using them. |
includeMatchDetailboolean | false | Fetch match detail pagesAdds round, surface, full player names, both rankings and head-to-head, at one extra request per match. A day is 200 to 600 matches, so leave it off unless you need those fields. |
rankingTourstring | "atp-men" | RankingATP (men) or WTA (women). |
rankingTypestring | "singles" | Ranking typeSingles, doubles, or the season race. Race tables only exist for the current season. |
rankingDatestring | "" | Ranking weekA historical ranking week as `YYYY-MM-DD`. It must be one of the publication dates the site offers; empty means the latest. |
playerUrlsstring[] | [] | Player URLsTennisExplorer player pages to read. Leave empty to take the players from the selected ranking instead. |
playerMatchHistoryYearsinteger | 0 | Years of match history per player0 reads the profile only. Each extra season is one extra request per player. |
maxItemsinteger | 10000 | Maximum resultsA hard cap on delivered rows; you are only charged for what is delivered. Raise it for backfills — a season is about 148,000 matches and the whole 1995 to 2026 archive is around 2.4 million. |
maxConcurrencyinteger | 8 | Concurrency8 is plenty for day-to-day use. Capped at 16, the highest level measured against the site, and capped again by the run's memory: a 512 MB run tops out around 10. |
proxyConfigurationobject | built-in pool | Proxy configurationOptional. The Actor already routes through its own residential pool, so leaving this empty just works. Set it only to use your own Apify proxy, or none. |
Output and fill rates
A field being in the schema is not the same as it having a value. The percentages below were counted on real runs; the sample sizes are in Measurements. Anything not listed here is not promised.
| Field | Filled | Meaning |
|---|---|---|
statusstring | 100% | `completed`, `retired` or `walkover`. Derived: the source labels none of them. About 4.6% of matches are not `completed`. |
start_time_utcdatetime | 100% | Pinned to UTC, and the field name says so. The site's own day boundary follows a timezone cookie. |
home_playerstring | 100% | With `away_player`. Full names, read from the underlying attribute rather than the truncated visible label — which matters on the 17.4% of rows that are doubles. |
setsstring | not measured | The raw scoreline as published, always, so a recount can be checked against it. |
home_setsinteger | not measured | With `away_sets`. Recounted from the scoreline on unfinished matches, where the source publishes a won/lost flag rather than a set count. |
tourstring | 100% | `atp-single`, `atp-double`, `wta-single` or `wta-double`. |
tournamentstring | not measured | The tournament the match belongs to. |
roundstring | not measured | With `surface`, both rankings and head-to-head. Only when match detail is enabled. |
odds_homenumber | not measured | With `odds_away`. Only when odds are explicitly enabled, which they are not by default. |
birth_yearinteger | not measured | On player rows. Returned instead of a full birthdate for players under 18, who also carry no photo URL. |
rankinteger | not measured | On ranking rows, with points and the publication week. |
Every key is always present. A field that exists but is empty comes back as explicit null, so a parser never has to guess.
Datasets
Different record types go to different datasets, so the main table never carries columns that are blank on most rows.
defaultOne row per match, player or ranking entry, depending on the entity type the run selected.billedERRORS (key-value store)Every failed request with its reason. Failed requests are never charged.never billed
Pricing
Pay per delivered result. Charges are applied as each row is produced rather than in a lump at the end, so an aborted run bills only for what it actually gave you.
| Event | Price | Notes |
|---|---|---|
match-resultResult | $0.0015 | One match, player or ranking row delivered to your dataset, for the first 10,000 rows of a run. |
match-result-bulkResult (bulk) | $0.0003 | Rows beyond the first 10,000 in the same run, so historical backfills stay affordable. A full season of ~148,000 matches is about $56. |
match-detailMatch detail (surcharge) | $0.0024 | Per match enriched with round, surface, full names, both rankings and head-to-head. One extra request each. |
Measurements
Each figure is shown with the method that produced it. A benchmark without a method is a marketing claim wearing a number's clothes.
Status validation
2,364 matches across nine tournaments
Formats deliberately mixed, including Davis Cup, which puts best-of-3 and best-of-5 on the same page. The derivation does not need to be told which format a match is.
Matches affected by status
about 4.6%
Retirements, walkovers and abandoned matches together, from the same validation set. The remaining 95.4% are `completed`.
Timezone drift
227 matches against 269, same date
The same requested date in two timezones, because the site cuts its day from a cookie. Every request is pinned to UTC to remove it.
Truncated names
17.4% of all rows
Doubles rows, where the visible team label is truncated. The full name is read from the underlying attribute instead.
Season size
about 148,000 matches
Measured across a full season, not extrapolated. One day of results is one request; a season is 365.
Archive depth
1995 to 2026, around 2.4 million matches
The same input shape works for any past date, so a backfill is a date range rather than a different mode.
What it will not do
Stated plainly so you can judge fit before spending anything.
- `retired` covers retirement, injury, default and an abandoned dead rubber together, because the source does not distinguish them.
- Players under 18 are returned with `birth_year` rather than a full birthdate, and without a photo URL.
- One entity type per run. Results, schedule, players and rankings need separate runs.
- Race ranking tables only exist for the current season.
- A historical ranking week must be one of the publication dates the site offers; an arbitrary date will not resolve.
- Match detail costs one extra request per match, and a day is 200 to 600 matches, so it is off by default.
- `maxConcurrency` is capped at 16, which is the highest level actually measured against the site, and is capped again by the memory the run was given — a 512 MB run tops out around 10.
Privacy
- Match results, schedules and ranking tables are sporting facts published for the public to read.
- Players under 18 are returned with `birth_year` instead of a full birthdate, and without a photo URL. That is a deliberate reduction, not a gap in the source.
- Betting odds are off by default: they come from bookmakers rather than from the site, and redistributing them sits in a different position from redistributing scores.
- This Actor is independent and is not affiliated with, endorsed by, or connected to TennisExplorer or its operator.
- Removal requests: privacy@actorstack.dev
See also the data removal process.
Frequently asked questions
How do I tell a retirement from a completed match?
Does `retired` mean an injury?
Why are match counts different between tools?
How far back does the history go?
What does a season of history cost?
Why are betting odds off by default?
Can I get results, rankings and players in one run?
Guides for this Actor
- Scrape tennis resultsA walkthrough of extracting tennis data: one entity type per run, the day boundary that has to be pinned, and the two enrichments that cost an extra request each.
- The tennis API questionThe tours publish rankings and draws for readers, not as APIs. Commercial feeds are licensed per use. What is left is reading a results site, and what that does and does not entitle you to.
- Deriving match statusAn unfinished match just shows a scoreline that never closes. Working out which ones those are, without being told whether a match is best-of-3 or best-of-5, across 2,364 matches.
- The moving day boundaryTennisExplorer cuts its day using a timezone cookie, so two runs of the same date can legitimately disagree. Pinning every request to UTC is what makes a dataset reproducible.
- Truncated doubles namesThe visible team label in doubles is cut short — `Roger-Vas` rather than the two full surnames. The complete names are in the markup, and that is 17.4% of all rows.
- Backfilling historyHistory goes back to 1995 with the same input. What a season actually costs, why the price has two tiers, and the settings that decide how long it takes.
- RankingsSingles, doubles and the season race, for the current week or a past one — with the constraint that a historical week has to be a date the source actually published.
- Player profilesProfiles with as many past seasons as you ask for, one request per season each — and a deliberate reduction for players under 18.
- Field referenceWhat each of the four entity types returns, which fields are derived rather than read, and the one column that is recounted when the source publishes a flag instead.
- Why odds are offThe same results page carries match scores and bookmaker odds, and the two sit in different positions. Why one is on by default and the other is not.