Four entity types, four shapes. Two fields in the match shapes are derived rather than read, and both are derived because the source publishes something that would otherwise be misread.
Results and schedule rows
| Field | Notes |
|---|---|
match_id | Stable within the date and tour. |
start_time_utc | Pinned to UTC. The name states the assumption. |
tour | atp-single, atp-double, wta-single or wta-double. |
tournament | The tournament the match belongs to. |
home_player / away_player | Full names, read from the attribute rather than the visible label. |
sets | The raw scoreline, always, exactly as published. |
home_sets / away_sets | Recounted from the scoreline on unfinished matches. |
status | Derived: completed, retired or walkover. |
The two derived fields
status is derived because the source labels no retirements or walkovers anywhere — how, and against what sample. About 4.6% of matches are affected.
The set counts are derived on a subset of rows, for the reason below.
Why set counts are recounted
So home_sets and away_sets are recounted from the scoreline on those rows, and sets always carries the raw string — which is what lets you check the recount rather than trust it.
Fields that need match detail
Round, surface, the players' rankings at the time and the head-to-head record. Each costs one extra request per match, so on a 400-match day the request count goes from 1 to 401 and the surcharge is $0.0024 per match.
Player and ranking rows
Player rows carry profile fields and, optionally, match history by season. Under-18 players are returned with birth_year instead of a full birthdate and no photo URL — why. Ranking rows carry position, points and the publication week.
Odds, and why they are empty by default
odds_home and odds_away exist and are null unless includeOdds is switched on. That default is deliberate — scores and odds are not the same kind of data.


