The same date returned 227 matches once and 269 the next time
TennisExplorer cuts its day using a timezone cookie, so two runs of the same date can legitimately disagree. Pinning every request to UTC is what makes a dataset reproducible.
TennisExplorer decides which matches belong to which calendar day from a timezone cookie, so the same requested date returned 227 matches in one timezone and 269 in another during testing. Neither answer is wrong; they are answers to different questions. A dataset built without pinning that setting is not reproducible — two runs of the same date disagree and nothing in the output explains why. This Actor pins every request to UTC and names every time field `start_time_utc`, so the boundary is visible in the schema rather than assumed by the reader.
Key points
The same date returned 227 matches in one timezone and 269 in another, measured directly.
The difference is matches near midnight, which move between days as the boundary moves.
Neither count is an error: both are correct for their own day boundary, which is why nothing surfaces as a failure.
Every request is pinned to UTC, so two runs of the same date return the same set of matches.
The field is named `start_time_utc` rather than `start_time`, so the assumption is stated in the column name.
A dataset is reproducible when asking the same question twice gives the same answer. This is a case where it did not, and the reason is not a bug anywhere.
The measurement
Request
Matches returned
The date, in one timezone
Matches227
The same date, in another
Matches269
Forty-two matches of difference, for the same requested calendar date.
Why a day boundary moves at all
The site presents match times in the visitor's timezone, which is a sensible thing to do for a reader, and it decides which calendar day a match belongs to using that same setting. A match starting at 23:30 in one zone starts at 01:30 the next day in another, so it appears on a different date.
With a full calendar of matches across the world, the number of matches near a boundary is not small.
Why nothing looks broken
Both responses are correct. Neither returns an error, a warning or a duplicate. A scraper that does not pin the setting simply inherits whatever the request happened to negotiate, and produces a dataset whose day assignment depends on conditions nobody recorded.
Pinning it to UTC
Every request this Actor makes is pinned to UTC. Not configurable, because a configurable day boundary reintroduces exactly the problem: two datasets built with different settings look identical and are not comparable.
Putting the assumption in the column name
The field is start_time_utc, not start_time. A column name is the cheapest documentation there is, and it travels with the data into every export, notebook and database somebody copies it into.
Frequently asked questions
▸Why do two runs of the same date return different matches?
Because TennisExplorer cuts its day from a timezone cookie, and matches near midnight move between days when the boundary moves. Measured: 227 matches against 269 for the same requested date in two timezones.
▸Which day boundary does this Actor use?
UTC, on every request, always. That makes two runs of the same date return the same set of matches, which is the property a reproducible dataset needs.
▸Does that mean some matches are missing?
No — it means each match is assigned to exactly one day, consistently. A match at 23:30 local time somewhere sits on the UTC day it started on, and it is in your data under that date every time you ask.
Sources
Every URL below was requested and returned a page on the date shown.
The visible team label in doubles is cut short — `Roger-Vas` rather than the two full surnames. The complete names are in the markup, and that is 17.4% of all rows.
An unfinished match just shows a scoreline that never closes. Working out which ones those are, without being told whether a match is best-of-3 or best-of-5, across 2,364 matches.
A walkthrough of extracting tennis data: one entity type per run, the day boundary that has to be pinned, and the two enrichments that cost an extra request each.