A regular expression that works on every row you looked at is not the same as one that works on every row.
The three shapes, and what each one is
| Link | What it is | link_kind |
|---|---|---|
articulo.mercadolibre.…/MLA-… | An individual seller's publication | publication |
/p/MLA… | A catalog product page shared by several sellers | catalog |
/up/MLAU… | A user product | user_product |
All three can appear in the same list of search results.
Why a URL pattern looks like it works
Write a pattern for the first two shapes and it extracts a clean id from most rows. The rows it cannot match are simply dropped, and the rows that remain are all valid — so the dataset passes every check you would think to run. The third shape can be 40% of a page.
Where the id actually comes from
The results page embeds its own search results array, with the item id already separated from however the link happens to be written. Reading item_id from there makes the key independent of URL formatting — which is a good idea in any case, because a marketplace can change a link shape without announcing it and has no reason to keep a fourth from appearing.
Publishing the shape as a field
link_kind and is_catalog come back on every row. That turns the distinction from a parsing detail into something a consumer can filter or audit on — useful, because a catalog page and a seller publication are genuinely different objects with different pricing behaviour.
One key doing three jobs
item_id is the primary key of the dataset, the cache key for the 24-hour listing cache, and the deduplication key across runs. An unstable derivation would corrupt all three at once — and the failure would look like a slightly small dataset rather than an error, which is the same shape of problem as a currency stored at the wrong level. For how the rows are gathered in the first place, see the run guide.


