ActorStack.dev

Three link shapes on one results page, and the 40% a URL pattern drops

Mercado Libre results mix `articulo.…/MLA-…`, `/p/MLA…` and `/up/MLAU…` links. Deriving the item id from the URL looks fine until the third shape appears, which can be nearly half a page.

By Oswaldo Carabano5 min read

Short answer

One Mercado Libre results page can contain three different link shapes at once: `articulo.mercadolibre.…/MLA-…` for a publication, `/p/MLA…` for a catalog product, and `/up/MLAU…` for a user product. A regular expression written against the first two silently drops the third, which can be 40% of a page, and the loss is invisible because the remaining rows are all valid. The Actor therefore takes `item_id` from the page's embedded search results array rather than parsing it out of the href, and publishes `link_kind` and `is_catalog` so a row says which of the three it came from. The same id is the primary key, the cache key and the dedupe key.

Key points

  • A single results page can mix `articulo.…/MLA-…`, `/p/MLA…` and `/up/MLAU…` links, which are publications, catalog products and user products respectively.
  • A URL pattern written for the first two shapes drops the third silently, and that third shape can account for 40% of a page.
  • Taking `item_id` from the page's embedded search results array rather than from the href makes the key independent of how the link happens to be written.
  • `link_kind` and `is_catalog` are published on every row, so a dataset can be filtered or audited by which shape a listing came from.
  • The same `item_id` serves as primary key, cache key and dedupe key, so an unstable derivation would corrupt all three at once.
On this page5 sections

A regular expression that works on every row you looked at is not the same as one that works on every row.

The three shapes, and what each one is

LinkWhat it islink_kind
articulo.mercadolibre.…/MLA-…An individual seller's publicationpublication
/p/MLA…A catalog product page shared by several sellerscatalog
/up/MLAU…A user productuser_product

All three can appear in the same list of search results.

Why a URL pattern looks like it works

Write a pattern for the first two shapes and it extracts a clean id from most rows. The rows it cannot match are simply dropped, and the rows that remain are all valid — so the dataset passes every check you would think to run. The third shape can be 40% of a page.

Where the id actually comes from

The results page embeds its own search results array, with the item id already separated from however the link happens to be written. Reading item_id from there makes the key independent of URL formatting — which is a good idea in any case, because a marketplace can change a link shape without announcing it and has no reason to keep a fourth from appearing.

link_kind and is_catalog come back on every row. That turns the distinction from a parsing detail into something a consumer can filter or audit on — useful, because a catalog page and a seller publication are genuinely different objects with different pricing behaviour.

One key doing three jobs

item_id is the primary key of the dataset, the cache key for the 24-hour listing cache, and the deduplication key across runs. An unstable derivation would corrupt all three at once — and the failure would look like a slightly small dataset rather than an error, which is the same shape of problem as a currency stored at the wrong level. For how the rows are gathered in the first place, see the run guide.

Frequently asked questions

What are the three Mercado Libre link shapes?
`articulo.mercadolibre.…/MLA-…` is an individual publication, `/p/MLA…` is a catalog product page shared by several sellers, and `/up/MLAU…` is a user product. All three can appear in the same list of search results.
Why not parse the item id out of the URL?
Because a pattern that covers the two obvious shapes drops the third without any error, and the third can be 40% of a page. The remaining rows all look valid, so the loss shows up as a quietly incomplete dataset rather than as a failure.
What is item_id used for?
It is the primary key of the dataset, the cache key for the 24-hour listing cache and the deduplication key across runs. Deriving it unreliably would corrupt all three at once, which is why it comes from the page's own embedded results array.

Sources

Every URL below was requested and returned a page on the date shown.

  1. Operator claimchecked 18 Sept 2026
    Mercado Libre Scraper & API — Actor README and input schemaActorStack / Apify Store
  2. Operator claimchecked 18 Sept 2026
    Items y búsquedas — Mercado Libre DevelopersMercado Libre
A white measuring tape curving across a dark background, showing the numbers 15 to 45.
Mercado LibreExplainer

Stock and sales ranges

The site shows `+25 vendidos` and the official API answers `RANGO_1_50`. Any dataset with an exact sales integer in it invented that integer.

4 min
Aerial view of a Latin American city centre at night, streets picked out in light.
Mercado LibreGuide

How to scrape Mercado Libre

A working method for extracting listings and product pages from any of the 17 Mercado Libre marketplaces, including the two decisions — pagination and detail pages — that decide what a run costs.

9 min
A laptop screen showing a plain text-mode terminal with a command prompt.
Mercado LibreComparison

API alternative

The official Mercado Libre API requires OAuth and returns 403 to an anonymous request. A comparison of what the API gives an authorised caller, what scraping gives anyone, and which fields exist in only one of the two.

7 min