Mercado Libre serves pagination only in the form its robots.txt disallows
Measured in September 2026: the plain paginated path returns the anti-bot wall and the `_NoIndex_True` form returns results — and robots.txt disallows the second by name. There is currently no polite pagination available.
Mercado Libre's search robots.txt, served from `listado.mercadolibre.com.ar` and checked on 18 September 2026, disallows both `/*_Desde_` and `/*_NoIndex_True`. Measured in Argentina on 13 September 2026, the plain `/iphone_Desde_49` path returned the anti-bot wall on 3 of 3 attempts while `/iphone_Desde_49_NoIndex_True` served results on 3 of 3, and the same split was measured in Mexico. So the only URL form that returns page 2 is the one robots.txt names, and there is no version of polite pagination available today. The Actor does not resolve that for the operator: it counts pages fetched against the policy and pages skipped for it separately, and `respectRobots` decides which happens.
Key points
Mercado Libre's search robots.txt disallows both `/*_Desde_` and `/*_NoIndex_True`, checked on the `listado` host on 18 September 2026.
The plain paginated path returned the anti-bot wall on 3 of 3 attempts in Argentina while the disallowed form served results on 3 of 3, with the same split measured in Mexico.
Search results live on the `listado` subdomain, and its robots.txt is a different file from the one on `www` — checking the wrong host finds neither rule.
Strict robots compliance now means page 1 per query and little else, because sub-category links are served in the disallowed form too.
Pages fetched against the policy and pages skipped for it are counted in two separate run statistics, so the choice is visible after the fact rather than implied.
Sometimes the compliant option and the working option are the same URL, and sometimes a site removes that overlap. Mercado Libre has removed it.
What the robots.txt actually says
Two rules matter, and both were present when the file was checked on 18 September 2026:
listado.mercadolibre.com.ar/robots.txt — the two relevant lines
Disallow: /*_NoIndex_True
Disallow: /*_Desde_
_Desde_ is the plain pagination offset. _NoIndex_True is a second URL form of the same pages. Both are disallowed, by name.
The file is not on the host most people check
Search results live on the listado subdomain, and it serves its own robots.txt. The file on www.mercadolibre.com.ar is a different document covering the storefront, and neither pagination rule appears in it. Checking the wrong host finds nothing and concludes, wrongly, that pagination is unrestricted.
What happens on each URL form
URL form
Allowed by robots.txt
Argentina, 13 Sep 2026
URL form/iphone_Desde_49
Allowed by robots.txtNo
ResultAnti-bot wall, 3 of 3 attempts
URL form/iphone_Desde_49_NoIndex_True
Allowed by robots.txtNo
ResultResults served, 3 of 3 attempts
The same split was measured in Mexico.
There is no polite pagination today
Both forms are disallowed, and only one of them works. There is therefore no configuration that paginates and stays inside the policy. respectRobots: true stays inside the allowed paths, which now means page 1 per query and little else, because sub-category links are served in the disallowed form as well.
Two counters instead of one decision
pages_disallowed_by_robots_fetched counts pages fetched against the policy. pages_truncated_by_robots_policy counts pages skipped for it. Keeping them separate means the choice made during a run is auditable afterwards instead of being implied by a row count. If robots.txt cannot be downloaded at all, the run fails rather than claiming a compliance it could not verify.
Yes. The robots.txt served from the `listado` search host disallows `/*_Desde_`, which is the plain paginated form, and `/*_NoIndex_True`, which is the form the site currently serves results under. Both rules were present when checked on 18 September 2026.
▸Why does the www robots.txt not mention those rules?
Because search results are served from a different host. `www.mercadolibre.com.ar/robots.txt` is a separate file that covers the storefront, and neither pagination rule appears in it. The rules live on `listado.mercadolibre.com.ar/robots.txt`.
▸Can I paginate and stay within robots.txt?
Not today, in any useful sense. The only URL form that returns page 2 is the one robots.txt disallows by name, so strict compliance means page 1 per query and little else, because sub-category links are served in the disallowed form as well.
▸How do I know which my run did?
Read the run statistics. Pages fetched against the policy are counted in `pages_disallowed_by_robots_fetched` and pages skipped for it in `pages_truncated_by_robots_policy`, so the two are separable after the run rather than buried in one total.
Sources
Every URL below was requested and returned a page on the date shown.
Venezuelan product pages show a dollar price and a bolívar price at once. Deriving the implied rate from the page's structured data rather than the rendered price is the difference between a stable number and one that drifts.
A working method for extracting listings and product pages from any of the 17 Mercado Libre marketplaces, including the two decisions — pagination and detail pages — that decide what a run costs.