The short answer

Which European marketplaces let AI search engines read their product pages? Most of them. I read the robots.txt of 82 European marketplace sites and 69 could be read. Of those, 51 let the search crawlers of OpenAI, Anthropic and Perplexity into their product pages. Ten keep all three out: the nine European Amazon storefronts and leboncoin. None blocks Googlebot.

What does Amazon block? 99 user agents by name, the same list on all ten Amazon storefronts I checked. That includes the training crawlers, and also the agents that browse for a shopper: Google’s shopping agent, ChatGPT-User, Claude-User, Perplexity-User, Meta’s fetcher and three Copilot tokens. It leaves Googlebot, Bingbot, Applebot and its own Amazonbot alone.

Did Amazon change its robots.txt to stop Meta’s Muse? No. The amazon.com file archived on 15 August and the amazon.de file archived on 10 September are identical to the files live today. Meta’s crawlers were already on the list. The block that began on 20 September happened where robots.txt cannot reach.

Does it show in AI answers? In the 10,329 citations of my seven-language study, OpenAI’s search never cited Amazon. Perplexity cited it 13 times, although Amazon disallows both of Perplexity’s bots. And 25 of the 26 Amazon citations, across all engines, pointed at search results or bestseller lists rather than at a product.

Why it matters if you sell on Amazon: a listing that exists only on Amazon cannot be read by ChatGPT search, Claude search or Perplexity’s crawler. That is Amazon’s decision, not yours, and most of Europe’s other marketplaces decided the opposite.


Why now

On the night of Sunday 20 September 2026, Amazon started blocking Meta’s Muse shopping agent. Shoppers who sent it to Amazon got a warning that “continued access by an unauthorized AI agent” violates Amazon’s Conditions of Use. Amazon said Meta had not told it that Muse would access the store, and that the agent does not identify itself.

On 4 August, the Ninth Circuit vacated the preliminary injunction that had kept Perplexity’s Comet agent off Amazon. The court held that Amazon was unlikely to win its computer-fraud claims, because it is the user, not the developer, who accesses the site. The case continues.

On the other side of the argument, Allegro, Poland’s largest marketplace, launched an app inside ChatGPT on 12 May.

Almost all of the coverage is American, and most of it asserts what marketplaces allow without checking. robots.txt is the one place where each marketplace writes its policy down, so I read them.


What I checked

  • 82 European hosts across 32 marketplaces, one per country domain: Amazon, eBay, Allegro, Zalando, Kaufland, eMAG, Vinted, Otto, bol, Cdiscount, MediaMarkt, leboncoin and 20 more, from pure marketplaces to retailers with marketplace sections and classifieds sites. amazon.com and ebay.com were read as US references and are left out of every European count.
  • 22 crawler tokens from the database behind my AI crawler tool, each tagged by purpose: search index, training, user-triggered agent, or control token.
  • Two pages per site: the home page and a representative product URL. The product page is the verdict that matters to a seller.
  • The standard’s own matching rules (RFC 9309): a group naming the bot beats *, the longest matching rule wins, and Allow wins a tie.
  • Unreadable means unreadable. A 403, a bot challenge or a failed TLS handshake is reported as such, never guessed. 13 hosts ended up here: all seven Kaufland domains, eMAG in Bulgaria and Hungary, two Decathlon, Conrad, and MediaMarkt Italy.

Run on 27 September 2026. The data is published; details are at the end.


The map

No marketplace’s country domains disagree on the three AI search crawlers, so one row stands for all of a marketplace’s readable sites. Verdicts are for product pages.

Marketplace Sites ChatGPT search ChatGPT-User GPTBot (training) Claude search Perplexity search Google-Extended Amazonbot
Amazon 9 Blocked Blocked Blocked Blocked Blocked Blocked Allowed
leboncoin 1 Blocked¹ Blocked¹ Blocked¹ Blocked¹ Blocked¹ Blocked Allowed
eBay 8 Allowed Allowed Blocked Allowed Blocked Allowed Blocked
Zalando 7 Allowed Allowed Blocked Allowed Allowed Blocked Allowed
Vinted 3 Allowed Allowed Blocked Allowed Allowed Blocked Blocked
Worten 1 Allowed Allowed Blocked² Allowed Allowed Allowed Allowed
MediaMarkt 3 Allowed Allowed Allowed Allowed Allowed Allowed Blocked on DE, NL
22 others³ 37 Allowed Allowed Allowed Allowed Allowed Allowed Allowed

¹ leboncoin blocks its listing pages, under /ad/, except car ads. More below. ² Worten keeps GPTBot and CCBot off its product pages only. ³ Allegro, eMAG (Romania), ManoMano, Etsy, AliExpress, Temu, Shein, Back Market, Rakuten, Otto, bol, Cdiscount, Fnac, Galaxus, Leroy Merlin, Carrefour, Empik, Kleinanzeigen, OLX, Wallapop, Subito, Marktplaats. Leroy Merlin was tested on its home page only.

Across the 69 readable European sites:

Crawler What it is Blocked on
GPTBot OpenAI’s training crawler 29
Google-Extended Google’s control token for AI use 20
PerplexityBot Perplexity’s search index 18
Amazonbot Amazon’s own crawler 13
OAI-SearchBot ChatGPT’s search index 10
Claude-SearchBot Claude’s search index 10
ChatGPT-User ChatGPT fetching a page for a user 10
Perplexity-User Perplexity fetching a page for a user 10
Googlebot Google Search 0

The pattern is not “Europe blocks AI”. It is one company blocking nearly everything, one classifieds site closing its listings, and a large middle that made a narrower choice: keep training out, let search in.


Amazon names the agents

Amazon’s robots.txt disallows 99 user agents outright, and the list is identical on all ten storefronts I checked, amazon.com included. Most of it is what you would expect: training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended, meta-externalagent, Bytespider) and AI search indexes (OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot).

What sets it apart is the agents: the user agents of software that browses on a person’s behalf, which is what a shopping agent is.

  • Google: GoogleAgent-Shopping, GoogleAgent-Mariner, Gemini-Deep-Research, Google-NotebookLM
  • Microsoft: Copilot, CopilotNative, CopilotSapphire
  • OpenAI, Anthropic, Perplexity: ChatGPT-User, Claude-User, Perplexity-User
  • Meta: meta-externalfetcher, next to meta-externalagent and meta-webindexer
  • Others: Manus-User, MistralAI-User, Devin, TwinAgent, CloudflareBrowserRenderingCrawler

What it leaves out matters as much. Googlebot, Bingbot and Applebot fall to the general rules, and so does Amazon’s own Amazonbot. Googlebot’s index is also what Google’s AI Overviews and AI Mode draw on, and blocking Google-Extended does not take a site out of them. Copilot’s web answers are grounded in Bing. So Amazon keeps its pages in Google’s and Microsoft’s AI answers through their search crawlers, and shuts out the same companies’ agents.

Apple is the odd one out. Neither Applebot nor Applebot-Extended is named, so as far as robots.txt goes, Apple may index Amazon and use its pages for AI training.

The file did not change for Muse

Meta’s crawlers were on the list before Muse was blocked. The Wayback Machine holds a copy of amazon.com’s robots.txt from 15 August 2026 and of amazon.de’s from 10 September. Both are identical to the files served on 27 September, byte for byte.

So the Muse block was not a robots.txt change. By Amazon’s account, Muse did not identify itself, and that is the one case robots.txt cannot handle: a rule binds only a crawler that announces its name and chooses to obey. What Amazon used instead is a warning shown to the shopper, citing its Conditions of Use.

robots.txt states intent. RFC 9309, the standard that defines it, says in so many words that its rules are not a form of access authorisation. Enforcement happens somewhere else.


eBay picks sides

eBay’s file is the most deliberate in the set. One group disallows GPTBot, ClaudeBot, anthropic-ai, PerplexityBot, CCBot, Bytespider, ChatGLM-Spider, meta-externalagent, Applebot-Extended and AmazonBot, and still lets all of them read its help pages, seller portal and blog.

Read company by company:

  • OpenAI: training crawler out, search crawler and ChatGPT-User in.
  • Anthropic: the same. Training out, Claude-SearchBot and Claude-User in.
  • Perplexity: index crawler out, user-triggered fetcher in.
  • Google: Google-Extended is not named, so it is in.
  • Amazon: out.

The file is versioned in a comment (v28.1_IT_August_2026 on ebay.it), and the PerplexityBot rule was already in the April version the Wayback Machine archived on 1 May.


Training out, search in

19 of the 69 sites block OpenAI’s training crawler and allow its search crawler: eBay (8), Zalando (7), Vinted (3) and Worten (1). It is the most common deliberate policy in the set, and each of them writes it differently.

  • Zalando names six training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, AI2Bot, Bytespider) and leaves every search crawler to the general rules.
  • Vinted says it twice. It disallows five crawlers by name (GPTBot, ClaudeBot, Google-Extended, Amazonbot, CCBot), and adds a Content-Signal line for everyone else: ai-train=no, search=yes, ai-input=yes. Content-Signal is Cloudflare’s robots.txt extension for saying what content may be used for. Vinted is the only marketplace in the set that uses it.
  • Worten keeps GPTBot and CCBot off /produtos/, its product pages, and names OAI-SearchBot only to allow it.

Finer lines

leboncoin lets AI read car ads, and nothing else. Its file has labelled groups: # LLMs / AI SEARCH / RETRIEVAL bots and # LLMs / AI TRAINING. The first names OAI-SearchBot, PerplexityBot, Claude-SearchBot, Applebot, DuckAssistBot, MistralAI-Index and four user-triggered agents. They may read the site but not /ad/, where the listings live, with one exception: Allow: /ad/voitures/.

Otto keeps AI off its reviews. It names GPTBot, ChatGPT-User, OAI-SearchBot, PerplexityBot and Perplexity-User for one purpose: to disallow /kundenbewertungen/, its customer reviews. Products stay open.

MediaMarkt lets OpenAI’s agent in and keeps Amazon’s out. The German file is the longest explicit allow list in the audit: ChatGPT-User, GPTBot, OAI-SearchBot, a token named Operator (the name OpenAI gave its browser agent), five Anthropic tokens, Google-Extended, Applebot, Meta’s two agents, Manus-User and both Perplexity bots, each with Allow: /. Then it disallows Amazonbot and NovaAct, the name of Amazon’s browser-agent SDK. The Dutch file does the same. The Spanish file names Operator but neither Amazon token, the only disagreement between a marketplace’s country domains anywhere in the table above.

Amazon’s crawler gets blocked back. Amazon blocks nearly every other company’s AI crawlers. 13 European sites return the favour and block Amazonbot: eBay’s eight, Vinted’s three, and MediaMarkt in Germany and the Netherlands.

Allegro allows everything. Its file names no AI crawler at all, which fits a marketplace that built an app inside ChatGPT.


What AI search engines actually cited

robots.txt says what a site wants. My seven-language study records what AI search engines did: 10,329 citations for commercial spare-parts questions in the same markets. Matching each cited URL to the platforms audited here:

Search engine Citations To audited platforms To Amazon To eBay
OpenAI search 450 2 (0.44%), both allegro.pl 0 0
Perplexity 3,169 76 (2.40%) 13 12
Exa (Claude, Gemini) 6,710 13 (0.19%) 13 0
All 10,329 91 (0.88%) 26 12

Four things in that table.

OpenAI’s search never cited Amazon, which blocks OAI-SearchBot. That fits, but it is weak evidence: OpenAI’s search cited almost no marketplace of any kind, 2 of 450.

Perplexity cited Amazon 13 times, although Amazon disallows both of its bots. For amazon.com, timing does not explain it: the copy archived on 15 August, inside the study’s collection window, already disallowed PerplexityBot and Perplexity-User. The nearest amazon.de copy is from 10 September, after collection, and matches today’s.

eBay’s 12 are a split case. Three were a seller-help page under /spaziovenditori, a path eBay explicitly opens to PerplexityBot. Six were category pages on ebay.it, which its file closes to PerplexityBot. The last three were on eBay’s Dominican Republic site, which I did not audit.

This data cannot show how Perplexity got the closed pages. It could be an index built before the rules, a third-party search provider, or fetches that do not announce themselves as PerplexityBot. Cloudflare accused Perplexity of undeclared crawling in August 2025, and Perplexity disputed it. Nothing here tells those explanations apart, and I am not claiming any of them.

Exa’s 13 Amazon citations break no rule. Exa’s crawler is not named in Amazon’s file, and the general group explicitly allows /-/es/, the Spanish-language section of amazon.com where all 13 sat, 12 of them on one bestseller page.

Almost nothing cited was a product. Of the 38 citations to Amazon and eBay, one pointed at a product page. The rest were search results, bestseller lists, category pages and eBay’s seller help. When AI search cites a marketplace, it cites the shelf, not the listing.


What this means if you sell on marketplaces

  1. An Amazon-only listing is invisible to the AI search crawlers that respect robots.txt. ChatGPT search, Claude search and Perplexity’s crawler are all disallowed. Google’s AI features, Copilot and Apple are not, because Amazon lets their search crawlers in. That split is Amazon’s choice, and nothing in Seller Central changes it.
  2. The rest of Europe decided the other way. Allegro in Poland, Czechia, Slovakia and Hungary, bol in the Netherlands, Otto in Germany, Cdiscount and Fnac in France, ManoMano in five markets, and eBay everywhere except for Perplexity’s index crawler. If AI visibility matters in your category, where you list is part of it now, market by market.
  3. Do not count on a marketplace listing being the page that gets cited. Even the open marketplaces took under 1% of citations in my data, and what got cited was category and search pages. Your own site, and the publishers and specialists who write about your category, take the rest.
  4. Consider the middle way for your own site. Training crawlers out, search crawlers and user agents in: what Zalando, eBay and Vinted chose. The AI crawler tool builds that robots.txt token by token.
  5. Decide per agent, not per company. One company can run a training crawler, a search crawler and a user agent under three names, and eBay treats OpenAI’s three differently. Amazon’s list is the most complete inventory of shopping-relevant agents I have seen. Read it before you write your own.
  6. Check that crawlers can read your robots.txt at all. Kaufland, Decathlon, Conrad and two of eMAG’s three sites answered my requests for robots.txt with a bot challenge. From a datacenter that may be deliberate. But if your bot management does the same to AI crawlers, they never see your policy, and the one you wrote does nothing.

Limits, stated plainly

  • Stated intent only. A site can block at the CDN whatever its robots.txt says, and user-triggered agents may not read robots.txt at all. By Amazon’s account, Muse did not.
  • A datacenter view. Fetches came from a cloud IP. 13 of 82 European hosts were unreadable from there, which says nothing certain about what a real crawler receives.
  • One product URL per site. Rules are path-based, and a site can treat categories differently, as leboncoin does. Leroy Merlin was tested on its home page only.
  • The tables cover 22 tokens. Agent names outside the crawler database, such as GoogleAgent-Shopping or NovaAct, come from reading the files, not from the matcher.
  • Timing. The study’s citations were collected on or before 16 August, and the audit ran on 27 September. amazon.com matches its 15 August archive, and eBay’s PerplexityBot rule dates back to at least May. The other sites were not checked against archives.
  • Small numbers, one category. 26 Amazon citations, all about appliance spare parts. They show that the citations happened, not how often they happen in general.
  • Reported, not verified by me. The Muse block comes from press reports, linked where it appears. The Ninth Circuit ruling links to the opinion itself and the Allegro launch to Allegro’s own announcement.

FAQ

Does Amazon block ChatGPT? In robots.txt, yes. Every Amazon storefront disallows OAI-SearchBot, GPTBot and ChatGPT-User: OpenAI’s search crawler, its training crawler and its user-triggered agent. In my citation data, OpenAI’s search never cited Amazon.

Why did Amazon block Meta’s Muse? According to press reports, Amazon said Meta did not disclose that Muse would access its store, and that the agent does not identify itself. Amazon’s robots.txt already disallowed Meta’s crawlers before the block and did not change around it. The block itself is a warning citing Amazon’s Conditions of Use.

Which marketplaces allow AI search engines? Of 69 readable European marketplace sites, 51 let the search crawlers of OpenAI, Anthropic and Perplexity read product pages, among them Allegro, Zalando, Vinted, Otto, bol, Cdiscount, Fnac, MediaMarkt and ManoMano. eBay allows OpenAI’s and Anthropic’s, but not Perplexity’s.

Does eBay block AI crawlers? It blocks the training crawlers of OpenAI, Anthropic and others, Perplexity’s index crawler and Amazonbot. It allows ChatGPT and Claude search and their user-triggered agents.

Can my Amazon products appear in ChatGPT? Not from Amazon’s own pages, if ChatGPT’s search follows Amazon’s robots.txt, which disallows it. ChatGPT can still find your products on your own site, on marketplaces that allow its crawler, and in articles that mention them.

Is robots.txt legally binding? No. RFC 9309 says its rules are not a form of access authorisation. robots.txt states what a site wants, and compliance is voluntary.

Can I check these numbers? Yes. Every verdict is published with the robots.txt it was read from, so each one can be checked against the file itself.


The takeaway

The story is not that European marketplaces block AI. Most do not. It is that the largest one does, completely, down to naming each company’s shopping agent, while keeping the search crawlers that feed Google’s and Microsoft’s AI answers. For a seller, choosing a marketplace now also means choosing which AI assistants can see you.


Method and data

The audit fetches each site’s robots.txt, applies RFC 9309 to 22 crawler tokens for the home page and a product page, and cross-checks the seven-language citation dataset. It writes one row per site and crawler, and keeps every robots.txt exactly as fetched, so each verdict can be checked against the file it came from.

The run behind this article is published as robots.csv, with the files as fetched in robots-txt.json. CC BY 4.0.


Emad Sharaki is a Senior SEO & GEO Strategist with 12+ years in search. He leads AI Visibility work at AUTODOC SE, tracking brand representation across six major language models and 35+ European markets, and maintains an enterprise database of 38+ verified AI crawlers used to govern crawler policy across seven European markets. He has spoken at WordCamp Porto, Porto WordPress Meetup and WordPress Day for E-Commerce, and has been quoted in Search Engine Journal. If you are starting on this, there is a learning path for AI search.