Glossary

Meta-ExternalAgent Meta's crawler for AI training and product indexing.

  • Noun
  • Crawler
  • Owned by Meta

The Meta crawler that collects pages for AI model training and product indexing, and not the bot Meta ties to citations in Meta AI answers. Prefer's free Crawler Access Checker confirms your search crawlers stay open while you decide on this one.

Updated1 Oct 2026Reading time6 min
Definition28 words

Meta-ExternalAgent is Meta's web crawler for training AI models and indexing content, controlled in robots.txt by the token meta-externalagent. Prefer's free generator writes your other AI bot rules.

Also writtenmeta-externalagentMeta external agentmeta-externalagent/1.1
Key facts ↓

Key facts

At a glance.

Meta-ExternalAgent at a glance

The entity facts an assistant lifts first, each with a checked-on date.

Checked1 Oct 2026
Operator
Meta · meta.com (docs on developers.facebook.com)
Type
Web crawler, AI model training and product indexing
User agent token
meta-externalagent
Full UA string
meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers) or meta-externalagent/1.1, as Meta publishes them (the link is relative on Meta's page)
Respects robots.txt
Yes. Changes can take up to 24 hours, as robots.txt may be cached
IP ranges
Not documented
First announced
Not documented (Meta's page shows 'Updated: May 21, 2026')
Affects Meta AI citations
Not documented. Meta ties citations to Meta-WebIndexer

How it works

What happens when Meta-ExternalAgent visits.

Three steps. The third is the one to read twice: one token covers two different uses, so you cannot split them.

  1. 01

    Reads robots.txt

    The crawler checks robots.txt for a meta-externalagent rule. Meta says its crawlers may cache the file for up to 24 hours, so a new rule can take that long to apply.

  2. 02

    Fetches pages it is allowed to

    Pages not disallowed for the token can be fetched. Meta's crawler page does not list any further filtering, such as skipping logged-in or paywalled pages.

  3. 03

    Uses the content for training or indexing

    Meta names two uses: training foundation AI models, and improving products by indexing content directly. It does not say which pages go to which use.

NoteBecause both uses sit behind one token, a Disallow for meta-externalagent opts you out of both. Meta documents a separate bot, Meta-WebIndexer, for Meta AI search results.

The Meta crawler family

Meta-ExternalAgent is one of several. They do different jobs.

Meta documents five crawlers on one page. Each is named separately in robots.txt, so each can be allowed or blocked on its own.

Meta's user agents

Purpose, what blocking it changes, and whether Meta ties it to Meta AI citations.

Checked1 Oct 2026
BotPurposeBlocking it meansTied to Meta AI citations
Meta-ExternalAgentThis term Crawls for training foundation AI models and for indexing content to improve products. It stops fetching your pages for both uses, within about 24 hours. ✕ Not documented
Meta-WebIndexer Navigates the web to improve Meta AI search result quality for users. Meta says allowing it helps Meta cite and link your content in Meta AI, so blocking works against that. ✓ Yes
Meta-ExternalFetcher Fetches individual links at a user's request, including for agentic AI tasks. A Disallow may not stop it. Meta says it may bypass robots.txt. ✕ Not documented
Meta-ExternalAds Improves advertising and other business-related products and services. Not documented on Meta's crawler page. ✕ Not documented
FacebookExternalHit Fetches pages to build link-share previews. Meta says it might bypass robots.txt when performing security or integrity checks. ✕ Not documented
BasisNames and purposes as documented by Meta at developers.facebook.com/docs/sharing/webmasters/web-crawlers on the checked-on date. Meta's robots.txt example gives the lowercase tokens meta-externalagent and meta-externalfetcher; the other bots are shown as Meta names them.

Why it matters for AEO

Allow or block: the decision in one sentence.

Should you block Meta-ExternalAgent?

For most brands that want AI visibility, allow it, and use Prefer's free Crawler Access Checker to confirm your search crawlers stay open whatever you choose here. Meta-ExternalAgent feeds model training and Meta's product indexing, so blocking it trades away how Meta's models learn your brand. Block it when the content is the product you sell, and keep Meta-WebIndexer allowed, since that is the bot Meta ties to Meta AI citations.

Allow Meta-ExternalAgent if
  • You want Meta's models to learn your brand, product names and positioning
  • Your content is marketing, documentation or editorial you already give away
  • You compete in a category where assistants are shaping buyer shortlists
Block Meta-ExternalAgent if
  • The content is the product: paid research, licensed data, a subscription archive
  • Legal or licensing terms forbid model training on the material
  • You block it by path, not by site, so public pages stay available

Meta does not document what blocking Meta-ExternalAgent does to Meta AI answers, so this guidance rests on what Meta says the bot is for. The robots.txt patterns below implement each side.

Allow or block it

Three robots.txt patterns that cover most cases.

Each pattern names the token explicitly. A wildcard block catches Meta-ExternalAgent too, which is how sites block it by accident.

robots.txt

Copy the pattern that matches your decision above.

# Opt out of Meta's training and product indexing crawler
User-agent: meta-externalagent
Disallow: /

# Meta ties Meta AI citations to this bot. robots.txt
# matches user-agent names case-insensitively (RFC 9309).
User-agent: Meta-WebIndexer
Allow: /
Block training, keep Meta AI search

Out of Meta's training and product indexing, still open to the bot Meta ties to Meta AI citations.

  • Check for a wildcard firstA User-agent: * group with Disallow: / already blocks Meta-ExternalAgent. A named group overrides it only for the named token.
  • Allow up to 24 hoursMeta says its crawlers may cache robots.txt for up to 24 hours, so a new rule is not instant.
  • One token, one botA rule for meta-externalagent does not cover meta-externalfetcher, and Meta says that fetcher may bypass robots.txt anyway.

Verify a visit

How to confirm it was really Meta-ExternalAgent.

Any script can claim this user agent, and Meta documents fewer checks than most operators. Here is what is and is not available.

  1. 01

    Match the user agent

    Look for meta-externalagent/1.1 in the request's user agent string. Meta says real requests look similar to the published strings.

  2. 02

    No published IP list

    Meta's page suggests allow-listing by IP address as the more secure option, but publishes no IP list or URL at the checked-on date. Without one, a log line alone cannot prove a request is genuine.

203.0.113.42 - - [01/Oct/2026:09:14:07 +0000] "GET /pricing/ HTTP/1.1" 200 18422 "-" "meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)"
An illustrative access-log line using the user agent exactly as Meta publishes it, relative link included. With no published IP list, the user agent is the only documented signal.

In context

The term in a sentence.

01

"Our robots.txt blocks meta-externalagent for the paid research archive only, so Meta's models can still learn from the public blog."

02

"We allow Meta-WebIndexer and block Meta-ExternalAgent, because Meta ties Meta AI citations to the indexer, not the training crawler."

Related questions

People also ask

The questions buyers ask next, taken from what assistants cluster with this one.

Does blocking Meta-ExternalAgent stop Meta AI from citing my site?

Meta does not say. Prefer tracks citations on ChatGPT, Gemini, Perplexity, Google AI Overviews and AI Mode rather than Meta AI, so go by Meta's page: it says allowing Meta-WebIndexer helps Meta cite and link your content in Meta AI's responses. To keep that path open, leave Meta-WebIndexer allowed.

See the crawler comparison →
What is the Meta-ExternalAgent user agent string?

Meta publishes two forms, meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers) and meta-externalagent/1.1, and says real requests look similar. Prefer quotes them exactly as published at the checked-on date; the robots.txt token is meta-externalagent.

Does Meta publish IP ranges for Meta-ExternalAgent?

No. Prefer checked Meta's crawler page on 1 Oct 2026 and found a suggestion to allow-list by IP address but no IP list or URL. Until Meta publishes one, the user agent string is the only documented way to identify the bot.

Is Meta-ExternalAgent the same as Meta AI?

No, and Prefer keeps the two apart for every operator it profiles. Meta AI is the assistant; Meta-ExternalAgent is a crawler that gathers web content for training models and indexing, while Meta ties Meta AI citations to a different bot, Meta-WebIndexer.

Questions

Asked plainly.

What does Meta-ExternalAgent do?

Meta says the Meta-ExternalAgent crawler 'crawls the web for use cases such as training foundation AI models or improving products by indexing content directly.' Prefer's free robots.txt generator covers 14 other named AI crawlers, so add a group for the token meta-externalagent by hand to decide whether it can reach your pages.

How do I block or allow Meta-ExternalAgent?

Add 'User-agent: meta-externalagent' followed by 'Disallow: /' to robots.txt to block it, or 'Allow: /' to let it in. Prefer's free robots.txt generator builds the groups for the other major AI crawlers, so you can paste this one alongside them. Meta says to allow up to 24 hours for the change, because its crawlers may cache robots.txt for that long.

Does blocking Meta-ExternalAgent remove me from Meta AI answers?

Meta does not say. Prefer tracks citations on ChatGPT, Gemini, Perplexity, Google AI Overviews and AI Mode, not Meta AI, so the answer here comes from Meta's own page: it ties citing and linking in Meta AI responses to a different bot, Meta-WebIndexer, and documents Meta-ExternalAgent only as a training and indexing crawler.

What is the difference between Meta-ExternalAgent and Meta-ExternalFetcher?

Meta-ExternalAgent crawls the web for training and indexing and carries no robots.txt bypass caveat; Meta-ExternalFetcher fetches individual links at a user's request and, Meta says, may bypass robots.txt. Prefer recommends naming each token in its own robots.txt group, because a rule for one does not cover the other.

Sources

Where these facts come from.

Every fact on this page traces to one of these. Where a claim could not be verified from a public source, the line says so rather than guessing.

  1. 01 Meta, web crawlers documentation Meta's page for Meta-ExternalAgent, Meta-ExternalFetcher, Meta-WebIndexer, Meta-ExternalAds and FacebookExternalHit, including user agents, purposes and robots.txt behaviour. Vendor docs 1 Oct 2026
  2. 02 Meta, web crawlers documentation (Markdown view) The same page in Meta's own Markdown view, used to cross-check the user agent strings verbatim. Vendor docs 1 Oct 2026
  3. 03 Robots Exclusion Protocol, RFC 9309 The standard that defines how user-agent tokens, Allow and Disallow rules are interpreted. Standard 1 Oct 2026

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.