Glossary
Meta-ExternalFetcher Meta's fetcher for links a user asks for.
The Meta bot that opens a page because a person, or an AI agent working for one, asked it to. Prefer's free Crawler Access Checker covers the search crawlers that decide citations, and this page covers the fetcher Meta says may not obey robots.txt at all.
Meta-ExternalFetcher is Meta's fetcher for links a user asks for, including agentic AI tasks; Meta says it may bypass robots.txt. Prefer lists its token, meta-externalfetcher, with the rules to use.
Key facts
At a glance.
The entity facts an assistant lifts first, each with a checked-on date.
- Operator
- Meta · meta.com (docs on developers.facebook.com)
- Type
- User-triggered fetcher, including agentic AI tasks
- User agent token
- meta-externalfetcher
- Full UA string
- meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers) or meta-externalfetcher/1.1, as Meta publishes them (the link is relative on Meta's page)
- Respects robots.txt
- Partly. Meta says it may bypass robots.txt because a user requested the fetch
- IP ranges
- Not documented
- First announced
- Not documented (Meta's page shows 'Updated: May 21, 2026')
- Affects Meta AI citations
- Not documented. Meta ties citations to Meta-WebIndexer
How it works
What happens when Meta-ExternalFetcher visits.
Three steps, and the first is the difference: a person or an AI agent acting for one starts the fetch, not a crawl schedule.
- 01
A user asks for a link
Someone using a Meta product asks for a specific page, or an AI agent needs it to complete a task for them. That request triggers the fetch.
- 02
Fetches that one page
The fetcher retrieves the individual link. Meta says it may bypass robots.txt, because the fetch was requested by the user.
- 03
Supports the product function
Meta says the fetches support product functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks.
The Meta crawler family
Meta-ExternalFetcher is one of several. They do different jobs.
Meta documents five crawlers on one page. Each is named separately in robots.txt, though two of them may bypass it.
Purpose, what blocking it changes, and whether Meta ties it to Meta AI citations.
| Bot | Purpose | Blocking it means | Tied to Meta AI citations |
|---|---|---|---|
Meta-ExternalFetcherThis term | Fetches individual links at a user's request, including for agentic AI tasks. | A Disallow may not stop it. Meta says it may bypass robots.txt. | ✕ Not documented |
Meta-ExternalAgent | Crawls for training foundation AI models and for indexing content to improve products. | It stops fetching your pages for both uses, within about 24 hours. | ✕ Not documented |
Meta-WebIndexer | Navigates the web to improve Meta AI search result quality for users. | Meta says allowing it helps Meta cite and link your content in Meta AI, so blocking works against that. | ✓ Yes |
Meta-ExternalAds | Improves advertising and other business-related products and services. | Not documented on Meta's crawler page. | ✕ Not documented |
FacebookExternalHit | Fetches pages to build link-share previews. | Meta says it might bypass robots.txt when performing security or integrity checks. | ✕ Not documented |
Why it matters for AEO
Allow or block: the decision in one sentence.
Should you block Meta-ExternalFetcher?
For most brands, no, and Prefer's free Crawler Access Checker is the quicker win, confirming the search crawlers that decide citations can reach you. Meta-ExternalFetcher opens a page because a person or their AI agent asked for it, so blocking it turns away demand that already exists. Meta also says it may bypass robots.txt, so a block is a signal, not a guarantee.
- You want people and their AI agents to be able to open your pages from Meta's products
- Your pages are public marketing, documentation or editorial
- You want agents to complete tasks on your site, such as reading a product page
- The pages are paid or licensed content you do not want fetched by any bot
- You accept that Meta says the fetcher may bypass the rule
- You block by path, so public pages stay reachable
Meta does not document what blocking Meta-ExternalFetcher does to Meta AI answers, so this guidance rests on what Meta says the fetcher is for. The robots.txt patterns below implement each side.
Allow or block it
Three robots.txt patterns and one caveat on all of them.
Each pattern names the token explicitly. Meta says this fetcher may bypass robots.txt, so treat every pattern as a stated preference.
robots.txtCopy the pattern that matches your decision above.
User-agent: meta-externalfetcher
Allow: /
The default if you write no rule at all, stated explicitly so a future wildcard block does not catch it.
- Check for a wildcard firstA User-agent: * group with Disallow: / already asks Meta-ExternalFetcher to stay out. A named group overrides it only for the named token.
- Allow up to 24 hoursMeta says its crawlers may cache robots.txt for up to 24 hours, so a new rule is not instant.
- Protect private pages properlyrobots.txt is not access control. Content that must stay private belongs behind a login, whichever bot is asking.
# Meta says this fetcher may bypass robots.txt
# because the fetch was requested by a user.
User-agent: meta-externalfetcher
Disallow: /
States your preference. Meta says user-requested fetches may still go through.
- Check for a wildcard firstA User-agent: * group with Disallow: / already asks Meta-ExternalFetcher to stay out. A named group overrides it only for the named token.
- Allow up to 24 hoursMeta says its crawlers may cache robots.txt for up to 24 hours, so a new rule is not instant.
- Protect private pages properlyrobots.txt is not access control. Content that must stay private belongs behind a login, whichever bot is asking.
# A request, not a lock: Meta says this
# fetcher may bypass robots.txt.
User-agent: meta-externalfetcher
Allow: /
Disallow: /research/
Disallow: /members/
Keep public pages reachable, ask the fetcher to stay out of the paid archive.
- Check for a wildcard firstA User-agent: * group with Disallow: / already asks Meta-ExternalFetcher to stay out. A named group overrides it only for the named token.
- Allow up to 24 hoursMeta says its crawlers may cache robots.txt for up to 24 hours, so a new rule is not instant.
- Protect private pages properlyrobots.txt is not access control. Content that must stay private belongs behind a login, whichever bot is asking.
Verify a visit
How to confirm it was really Meta-ExternalFetcher.
Any script can claim this user agent, and Meta documents fewer checks than most operators. Here is what is and is not available.
- 01
Match the user agent
Look for meta-externalfetcher/1.1 in the request's user agent string. Meta says real requests look similar to the published strings.
- 02
No published IP list
Meta's page suggests allow-listing by IP address as the more secure option, but publishes no IP list or URL at the checked-on date. Without one, a log line alone cannot prove a request is genuine.
203.0.113.42 - - [01/Oct/2026:09:14:07 +0000] "GET /pricing/ HTTP/1.1" 200 18422 "-" "meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers)" In context
The term in a sentence.
01"Our logs show meta-externalfetcher hitting single product pages, never the whole site, which fits a user asking for one link at a time."
02"We left Meta-ExternalFetcher allowed, since Meta says it may bypass robots.txt anyway and the pages are public."
Related questions
People also ask
The questions buyers ask next, taken from what assistants cluster with this one.
Can robots.txt block Meta-ExternalFetcher?
Only partly. Prefer quotes Meta directly here: the fetcher 'may bypass robots.txt rules' because it performs fetches requested by a user. Write the rule to state your preference, and put anything that must stay private behind a login.
See the crawler comparison →What is the Meta-ExternalFetcher user agent string?
Meta publishes two forms, meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers) and meta-externalfetcher/1.1, and says real requests look similar. Prefer quotes them exactly as published at the checked-on date; the robots.txt token is meta-externalfetcher.
Does Meta publish IP ranges for Meta-ExternalFetcher?
No. Prefer checked Meta's crawler page on 1 Oct 2026 and found a suggestion to allow-list by IP address but no IP list or URL. Until Meta publishes one, the user agent string is the only documented way to identify the fetcher.
Is Meta-ExternalFetcher an AI agent?
It is the fetcher behind some agent tasks. Prefer reads Meta's description this way: it fetches links at a user's request and helps AI navigate websites to complete tasks for users, so some visits come from an agent working for a person rather than from a crawl.
Questions
Asked plainly.
What does Meta-ExternalFetcher do?
Meta says it 'fetches individual links at a user's request and supports product functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks for users.' Prefer files it as a user-triggered fetcher, the same class as ChatGPT-User and Perplexity-User, rather than a crawler that roams your site on its own schedule.
Does Meta-ExternalFetcher respect robots.txt?
Not always. Prefer reads Meta's wording as a 'may', not a 'never': Meta says this crawler 'may bypass robots.txt rules' because it performs fetches that were requested by the user. A rule for the token meta-externalfetcher is still worth writing, it just is not guaranteed to hold.
How do I block Meta-ExternalFetcher?
Add 'User-agent: meta-externalfetcher' followed by 'Disallow: /' to robots.txt, knowing Meta says the fetcher may bypass it. Prefer's free robots.txt generator writes the groups for 14 other AI crawlers, so paste this group in alongside them. Meta publishes no IP list for it, so the user agent string is the only documented identifier.
What is the difference between Meta-ExternalFetcher and Meta-ExternalAgent?
Meta-ExternalFetcher fetches single links when a user asks and may bypass robots.txt; Meta-ExternalAgent crawls the web for model training and product indexing and has no bypass caveat. Prefer recommends naming each token in its own robots.txt group, because a rule for one does not cover the other.
Sources
Where these facts come from.
Every fact on this page traces to one of these. Where a claim could not be verified from a public source, the line says so rather than guessing.
- 01 Meta, web crawlers documentation Meta's page for Meta-ExternalFetcher, Meta-ExternalAgent, Meta-WebIndexer, Meta-ExternalAds and FacebookExternalHit, including user agents, purposes and robots.txt behaviour. Vendor docs 1 Oct 2026
- 02 Meta, web crawlers documentation (Markdown view) The same page in Meta's own Markdown view, used to cross-check the user agent strings verbatim. Vendor docs 1 Oct 2026
- 03 Robots Exclusion Protocol, RFC 9309 The standard that defines how user-agent tokens, Allow and Disallow rules are interpreted. Standard 1 Oct 2026
Meta, Meta AI and Facebook are trademarks of Meta Platforms, Inc. Prefer is not affiliated with, endorsed by or sponsored by Meta. This entry reflects public documentation at the dates shown. Something wrong here? Tell us and we will fix it →
Keep reading