On-Page AEO: Make Your Pages Answer-Ready

Six on-page fixes that make AI engines read, understand, and cite your pages: answer-first blocks, question headings, schema, entity clarity, crawlers.

Javed Khatri Javed Khatri Co-founder, Prefer

13 min read Chapter 05

The short answer

How do I optimize my website for AI search?

On-page AEO is the cheapest and fastest lever, though not the biggest: off-page won more citations in our study. Six fixes make your pages answer-ready: lead with a self-contained answer, write headings as questions, add the schema AI reads, become one clean entity, fix your internal links, and let the AI crawlers in. Do it once, do it well, then spend most ongoing effort off-page.

Key takeaways

  • On-page AEO is the cheapest and fastest lever, but not the biggest: off-page sources won more citations in our study, so do on-page once and well, then spend most ongoing effort off-page.
  • Lead every money page with a self-contained answer, because models lift claims that stand alone without the surrounding page for context.
  • robots.txt is the one file every major AI crawler honors; a default block makes even a perfect page invisible to live search.
  • On-page work reaches grounded (live-search) answers, not memory answers, which only change as your off-page reputation does.

By the end of this chapter you will have

Your ranked on-page fix list: the specific pages to rewrite answer-first, the schema to add, the entity to clean up, the links to fix, and the crawler check to run first. Plays 10 to 15.

On-page AEO is the cheapest lever in this playbook, and the one most people over-invest in. It is worth doing, and doing well, because a page a model cannot read or understand cannot be cited. But it is not the biggest lever: in our study of 1,237 citations, third-party sources won far more citations than vendors’ own sites. So the honest instruction for this chapter is: do these six fixes once, do them properly, then move most of your ongoing effort to the off-page chapter.

There is a second limit worth naming up front. On-page work reaches grounded answers, the ones where the engine actually reads live pages. It does very little for memory answers, which change only as your off-page reputation does (chapter 2 covers that split). On-page is how you make sure that when a model does read your page, it can find the answer, understand it, and know who you are.

The one move that matters most: lead with the answer#

If you do only one thing in this chapter, do this. Models lift claims that stand on their own. A page that buries its answer under three paragraphs of preamble gives the model nothing clean to quote; a page that opens with a self-contained answer hands it a citation.

Two versions of the same page answering "What is expense management software?". Left, answer buried: the page opens with generic filler and the real answer sits deep in the page, hard for a model to lift and often skipped. Right, answer first: the page opens with a bold self-contained answer ("Expense management software automates how a company tracks, approves, and reimburses employee spending."), which an engine can cite cleanly.
The same page, two ways. The version that opens with a self-contained answer is the one a model can lift cleanly. Illustrative example.

The test is simple: copy your opening block, paste it somewhere with no other context, and read it. If it still makes sense and names its own subject (not “it” or “this tool”), it is liftable. If it only makes sense after the paragraph above it, rewrite it.

The six fixes#

Here are the six on-page plays. Numbered 10 to 15, but start with the crawler check in Play 15: it takes ten minutes and there is no point fixing a page the engines are blocked from reading.

A page mockup with 6 numbered, annotated blocks. 1. Answer-first block: Open with a self-contained answer a model can lift. (Play 10) 2. Heading written as a question: Match how buyers actually ask. (Play 11) 3. Structured data: FAQPage, Organization, and HowTo schema. (Play 12) 4. Descriptive internal links: To pricing, a competitor comparison, and a guide. (Play 14) 5. Entity line: Brand name, one-sentence definition, and sameAs links to profiles. (Play 13) 6. Crawlers allowed: AI bots permitted in robots.txt. (Play 15) The browser chrome shows a crawlers-allowed marker.
The six on-page fixes, mapped onto one page. Each is a play below; together they make a page a model can read, parse, and attribute.
Play 10
AnyStart here

Rewrite your top pages answer-first

Put a self-contained answer at the top of every money page, so a model can lift it without reading the rest.

Why it works

Peer-reviewed GEO research found that making content more quotable and well-structured raised how often generative engines cited it, with fluency and quotation optimizations each lifting visibility by roughly a quarter (up to 28%). Self-contained answers are the practical version of that. (Aggarwal et al., KDD 2024, checked 2026-07-14)

Steps
  1. List your top 5 money pages (the ones tied to your money prompts).
  2. For each, write a 2 to 3 sentence answer to the page's core question and put it at the very top, before any preamble.
  3. Write it to stand alone: name the subject explicitly, no 'it' or 'this tool' that refers to something above.
  4. Bold the single most citable sentence so it survives extraction on its own.
  5. Keep the depth below the answer; the answer block earns the citation, the detail earns the reader.
Tools Your CMS. Free
Effort About 30 minutes per page (estimate, not measured)
Time to impact Grounded engines can pick it up after the next crawl (weeks, estimate)

Done when: Your top 5 money pages each open with a self-contained answer block.

Verify it worked: Copy each opening block, paste it with no context, and confirm it still stands alone. Then re-ask that page's money prompt monthly and watch for the lift.

Common failure mode: Opening with brand throat-clearing ('At Acme, we believe...') before the answer. The model quotes the answer, not the preamble, so put the answer first.

Play 11
Any

Write your headings as the questions buyers ask

Match your H2s to the natural-language questions people ask AI, so your sections map directly onto prompts.

Why it works

Buyers ask assistants full questions, and models match content to the question. Our study's winning pages answered those questions directly; question-shaped headings with a direct answer underneath are the on-page shape of that, and the same Q-and-A structure is why Reddit gets cited so heavily. (Prefer AI Citation Study, checked 2026-07-10)

Steps
  1. Pull the real questions from your 10 money prompts and from the engines' 'people also ask' boxes.
  2. Rewrite your money-page H2s as those questions, in the words buyers actually use.
  3. Answer each question in the first sentence under its heading, self-contained.
  4. Add an FAQ block of at least 3 verbatim buyer questions with direct answers, and mark it up with FAQPage schema (Play 12).
Tools Your money prompts and the engines' related-questions boxes. Free
Effort About 1 hour per page (estimate)
Time to impact Weeks, after re-crawl (estimate)

Done when: Your money pages use question-shaped H2s and carry a 3+ question FAQ block.

Verify it worked: Ask one of your H2 questions in a grounded engine and see whether your section is the kind of thing it lifts.

Common failure mode: Clever headings that match no query ('The Acme difference'). A model searching a buyer's question never matches marketing-speak.

Play 12
Any

Add the schema types AI reads

Mark up your pages so an engine can tell which block is a definition, which is a step-by-step, and which is a question and answer.

Why it works

Schema makes your content unambiguous to parse, so the right block gets lifted cleanly. It is clarity, not a guaranteed ranking: its value is that the model does not have to guess what each part of your page is.

Steps
  1. Add Organization schema sitewide, with sameAs links to your off-page profiles (ties to Play 13).
  2. Add FAQPage to your question-and-answer blocks, HowTo to your step lists, and DefinedTerm to your definitions.
  3. Add BreadcrumbList to every page so the model understands your site structure.
  4. Keep every field consistent with the visible text; never mark up something that is not on the page.
  5. Validate before you ship.
Tools Your CMS or a schema generator; validator.schema.org and Google's Rich Results Test
Effort A few hours sitewide (estimate)
Time to impact Immediate once re-crawled (estimate)

Done when: The five schema types validate with no errors on your money pages.

Verify it worked: Run each page through the Rich Results Test and validator.schema.org; both should pass clean.

Common failure mode: Schema that contradicts the visible page. Models and search engines both discount mismatched markup, and it can cost you trust.

Play 13
Any

Make your brand one clean entity

Make it unambiguous who you are, so engines resolve you correctly, especially in the memory answers on-page fixes otherwise cannot reach.

Why it works

ChatGPT answered 29 of 47 questions from memory in our study, and a model can only describe you from memory if it has a stable, consistent picture of who you are. One name and one definition, repeated everywhere and linked to your off-page records, is how that picture forms. (Prefer AI Citation Study, checked 2026-07-10)

Before you start: Ideally your entity records from Play 19, so the sameAs links have somewhere to point.

Steps
  1. Pick one canonical brand name and use it identically everywhere (no stray variations).
  2. Write one one-sentence definition of what you are and what category you are in.
  3. Put that definition on your homepage and about page, in the same words.
  4. Wire Organization schema sameAs to your Wikidata, Crunchbase, and LinkedIn records (Play 19).
  5. Keep the name and definition frozen; consistency over time is what builds the entity.
Tools Your CMS and your Play 19 profiles. Free
Effort About 2 hours (estimate)
Time to impact Slow: memory-side recognition builds over months (estimate)

Done when: Your name and definition are identical on-page and across your off-page profiles, with sameAs wired.

Verify it worked: Ask an engine 'What is [your brand]?' monthly and check it describes the real company, in terms close to your definition.

Common failure mode: A different tagline and description on every page. If the model sees five versions of what you are, it forms none of them.

Play 14
Any

Fix your internal links

Make every money page reachable and connected, so crawlers can find it and models can relate it to the rest of your site.

Why it works

A page nothing links to often does not get crawled at all, so it cannot be cited. We learned this the hard way (see below): a single navigation change once orphaned our entire content surface, and search crawled almost none of it.

Steps
  1. List your money pages and check each is within about two clicks of the homepage.
  2. Add descriptive contextual links between related pages (the anchor text should say what the target is, never 'click here').
  3. Link every page to at least one conversion page, so the reader always has a next step.
  4. Run an internal-link crawl and fix every orphaned money page.
Tools Your CMS; a crawler like Screaming Frog's free tier
Effort About half a day (estimate)
Time to impact Weeks, as the newly linked pages get crawled (estimate)

Done when: No money page is orphaned, and each sits within about two clicks of the homepage.

Verify it worked: Crawl the site and confirm zero orphaned money pages.

Common failure mode: Publishing pages nothing links to. Orphans are invisible to crawlers, and a page that is never crawled is never cited.

Play 15
AnyDo this first

Let the AI crawlers in

Make sure the engines are actually allowed to read your pages. This is the ten-minute check to run before anything else on this list.

Why it works

robots.txt is the one file all the major AI operators honor, and some content platforms and CDNs block the AI crawlers by default. A blocked bot cannot read a page, so every other on-page fix is wasted until the crawler is allowed in.

Steps
  1. Open yoursite.com/robots.txt and read it.
  2. Confirm none of these are disallowed: GPTBot, OAI-SearchBot (ChatGPT), ClaudeBot (Claude), PerplexityBot (Perplexity), Google-Extended (Gemini and AI Overviews).
  3. If you are unsure, add an explicit Allow for each of those user-agents.
  4. Publish an llms.txt too, as cheap hygiene, but expect little from it (see the honest note below).
  5. Re-check robots.txt after any platform migration or CDN change; those often reset it.
Tools Your robots.txt. Free
Effort About 15 minutes (estimate)
Time to impact Immediate: it unblocks everything else (estimate)

Done when: The five AI crawlers are allowed in robots.txt.

Verify it worked: Fetch yoursite.com/robots.txt and confirm there is no Disallow covering the AI user-agents.

Common failure mode: Leaving a blanket Disallow or a default CMS block in place. It is the most common way a good page stays invisible, and the easiest to miss.

A robots.txt allow-list permitting 5 AI crawlers, each with an Allow check: GPTBot (ChatGPT training), OAI-SearchBot (ChatGPT search), ClaudeBot (Claude), PerplexityBot (Perplexity), Google-Extended (Gemini & AI Overviews). robots.txt is the one file all major AI operators read. Blocking these bots, often the default, makes your pages invisible to live AI search. Publish llms.txt too, but expect little: Ahrefs found only 3% of llms.txt files were ever requested across 137,210 domains (June 2026), and Google says it does not use it.
The five AI crawlers to allow, and the honest truth about llms.txt: publish it, but expect nothing. Sources: robots.txt user-agents are the operators' own; llms.txt data from Ahrefs, June 2026. Ahrefs llms.txt study

Where this failed for us#

Play 14 exists because we lived its failure. Before launch, a change to our site’s navigation and footer stripped almost all of the internal links: the homepage ended up linking only to the about page and a memo. The result was that our entire content surface (dozens of pages) became an orphan island, connected to nothing the crawler could follow from the homepage. Search indexed the homepage and almost none of the rest, and we did not catch it until our weekly analytics review showed the whole surface uncrawled. The fix was a footer with real links to every section, which is now the crawl path into everything. The lesson: a page that nothing links to is a page that does not exist, to a crawler or a model. Check your internal links before you write another word of content.

In the wild: the docs that became a channel#

The clearest example of on-page AEO working is Vercel. Its documentation is the open, well-maintained, answer-ready reference for the exact tasks its buyers ask AI about, and ChatGPT referrals grew from under 1% of Vercel’s new signups to about 10% in about seven months. The mechanism is everything in this chapter at once: the docs are crawlable, they answer the question directly, and they are unambiguously Vercel’s. The full teardown is in the teardowns chapter. The lesson for your own pages: treat your docs and money pages as the answer to your category’s questions, not as brochures.

Build your plan · This chapter's artifact

Your ranked on-page fix list

Fill this in as you run the plays. It becomes the on-page column of your 90-day plan, and it saves locally as you type.

See your plan so far →

Everything you type saves in this browser and assembles into one document on the Your plan page, where you can copy or download it. Nothing is sent anywhere. A duplicatable Notion and Google Sheets version ships with the companion pack.

Where this goes next#

Your on-page fix list is one column of your 90-day plan. On its own it makes your pages readable and citable; paired with the content plays it makes them worth citing, and paired with the off-page plays it makes sure something links to them in the first place. Grade any page with our free answer readiness scorer and generate clean markup with the schema generator; when you want it done at scale, Prefer optimizes your existing pages and creates new answer-ready content. Do the crawler check today, rewrite one page answer-first, and you have started. Everything you fill in above lands on your plan page.

People also ask

Frequently asked questions.

Updated 18 July 2026

Should I allow or block AI crawlers like GPTBot?

If you want to be cited in AI answers, allow them. robots.txt is the one file every major AI operator honors, and blocking the AI bots (which some CMSs and CDNs do by default) makes your pages invisible to live AI search no matter how good they are. Allow GPTBot and OAI-SearchBot (ChatGPT), ClaudeBot (Claude), PerplexityBot (Perplexity), and Google-Extended (Gemini and AI Overviews). Blocking them only makes sense if you have a specific reason to keep your content out of AI answers entirely.

Is on-page or off-page AEO more important?

Off-page wins more citations, but on-page comes first. In our study of 1,237 citations, third-party sources (listicles, reviews, communities) accounted for far more citations than vendors' own sites, so off-page is the bigger lever over time. But on-page is the foundation: if a model cannot crawl your page, cannot find the answer on it, or cannot tell who you are, no amount of off-page work converts a visit into a citation of you. Do the on-page fixes once, do them well, then spend most of your ongoing effort off-page.

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.