Chapter 05

On-Page AEO: Make Your Pages Answer-Ready

Six on-page fixes that make AI engines read, understand, and cite your pages: answer-first blocks, question headings, schema, entity clarity, crawlers.

By the end of this chapter you will have done

Your ranked on-page fix list: the specific pages to rewrite answer-first, the schema to add, the entity to clean up, the links to fix, and the crawler check to run first. Plays 10 to 15.

TL;DR

On-page AEO is the cheapest and fastest lever, though not the biggest: off-page won more citations in our study. Six fixes make your pages answer-ready: lead with a self-contained answer, write headings as questions, add the schema AI reads, become one clean entity, fix your internal links, and let the AI crawlers in. Do it once, do it well, then spend most ongoing effort off-page.

  • On-page AEO is the cheapest and fastest lever, but not the biggest: off-page sources won more citations in our study, so do on-page once and well, then spend most ongoing effort off-page.
  • Lead every money page with a self-contained answer, because models lift claims that stand alone without the surrounding page for context.
  • robots.txt is the one file every major AI crawler honors; a default block makes even a perfect page invisible to live search.
  • On-page work reaches grounded (live-search) answers, not memory answers, which only change as your off-page reputation does.

On-page AEO is the cheapest lever in this playbook, and the one most people over-invest in. It is worth doing, and doing well, because a page a model cannot read or understand cannot be cited. But it is not the biggest lever: in our study of 1,237 citations, third-party sources won far more citations than vendors’ own sites. So the honest instruction for this chapter is: do these six fixes once, do them properly, then move most of your ongoing effort to the off-page chapter.

There is a second limit worth naming up front. On-page work reaches grounded answers, the ones where the engine actually reads live pages. It does very little for memory answers, which change only as your off-page reputation does (chapter 2 covers that split). On-page is how you make sure that when a model does read your page, it can find the answer, understand it, and know who you are.

The one move that matters most: lead with the answer#

If you do only one thing in this chapter, do this. Models lift claims that stand on their own. A page that buries its answer under three paragraphs of preamble gives the model nothing clean to quote; a page that opens with a self-contained answer hands it a citation.

Two versions of the same page answering 'What is expense management software?'. On the left, answer buried: the page opens with generic filler and the real answer sits in paragraph four, wrapped in preamble, which is hard for a model to lift and often skipped. On the right, answer first: the page opens with a bold self-contained sentence defining expense management software, which is easy for an engine to cite. The answer-first version is marked self-contained and easy to cite.
The same page, two ways. The version that opens with a self-contained answer is the one a model can lift cleanly. Illustrative example.

The test is simple: copy your opening block, paste it somewhere with no other context, and read it. If it still makes sense and names its own subject (not “it” or “this tool”), it is liftable. If it only makes sense after the paragraph above it, rewrite it.

The six fixes#

Here are the six on-page plays. Numbered 10 to 15, but start with the crawler check in Play 15: it takes ten minutes and there is no point fixing a page the engines are blocked from reading.

A page mockup annotated with the six on-page fixes. At the top, an answer-first block (Play 10). Below it, a heading written as a question (Play 11). A structured-data badge showing FAQPage, Organization, and HowTo schema (Play 12). A row of descriptive internal links to pricing, a competitor comparison, and a guide (Play 14). An entity line stating the brand name, its one-sentence definition, and sameAs links to profiles (Play 13). A 'crawlers allowed' marker in the browser chrome (Play 15). Each fix is listed on the right with its play number.
The six on-page fixes, mapped onto one page. Each is a play below; together they make a page a model can read, parse, and attribute.
Play 10
AnyStart here

Rewrite your top pages answer-first

Put a self-contained answer at the top of every money page, so a model can lift it without reading the rest.

Why it works

Peer-reviewed GEO research found that making content more quotable and well-structured raised how often generative engines cited it, with fluency and quotation optimizations each lifting visibility by roughly a quarter (up to 28%). Self-contained answers are the practical version of that. (Aggarwal et al., KDD 2024, checked 2026-07-14)

Steps
  1. List your top 5 money pages (the ones tied to your money prompts).
  2. For each, write a 2 to 3 sentence answer to the page's core question and put it at the very top, before any preamble.
  3. Write it to stand alone: name the subject explicitly, no 'it' or 'this tool' that refers to something above.
  4. Bold the single most citable sentence so it survives extraction on its own.
  5. Keep the depth below the answer; the answer block earns the citation, the detail earns the reader.
Tools Your CMS. Free
Effort About 30 minutes per page (estimate, not measured)
Time to impact Grounded engines can pick it up after the next crawl (weeks, estimate)

Done when: Your top 5 money pages each open with a self-contained answer block.

Verify it worked: Copy each opening block, paste it with no context, and confirm it still stands alone. Then re-ask that page's money prompt monthly and watch for the lift.

Common failure mode: Opening with brand throat-clearing ('At Acme, we believe...') before the answer. The model quotes the answer, not the preamble, so put the answer first.

Play 11
Any

Write your headings as the questions buyers ask

Match your H2s to the natural-language questions people ask AI, so your sections map directly onto prompts.

Why it works

Buyers ask assistants full questions, and models match content to the question. Our study's winning pages answered those questions directly; question-shaped headings with a direct answer underneath are the on-page shape of that, and the same Q-and-A structure is why Reddit gets cited so heavily. (Prefer AI Citation Study, checked 2026-07-10)

Steps
  1. Pull the real questions from your 10 money prompts and from the engines' 'people also ask' boxes.
  2. Rewrite your money-page H2s as those questions, in the words buyers actually use.
  3. Answer each question in the first sentence under its heading, self-contained.
  4. Add an FAQ block of at least 3 verbatim buyer questions with direct answers, and mark it up with FAQPage schema (Play 12).
Tools Your money prompts and the engines' related-questions boxes. Free
Effort About 1 hour per page (estimate)
Time to impact Weeks, after re-crawl (estimate)

Done when: Your money pages use question-shaped H2s and carry a 3+ question FAQ block.

Verify it worked: Ask one of your H2 questions in a grounded engine and see whether your section is the kind of thing it lifts.

Common failure mode: Clever headings that match no query ('The Acme difference'). A model searching a buyer's question never matches marketing-speak.

Play 12
Any

Add the schema types AI reads

Mark up your pages so an engine can tell which block is a definition, which is a step-by-step, and which is a question and answer.

Why it works

Schema makes your content unambiguous to parse, so the right block gets lifted cleanly. It is clarity, not a guaranteed ranking: its value is that the model does not have to guess what each part of your page is.

Steps
  1. Add Organization schema sitewide, with sameAs links to your off-page profiles (ties to Play 13).
  2. Add FAQPage to your question-and-answer blocks, HowTo to your step lists, and DefinedTerm to your definitions.
  3. Add BreadcrumbList to every page so the model understands your site structure.
  4. Keep every field consistent with the visible text; never mark up something that is not on the page.
  5. Validate before you ship.
Tools Your CMS or a schema generator; validator.schema.org and Google's Rich Results Test
Effort A few hours sitewide (estimate)
Time to impact Immediate once re-crawled (estimate)

Done when: The five schema types validate with no errors on your money pages.

Verify it worked: Run each page through the Rich Results Test and validator.schema.org; both should pass clean.

Common failure mode: Schema that contradicts the visible page. Models and search engines both discount mismatched markup, and it can cost you trust.

Play 13
Any

Make your brand one clean entity

Make it unambiguous who you are, so engines resolve you correctly, especially in the memory answers on-page fixes otherwise cannot reach.

Why it works

ChatGPT answered 29 of 47 questions from memory in our study, and a model can only describe you from memory if it has a stable, consistent picture of who you are. One name and one definition, repeated everywhere and linked to your off-page records, is how that picture forms. (Prefer AI Citation Study, checked 2026-07-10)

Before you start: Ideally your entity records from Play 19, so the sameAs links have somewhere to point.

Steps
  1. Pick one canonical brand name and use it identically everywhere (no stray variations).
  2. Write one one-sentence definition of what you are and what category you are in.
  3. Put that definition on your homepage and about page, in the same words.
  4. Wire Organization schema sameAs to your Wikidata, Crunchbase, and LinkedIn records (Play 19).
  5. Keep the name and definition frozen; consistency over time is what builds the entity.
Tools Your CMS and your Play 19 profiles. Free
Effort About 2 hours (estimate)
Time to impact Slow: memory-side recognition builds over months (estimate)

Done when: Your name and definition are identical on-page and across your off-page profiles, with sameAs wired.

Verify it worked: Ask an engine 'What is [your brand]?' monthly and check it describes the real company, in terms close to your definition.

Common failure mode: A different tagline and description on every page. If the model sees five versions of what you are, it forms none of them.

Play 14
Any

Fix your internal links

Make every money page reachable and connected, so crawlers can find it and models can relate it to the rest of your site.

Why it works

A page nothing links to often does not get crawled at all, so it cannot be cited. We learned this the hard way (see below): a single navigation change once orphaned our entire content surface, and search crawled almost none of it.

Steps
  1. List your money pages and check each is within about two clicks of the homepage.
  2. Add descriptive contextual links between related pages (the anchor text should say what the target is, never 'click here').
  3. Link every page to at least one conversion page, so the reader always has a next step.
  4. Run an internal-link crawl and fix every orphaned money page.
Tools Your CMS; a crawler like Screaming Frog's free tier
Effort About half a day (estimate)
Time to impact Weeks, as the newly linked pages get crawled (estimate)

Done when: No money page is orphaned, and each sits within about two clicks of the homepage.

Verify it worked: Crawl the site and confirm zero orphaned money pages.

Common failure mode: Publishing pages nothing links to. Orphans are invisible to crawlers, and a page that is never crawled is never cited.

Play 15
AnyDo this first

Let the AI crawlers in

Make sure the engines are actually allowed to read your pages. This is the ten-minute check to run before anything else on this list.

Why it works

robots.txt is the one file all the major AI operators honor, and some content platforms and CDNs block the AI crawlers by default. A blocked bot cannot read a page, so every other on-page fix is wasted until the crawler is allowed in.

Steps
  1. Open yoursite.com/robots.txt and read it.
  2. Confirm none of these are disallowed: GPTBot, OAI-SearchBot (ChatGPT), ClaudeBot (Claude), PerplexityBot (Perplexity), Google-Extended (Gemini and AI Overviews).
  3. If you are unsure, add an explicit Allow for each of those user-agents.
  4. Publish an llms.txt too, as cheap hygiene, but expect little from it (see the honest note below).
  5. Re-check robots.txt after any platform migration or CDN change; those often reset it.
Tools Your robots.txt. Free
Effort About 15 minutes (estimate)
Time to impact Immediate: it unblocks everything else (estimate)

Done when: The five AI crawlers are allowed in robots.txt.

Verify it worked: Fetch yoursite.com/robots.txt and confirm there is no Disallow covering the AI user-agents.

Common failure mode: Leaving a blanket Disallow or a default CMS block in place. It is the most common way a good page stays invisible, and the easiest to miss.

A robots.txt allow-list showing five AI crawlers each permitted: GPTBot (ChatGPT training), OAI-SearchBot (ChatGPT search), ClaudeBot (Claude), PerplexityBot (Perplexity), and Google-Extended (Gemini and AI Overviews), each with an Allow check. A note explains robots.txt is the one file all major AI operators read, and blocking these bots (often the default) makes pages invisible to live search. A second note says to publish llms.txt too but expect little: Ahrefs found only 3% of llms.txt files ever get requested across 137,210 domains in June 2026, and Google says it does not use it.
The five AI crawlers to allow, and the honest truth about llms.txt: publish it, but expect nothing. Sources: robots.txt user-agents are the operators' own; llms.txt data from Ahrefs, June 2026. Ahrefs llms.txt study

Where this failed for us#

Play 14 exists because we lived its failure. Before launch, a change to our site’s navigation and footer stripped almost all of the internal links: the homepage ended up linking only to the about page and a memo. The result was that our entire content surface (dozens of pages) became an orphan island, connected to nothing the crawler could follow from the homepage. Search indexed the homepage and almost none of the rest, and we did not catch it until our weekly analytics review showed the whole surface uncrawled. The fix was a footer with real links to every section, which is now the crawl path into everything. The lesson: a page that nothing links to is a page that does not exist, to a crawler or a model. Check your internal links before you write another word of content.

In the wild: the docs that became a channel#

The clearest example of on-page AEO working is Vercel. Its documentation is the open, well-maintained, answer-ready reference for the exact tasks its buyers ask AI about, and ChatGPT referrals grew from under 1% of Vercel’s new signups to about 10% in about seven months. The mechanism is everything in this chapter at once: the docs are crawlable, they answer the question directly, and they are unambiguously Vercel’s. The full teardown is in the teardowns chapter. The lesson for your own pages: treat your docs and money pages as the answer to your category’s questions, not as brochures.

Build your plan · This chapter's artifact

Your ranked on-page fix list

Fill this in as you run the plays. It becomes the on-page column of your 90-day plan, and it saves locally as you type.

See your plan so far →

Everything you type saves in this browser and assembles into one document on the Your plan page, where you can copy or download it. Nothing is sent anywhere. A duplicatable Notion and Google Sheets version ships with the companion pack.

Where this goes next#

Your on-page fix list is one column of your 90-day plan. On its own it makes your pages readable and citable; paired with the content plays it makes them worth citing, and paired with the off-page plays it makes sure something links to them in the first place. Do the crawler check today, rewrite one page answer-first, and you have started. Everything you fill in above lands on your plan page.

People also ask

Frequently asked questions

How do I optimize my website for AI search?

Make each page easy for a model to read, understand, and lift. In practice that is six fixes: confirm the AI crawlers are allowed in robots.txt, lead every money page with a self-contained answer, write your headings as the questions buyers ask, add the schema types AI reads (Organization, FAQPage, HowTo, DefinedTerm, BreadcrumbList), present your brand as one clean entity, and fix your internal links so no money page is orphaned. On-page work is the cheapest lever, but it mostly influences live-search answers, so pair it with the off-page work that earns citations.

Does schema markup help with AI search?

It helps by making your content unambiguous to parse, not by guaranteeing a citation. Schema tells an engine which block is a definition, which is a step-by-step, and which is a question and answer, so the right piece can be lifted cleanly. The types that matter most are Organization (with sameAs), FAQPage, HowTo, DefinedTerm, and BreadcrumbList. One rule: schema must match the visible text. Marking up content that is not on the page is discounted by both search engines and models, and can hurt trust.

Should I allow or block AI crawlers like GPTBot?

If you want to be cited in AI answers, allow them. robots.txt is the one file every major AI operator honors, and blocking the AI bots (which some CMSs and CDNs do by default) makes your pages invisible to live AI search no matter how good they are. Allow GPTBot and OAI-SearchBot (ChatGPT), ClaudeBot (Claude), PerplexityBot (Perplexity), and Google-Extended (Gemini and AI Overviews). Blocking them only makes sense if you have a specific reason to keep your content out of AI answers entirely.

Is on-page or off-page AEO more important?

Off-page wins more citations, but on-page comes first. In our study of 1,237 citations, third-party sources (listicles, reviews, communities) accounted for far more citations than vendors' own sites, so off-page is the bigger lever over time. But on-page is the foundation: if a model cannot crawl your page, cannot find the answer on it, or cannot tell who you are, no amount of off-page work converts a visit into a citation of you. Do the on-page fixes once, do them well, then spend most of your ongoing effort off-page.

Go deeper

See how AI search cites you, free.

Run your AI visibility audit and see exactly where you stand across every model.

Run your free audit
Diagnostic in 10 mins · No card, no commitment