Here is the checklist: 25 points across five layers, technical, on-page, entity, off-site, and measurement, ordered so each layer makes the next one work. Run the technical items first, because AI crawlers that cannot read your site make every other item pointless.
Treat it as an audit, not a syllabus. Most brands pass half the items already; the score matters less than which layer holds your failures, since clustered failures point to one root cause.
Technical: can AI read you at all
GPTBot, OAI-SearchBot, ClaudeBot, and PerplexityBot fetch raw HTML and do not execute JavaScript, which drives the first five items.
- Key pages are server-rendered or static, because client-side content is invisible to AI crawlers
- robots.txt permits the AI bots you want, since accidental blocks from old configs are common
- Content appears with JavaScript disabled: the fifteen-second curl test on your top five pages
- Prices, specs, and key facts live in HTML text, not images, PDFs, or script-injected widgets
- Server logs show AI bot visits, confirming the crawlers are actually fetching your pages
On-page: is anything worth citing
AI engines cite the clearest extractable answer, so structure your pages so the answer is impossible to miss. This is the layer the covered question "how to structure content so AI cites it" lives in.
- Every key page answers its main question in the first two sentences, no wind-up
- H2s are phrased as the questions buyers actually ask, one question per section
- Each section is self-contained and readable out of context, since retrieval pulls passages, not pages
- Facts outnumber adjectives: numbers, names, and dates survive summarization, superlatives do not
- Comparison content exists for your vs, alternatives, and best-of prompts, with the honest competitor treatment covered in how to get AI to recommend your brand over competitors
- Dates are visible and honest, because retrieval-based engines favor demonstrably fresh sources
Entity: does AI know what you are
Models recommend entities they can resolve. Inconsistency across sources is the quiet killer here.
- One canonical one-line description used verbatim on your site, LinkedIn, Crunchbase, and directories
- Organization schema on the homepage with sameAs links to major profiles
- Product or Service schema on offer pages, mirroring visible facts including price
- No stale identities: old names, dead products, and pre-pivot descriptions cleaned from ranking pages
- A category label you can own, specific enough that AI can slot you confidently into shortlists
Off-site: does anyone corroborate you
AI weights independent evidence over self-description, so mention frequency is largely earned elsewhere.
- Profiles claimed and complete on the review platforms cited in your category
- Steady flow of genuine reviews, since volume and recency both feed AI's confidence
- Placement on the best-of listicles AI engines repeatedly cite for your buyer prompts
- Presence in third-party comparisons, not only your own vs pages
- Authentic community mentions where buyers ask for advice, earned rather than planted
Measurement: do you know if any of it works
No provider publishes prompt logs, so visibility is measured by sampling, and undisciplined sampling measures nothing.
- A baseline scan exists before optimization began, or run a free scan now to create one
- A fixed panel of unbranded buyer prompts, held constant between scans so trends are real
- Coverage across providers, not just ChatGPT, since visibility diverges sharply by engine
- AI referral traffic segmented in analytics: chatgpt.com, perplexity.ai, gemini.google.com
- A re-scan cadence, weekly during active work, monthly at steady state
Score yourself against the checklist in one scan
The free scan covers the measurement layer and the crawler checks for you: GEO score, competitor matrix, per-provider breakdown, and an AI readability check in one pass.
Frequently asked questions
How do I structure content so AI cites it?
Answer the question completely in the first two sentences under a question-shaped heading, keep each section self-contained, use lists and tables for anything enumerable or comparable, and pack in verifiable facts. Serve it as raw HTML, because AI crawlers do not execute JavaScript.
Which checklist layer should I fix first?
Technical, always. If AI crawlers cannot read your pages, on-page, entity, and off-site work cannot pay off. After that, fix whichever layer holds your densest cluster of failures, since clustered gaps usually share one root cause.
How long does the full checklist take to implement?
Technical and entity items are typically days to a few weeks. On-page restructuring runs in content sprints over one to two months, and off-site placement is an ongoing campaign with wins landing over one to two quarters. Measurement takes an afternoon to set up and pays off immediately.
Do I need to complete all 25 points to see results?
No. Retrieval-driven surfaces like Perplexity respond to a readable site plus a few strongly citable pages within weeks. The full checklist is what durable, cross-provider visibility looks like, but early wins come from the technical layer plus answer-first restructuring of your money pages.
Related answers