seo-geo
Optimize content for AI Overviews (formerly SGE), ChatGPT web search, Perplexity, and other AI-powered search experiences. Generative Engine Optimization (GEO) analysis including brand mention signals, AI crawler accessibility, llms.txt compliance, passage-level citability scoring, and platform-spec
By agricidaniel · 5,531 installs
npx skills add agricidaniel/claude-seo --skill seo-geo
Source repository · Upstream listing
AI Search / GEO Optimization (May 2026)
Primary Source: Google's AI Optimization Guide
Google's official position, published under Search Central docs:
"Optimizing for generative AI search is still SEO from Google's
perspective. AEO and GEO are rebranded labels for the same work."
Read references/google ai optimization guide.md for the full synthesis,
myth busting list ( llms.txt , chunking, AI rephrasing, mention farming,
all rejected by Google as ineffective), and the Who/How/Why test for
content quality.
Audits should frame GEO findings as SEO fundamentals applied to AI search
surfaces , not as a separate optimization discipline. When community
recommendations contradict Google's primary source, defer to Google and note
the contradiction in the report.
Key Statistics
Metric Value Source
AI Overviews reach 2.5 billion+ monthly active users, reported from Google I/O 2026 keynote coverage; not confirmed on a Google owned source; 200+ countries Third party I/O reporting
AI Overviews query coverage ~50% of queries (third party measurement; varies by country) Industry data
AI Mode monthly users 1B+, reported from Google I/O 2026 keynote coverage; not confirmed on a Google owned source Third party I/O reporting
AI Mode model custom version of Gemini 2.5 Google
AI referred sessions growth 527% (Jan May 2025) SparkToro
ChatGPT weekly active users 900 million OpenAI
Perplexity monthly queries 500+ million Perplexity
Critical Insight: Brand Mentions Backlinks
Brand mentions correlate 3x more strongly with AI visibility than backlinks.
(Ahrefs December 2025 study of 75,000 brands)
Signal Correlation with AI Citations
YouTube mentions ~0.737 (strongest)
Reddit mentions High
Wikipedia presence High
LinkedIn presence Moderate
Domain Rating (backlinks) ~0.266 (weak)
Only 11% of domains are cited by both ChatGPT and Google AI Overviews for the same query, so platform specific optimization is essential.
GEO Analysis Criteria (Updated)
1. Citability Score (25%)
Optimal passage length: 134 167 words for AI citation. And ~44% of AI
citations come from the first 30% of a page (SE Ranking study), front load
your most citable, self contained answer rather than burying it below the fold.
Strong signals:
Clear, quotable sentences with specific facts/statistics
Self contained answer blocks (can be extracted without context)
Direct answer in first 40 60 words of section
Claims attributed with specific sources
Definitions following "X is..." or "X refers to..." patterns
Unique data points not found elsewhere
Weak signals:
Vague, general statements
Opinion without evidence
Buried conclusions
No specific data points
2. Structural Readability (20%)
92% of AI Overview citations come from top 10 ranking pages , but 47% come from pages ranking below position 5, demonstrating different selection logic.
Strong signals:
Clean H1 H2 H3 heading hierarchy
Question based headings (matches query patterns)
Short paragraphs (2 4 sentences)
Tables for comparative data
Ordered/unordered lists for step by step or multi item content
FAQ sections with clear Q&A format
Weak signals:
Wall of text with no structure
Inconsistent heading hierarchy
No lists or tables
Information buried in paragraphs
3. Multi Modal Content (15%)
Content with multi modal elements sees 156% higher selection rates .
Check for:
Text + relevant images
Video content (embedded or linked)
Infographics and charts
Interactive elements (calculators, tools)
Structured data supporting media
4. Authority & Brand Signals (20%)
Strong signals:
Author byline with credentials
Publication date and last updated date
Recency , content under 3 months old is ~3x more likely to be cited in AI answers; pages left stale 6+ months lose citation eligibility (SE Ranking, 1.3M citation study). A scheduled refresh program is one of the highest leverage GEO plays.
Citations to primary sources (studies, official docs, data)
Organization credentials and affiliations
Expert quotes with attribution
Entity presence in Wikipedia, Wikidata
Mentions on Reddit, YouTube, LinkedIn
Weak signals:
Anonymous authorship
No dates
No sources cited
No brand presence across platforms
5. Technical Accessibility (20%)
AI crawlers do NOT execute JavaScript. Server side rendering is critical.
Check for:
Server side rendering (SSR) vs client only content
AI crawler access in robots.txt
llms.txt file presence and configuration
RSL 1.0 licensing terms
AI Crawler Detection
Check robots.txt for these AI crawlers:
Crawler Owner Purpose Obeys robots.txt?
GPTBot OpenAI Model training only (NOT ChatGPT Search) yes
OAI SearchBot OpenAI ChatGPT Search citability (the crawler that decides it) yes
ChatGPT User OpenAI ChatGPT browsing (user triggered) no (user triggered)
ClaudeBot Anthropic Model training only (NOT Claude's search features) yes
Claude SearchBot Anthropic Claude/Claude.ai search result citability (the crawler that decides it) yes
Claude User Anthropic Claude browsing on a user's behalf (user triggered) no (user triggered)
PerplexityBot Perplexity Perplexity AI search yes
CCBot Common Crawl Training data (often blocked) yes
Bytespider ByteDance TikTok/Douyin AI yes
cohere ai Cohere Cohere models yes
Google Extended Google Gemini/Vertex training & grounding only (NOT Google Search) yes
Google CloudVertexBot Google Site owner requested Vertex AI Agent crawls yes
Google Agent Google Agentic browsing (Project Mariner), acts for a user no (user triggered)
Google NotebookLM Google Fetches individual user added source URLs no (user triggered)
Google Messages Google User triggered fetch no (user triggered)
Applebot Extended Apple Apple Intelligence / generative AI training data opt out only (NOT Siri, Spotlight, or Safari search; does not itself crawl, it labels content already fetched by Applebot) yes
Sources: [OpenAI crawlers](https://platform.openai.com/docs/bots),
[Google crawlers overview](https://developers.google.com/search/docs/crawling indexing/overview google crawlers),
[Anthropic crawler support article](https://support.anthropic.com/en/articles/8896518 does anthropic crawl data from the web and how can site owners block the crawler),
[Apple Applebot Extended support article](https://support.apple.com/en us/119829).
Anthropic's current crawler support article documents only ClaudeBot, Claude User,
and Claude SearchBot; it does not list anthropic ai , so the previously unverified
anthropic ai row has been removed rather than kept as a guess.
Recommendation: Allow OAI SearchBot, Claude SearchBot, and PerplexityBot for AI
search visibility. GPTBot, ClaudeBot, CCBot, and Applebot Extended are training only
signals allow or block them on licensing preference, not on search visibility
grounds.
Check the right bot for the claim you are making
Two pairs are routinely conflated. Each claim below may only be supported by its own
bot's robots.txt status check them separately and report them separately.
Claim you want to make Bot to check Bot that does NOT support this claim
"Content is citable in ChatGPT Search" OAI SearchBot GPTBot
"Content is available for OpenAI model training" GPTBot OAI SearchBot
"Content can be used for Gemini/Vertex training & grounding" Google Extended Googlebot
"Content is eligible for Google Search / AI Overviews" Googlebot Google Extended
"Content is citable in Claude's search features" Claude SearchBot ClaudeBot
"Content is available for Anthropic model training" ClaudeBot Claude SearchBot
"Content can be used for Apple Intelligence training" Applebot Extended Applebot
"Content is discoverable via Siri, Spotlight, or Safari search" Applebot Applebot Extended
Google Extended governs Gemini and Vertex AI training and grounding use only.
It does not affect inclusion in ordinary Google Search, or in AI Overviews and AI
Mode, both of which are served from the Googlebot index. Never score
Google Extended as a "Google Search readiness" signal, and never cite a blocked
Google Extended as evidence that a site is missing from Google Search.
OAI SearchBot is the crawler that determines ChatGPT Search citability.
GPTBot is OpenAI's separate training crawler. Checking GPTBot access tells
you nothing about whether ChatGPT Search can cite the page. A site that blocks
GPTBot and allows OAI SearchBot is fully citable in ChatGPT Search.
Claude SearchBot is the crawler that determines citability in Claude's own
search features. ClaudeBot is Anthropic's separate training crawler (per
Anthropic's crawler support article). Checking ClaudeBot access tells you
nothing about Claude search citability, and vice versa; report each separately.
Applebot Extended is a training data opt out signal, not a crawler that
fetches pages itself. Per Apple's support article, disallowing
Applebot Extended opts a site out of Apple Intelligence / generative model
training use, but the page remains discoverable through Siri, Spotlight, and
Safari as long as Applebot itself is allowed. Never cite a blocked
Applebot Extended as evidence a site is missing from Apple's search surfaces.
Do not use these names interchangeably in report prose. When reporting crawler access,
name the specific user agent that was checked and the specific capability it governs.
User triggered fetchers ignore robots.txt by design (Google Agent, Google NotebookLM, Google Messages, ChatGPT User). robots.txt cannot block them, use server side access controls. Google's canonical crawling/robots reference moved to developers.google.com/crawling (migrated 2025 11 20); IP range files now live at /crawling/ipranges/ and googlebot.json was renamed common crawlers.json . Emerging: Web Bot Auth (RFC 9421) lets bots authenticate via a Signature Agent header + key directory (used by Google Agent); reverse DNS verification remains the fallback.
llms.txt Standard
Read references/llmstxt evidence.md for the primary source evidence (Mueller, Illyes, SE Ranking 300k domain study, OtterlyAI server log audit) on why /llms.txt is not currently a citation lever for major AI search systems. claude seo reports presence but assigns no citation ranking weight.
Google now states this explicitly. Google's AI optimization guide, introduced
2026 05 15 and clarified 2026 06 15, says llms.txt and other AI text files are
not needed for Google Search and do not help or hurt visibility or rankings.
They may still serve non Google systems. Never recommend llms.txt as a Google
ranking or citation lever. Source:
developers.google.com/search/docs/fundamentals/ai optimization guide
The emerging llms.txt standard provides AI crawlers with structured content guidance.
Location: /llms.txt (root of domain)
Format:
Check for:
Presence of /llms.txt
Structured content guidance
Key page highlights
Contact/authority information
RSL 1.0 (Really Simple Licensing)
New standard (December 2025) for machine readable AI licensing terms.
Backed by: Reddit, Yahoo, Medium, Quora, Cloudflare, Akamai, Creative Commons
Check for: RSL implementation and appropriate licensing terms.
Platform Specific Optimization
Platform Key Citation Sources Optimization Focus
Google AI Overviews Strongly ranking correlated, cites pages that already rank well Traditional SEO + passage optimization
Google AI Mode (custom version of Gemini 2.5)