seo-geo

Optimize content for AI Overviews (formerly SGE), ChatGPT web search, Perplexity, and other AI-powered search experiences. Generative Engine Optimization (GEO) analysis including brand mention signals, AI crawler accessibility, llms.txt compliance, passage-level citability scoring, and platform-spec

By agricidaniel · 5,531 installs

npx skills add agricidaniel/claude-seo --skill seo-geo

Source repository · Upstream listing

AI Search / GEO Optimization (May 2026) Primary Source: Google's AI Optimization Guide Google's official position, published under Search Central docs: "Optimizing for generative AI search is still SEO from Google's perspective. AEO and GEO are rebranded labels for the same work." Read references/google ai optimization guide.md for the full synthesis, myth busting list ( llms.txt , chunking, AI rephrasing, mention farming, all rejected by Google as ineffective), and the Who/How/Why test for content quality. Audits should frame GEO findings as SEO fundamentals applied to AI search surfaces , not as a separate optimization discipline. When community recommendations contradict Google's primary source, defer to Google and note the contradiction in the report. Key Statistics Metric Value Source AI Overviews reach 2.5 billion+ monthly active users, reported from Google I/O 2026 keynote coverage; not confirmed on a Google owned source; 200+ countries Third party I/O reporting AI Overviews query coverage ~50% of queries (third party measurement; varies by country) Industry data AI Mode monthly users 1B+, reported from Google I/O 2026 keynote coverage; not confirmed on a Google owned source Third party I/O reporting AI Mode model custom version of Gemini 2.5 Google AI referred sessions growth 527% (Jan May 2025) SparkToro ChatGPT weekly active users 900 million OpenAI Perplexity monthly queries 500+ million Perplexity Critical Insight: Brand Mentions Backlinks Brand mentions correlate 3x more strongly with AI visibility than backlinks. (Ahrefs December 2025 study of 75,000 brands) Signal Correlation with AI Citations YouTube mentions ~0.737 (strongest) Reddit mentions High Wikipedia presence High LinkedIn presence Moderate Domain Rating (backlinks) ~0.266 (weak) Only 11% of domains are cited by both ChatGPT and Google AI Overviews for the same query, so platform specific optimization is essential. GEO Analysis Criteria (Updated) 1. Citability Score (25%) Optimal passage length: 134 167 words for AI citation. And ~44% of AI citations come from the first 30% of a page (SE Ranking study), front load your most citable, self contained answer rather than burying it below the fold. Strong signals: Clear, quotable sentences with specific facts/statistics Self contained answer blocks (can be extracted without context) Direct answer in first 40 60 words of section Claims attributed with specific sources Definitions following "X is..." or "X refers to..." patterns Unique data points not found elsewhere Weak signals: Vague, general statements Opinion without evidence Buried conclusions No specific data points 2. Structural Readability (20%) 92% of AI Overview citations come from top 10 ranking pages , but 47% come from pages ranking below position 5, demonstrating different selection logic. Strong signals: Clean H1 H2 H3 heading hierarchy Question based headings (matches query patterns) Short paragraphs (2 4 sentences) Tables for comparative data Ordered/unordered lists for step by step or multi item content FAQ sections with clear Q&A format Weak signals: Wall of text with no structure Inconsistent heading hierarchy No lists or tables Information buried in paragraphs 3. Multi Modal Content (15%) Content with multi modal elements sees 156% higher selection rates . Check for: Text + relevant images Video content (embedded or linked) Infographics and charts Interactive elements (calculators, tools) Structured data supporting media 4. Authority & Brand Signals (20%) Strong signals: Author byline with credentials Publication date and last updated date Recency , content under 3 months old is ~3x more likely to be cited in AI answers; pages left stale 6+ months lose citation eligibility (SE Ranking, 1.3M citation study). A scheduled refresh program is one of the highest leverage GEO plays. Citations to primary sources (studies, official docs, data) Organization credentials and affiliations Expert quotes with attribution Entity presence in Wikipedia, Wikidata Mentions on Reddit, YouTube, LinkedIn Weak signals: Anonymous authorship No dates No sources cited No brand presence across platforms 5. Technical Accessibility (20%) AI crawlers do NOT execute JavaScript. Server side rendering is critical. Check for: Server side rendering (SSR) vs client only content AI crawler access in robots.txt llms.txt file presence and configuration RSL 1.0 licensing terms AI Crawler Detection Check robots.txt for these AI crawlers: Crawler Owner Purpose Obeys robots.txt? GPTBot OpenAI Model training only (NOT ChatGPT Search) yes OAI SearchBot OpenAI ChatGPT Search citability (the crawler that decides it) yes ChatGPT User OpenAI ChatGPT browsing (user triggered) no (user triggered) ClaudeBot Anthropic Model training only (NOT Claude's search features) yes Claude SearchBot Anthropic Claude/Claude.ai search result citability (the crawler that decides it) yes Claude User Anthropic Claude browsing on a user's behalf (user triggered) no (user triggered) PerplexityBot Perplexity Perplexity AI search yes CCBot Common Crawl Training data (often blocked) yes Bytespider ByteDance TikTok/Douyin AI yes cohere ai Cohere Cohere models yes Google Extended Google Gemini/Vertex training & grounding only (NOT Google Search) yes Google CloudVertexBot Google Site owner requested Vertex AI Agent crawls yes Google Agent Google Agentic browsing (Project Mariner), acts for a user no (user triggered) Google NotebookLM Google Fetches individual user added source URLs no (user triggered) Google Messages Google User triggered fetch no (user triggered) Applebot Extended Apple Apple Intelligence / generative AI training data opt out only (NOT Siri, Spotlight, or Safari search; does not itself crawl, it labels content already fetched by Applebot) yes Sources: [OpenAI crawlers](https://platform.openai.com/docs/bots), [Google crawlers overview](https://developers.google.com/search/docs/crawling indexing/overview google crawlers), [Anthropic crawler support article](https://support.anthropic.com/en/articles/8896518 does anthropic crawl data from the web and how can site owners block the crawler), [Apple Applebot Extended support article](https://support.apple.com/en us/119829). Anthropic's current crawler support article documents only ClaudeBot, Claude User, and Claude SearchBot; it does not list anthropic ai , so the previously unverified anthropic ai row has been removed rather than kept as a guess. Recommendation: Allow OAI SearchBot, Claude SearchBot, and PerplexityBot for AI search visibility. GPTBot, ClaudeBot, CCBot, and Applebot Extended are training only signals allow or block them on licensing preference, not on search visibility grounds. Check the right bot for the claim you are making Two pairs are routinely conflated. Each claim below may only be supported by its own bot's robots.txt status check them separately and report them separately. Claim you want to make Bot to check Bot that does NOT support this claim "Content is citable in ChatGPT Search" OAI SearchBot GPTBot "Content is available for OpenAI model training" GPTBot OAI SearchBot "Content can be used for Gemini/Vertex training & grounding" Google Extended Googlebot "Content is eligible for Google Search / AI Overviews" Googlebot Google Extended "Content is citable in Claude's search features" Claude SearchBot ClaudeBot "Content is available for Anthropic model training" ClaudeBot Claude SearchBot "Content can be used for Apple Intelligence training" Applebot Extended Applebot "Content is discoverable via Siri, Spotlight, or Safari search" Applebot Applebot Extended Google Extended governs Gemini and Vertex AI training and grounding use only. It does not affect inclusion in ordinary Google Search, or in AI Overviews and AI Mode, both of which are served from the Googlebot index. Never score Google Extended as a "Google Search readiness" signal, and never cite a blocked Google Extended as evidence that a site is missing from Google Search. OAI SearchBot is the crawler that determines ChatGPT Search citability. GPTBot is OpenAI's separate training crawler. Checking GPTBot access tells you nothing about whether ChatGPT Search can cite the page. A site that blocks GPTBot and allows OAI SearchBot is fully citable in ChatGPT Search. Claude SearchBot is the crawler that determines citability in Claude's own search features. ClaudeBot is Anthropic's separate training crawler (per Anthropic's crawler support article). Checking ClaudeBot access tells you nothing about Claude search citability, and vice versa; report each separately. Applebot Extended is a training data opt out signal, not a crawler that fetches pages itself. Per Apple's support article, disallowing Applebot Extended opts a site out of Apple Intelligence / generative model training use, but the page remains discoverable through Siri, Spotlight, and Safari as long as Applebot itself is allowed. Never cite a blocked Applebot Extended as evidence a site is missing from Apple's search surfaces. Do not use these names interchangeably in report prose. When reporting crawler access, name the specific user agent that was checked and the specific capability it governs. User triggered fetchers ignore robots.txt by design (Google Agent, Google NotebookLM, Google Messages, ChatGPT User). robots.txt cannot block them, use server side access controls. Google's canonical crawling/robots reference moved to developers.google.com/crawling (migrated 2025 11 20); IP range files now live at /crawling/ipranges/ and googlebot.json was renamed common crawlers.json . Emerging: Web Bot Auth (RFC 9421) lets bots authenticate via a Signature Agent header + key directory (used by Google Agent); reverse DNS verification remains the fallback. llms.txt Standard Read references/llmstxt evidence.md for the primary source evidence (Mueller, Illyes, SE Ranking 300k domain study, OtterlyAI server log audit) on why /llms.txt is not currently a citation lever for major AI search systems. claude seo reports presence but assigns no citation ranking weight. Google now states this explicitly. Google's AI optimization guide, introduced 2026 05 15 and clarified 2026 06 15, says llms.txt and other AI text files are not needed for Google Search and do not help or hurt visibility or rankings. They may still serve non Google systems. Never recommend llms.txt as a Google ranking or citation lever. Source: developers.google.com/search/docs/fundamentals/ai optimization guide The emerging llms.txt standard provides AI crawlers with structured content guidance. Location: /llms.txt (root of domain) Format: Check for: Presence of /llms.txt Structured content guidance Key page highlights Contact/authority information RSL 1.0 (Really Simple Licensing) New standard (December 2025) for machine readable AI licensing terms. Backed by: Reddit, Yahoo, Medium, Quora, Cloudflare, Akamai, Creative Commons Check for: RSL implementation and appropriate licensing terms. Platform Specific Optimization Platform Key Citation Sources Optimization Focus Google AI Overviews Strongly ranking correlated, cites pages that already rank well Traditional SEO + passage optimization Google AI Mode (custom version of Gemini 2.5)