seo-technical
Technical SEO audit across 9 categories: crawlability, indexability, security, URL structure, mobile, Core Web Vitals, structured data, JavaScript rendering, and IndexNow protocol. Use when user says "technical SEO", "crawl issues", "robots.txt", "Core Web Vitals", "site speed", or "security headers
By agricidaniel · 5,548 installs
npx skills add agricidaniel/claude-seo --skill seo-technical
Source repository · Upstream listing
Technical SEO Audit
Categories
1. Crawlability
robots.txt: exists, valid, not blocking important resources
XML sitemap: run "${CLAUDE PLUGIN ROOT}/scripts/claude seo" run sitemap discovery.py <url json ; require a
valid entry in found , and report stale or unsafe robots.txt declarations
separately from working fallback locations
Noindex tags: intentional vs accidental
Crawl depth: important pages within 3 clicks of homepage
JavaScript rendering: check if critical content requires JS execution
Crawl budget: for large sites ( 10k pages), efficiency matters
Googlebot fetch limits : Googlebot fetches the first 2MB of HTML and first 64MB of a PDF (uncompressed; 15MB is the broader crawler infra default). Long standing, not a 2026 change, but inline base64 images, oversized inline CSS/JS, or bloated nav can push critical content/JSON LD past the cap and out of the index. Keep key content + structured data within the first 2MB.
Crawl rate auto adjusts (backs off on 5xx/slow responses); there is no manual crawl rate control (the legacy Search Console setting was removed Jan 2024). Influence crawling via sitemaps, server responsiveness, and robots controls.
Google's canonical crawling/robots reference moved to developers.google.com/crawling (migrated 2025 11 20); IP range files relocated to /crawling/ipranges/ and googlebot.json was renamed common crawlers.json .
AMP has no separate ranking advantage. Since 2026 07 01, Google Search sends
users directly to publisher hosted AMP URLs, so do not recommend AMP Cache,
AMP Viewer, or signed exchange maintenance. Audit AMP against the same content,
action parity, and quality requirements as other pages.
AI Crawler Management
As of 2025 2026, AI companies actively crawl the web to train models and power AI search. Managing these crawlers via robots.txt is a critical technical SEO consideration.
Known AI crawlers:
Crawler Company robots.txt token Purpose
GPTBot OpenAI GPTBot Model training (NOT ChatGPT Search)
OAI SearchBot OpenAI OAI SearchBot ChatGPT Search citability
ChatGPT User OpenAI ChatGPT User Real time browsing (user triggered)
ClaudeBot Anthropic ClaudeBot Model training (NOT Claude search citability)
Claude SearchBot Anthropic Claude SearchBot Claude search result citability
PerplexityBot Perplexity PerplexityBot Search index + training
Bytespider ByteDance Bytespider Model training
Google Extended Google Google Extended Gemini training (NOT search)
Applebot Extended Apple Applebot Extended Apple Intelligence training opt out (NOT Siri/Spotlight/Safari)
CCBot Common Crawl CCBot Open dataset
Key distinctions:
Blocking Google Extended prevents Gemini training use but does NOT affect Google Search indexing or AI Overviews (those use Googlebot )
Blocking GPTBot prevents OpenAI training but does NOT affect ChatGPT Search
citability, which is governed by OAI SearchBot , nor user triggered browsing
( ChatGPT User ). Check OAI SearchBot for any citability claim; GPTBot
status is evidence about training use only
Blocking ClaudeBot prevents Anthropic model training but does NOT affect
citability in Claude's own search features, which is governed by
Claude SearchBot (per Anthropic's crawler support article). Check
Claude SearchBot for any Claude search citability claim; ClaudeBot status
is evidence about training use only
Blocking Applebot Extended opts out of Apple Intelligence / generative model
training use but does NOT affect discoverability via Siri, Spotlight, or Safari,
which follows Applebot (per Apple's support article); Applebot Extended does
not itself crawl
~3 5% of websites now use AI specific robots.txt rules
Example, selective AI crawler blocking:
Recommendation: Consider your AI visibility strategy before blocking. Being cited by AI systems drives brand awareness and referral traffic. Cross reference the seo geo skill for the full AI crawler/fetcher taxonomy.
User triggered fetchers ignore robots.txt by design. Google now documents Google Agent (Project Mariner, agentic browsing) plus Google NotebookLM and Google Messages as user triggered fetchers that cannot be blocked via robots.txt . Use server side access controls instead. By contrast, Google Extended and Google CloudVertexBot obey robots.txt. Emerging: Web Bot Auth (RFC 9421) lets bots authenticate cryptographically via a Signature Agent header + key directory at agent.bot.goog (used by Google Agent); reverse DNS verification remains the fallback.
2. Indexability
Canonical tags: self referencing, no conflicts with noindex
Duplicate content: near duplicates, parameter URLs, www vs non www
Canonicalization fixes can take time: Google may retain corrected pages in a
duplicate cluster for up to two weeks while re evaluating them. Do not
interpret an unchanged canonical immediately after a fix as proof that the
fix failed.
Thin content: pages below minimum word counts per type
Pagination: rel=next/prev or load more pattern
Hreflang: correct for multi language/multi region sites
Index bloat: unnecessary pages consuming crawl budget
3. Security
HTTPS: enforced, valid SSL certificate, no mixed content
Security headers:
Content Security Policy (CSP)
Strict Transport Security (HSTS)
X Frame Options
X Content Type Options
Referrer Policy
HSTS preload: check preload list inclusion for high security sites
Back button hijacking (spam policy violation, malicious practices): flag pages that defeat the Back button via history.pushState / replaceState (including scripts injected by third party ad/library platforms). Added to Google's spam policies 2026 04 13; enforcement live since 2026 06 15 (manual actions + automated demotions): treat as Critical.
4. URL Structure
Clean URLs: descriptive, hyphenated, no query parameters for content
Hierarchy: logical folder structure reflecting site architecture
Redirects: no chains (max 1 hop), 301 for permanent moves
URL length: flag 100 characters
Trailing slashes: consistent usage
5. Mobile Optimization & Page Experience
Responsive design: viewport meta tag, responsive CSS
Touch targets: minimum 48x48px with 8px spacing
Font size: minimum 16px base
No horizontal scroll
Mobile first indexing: Googlebot Smartphone is the primary crawler (rollout completed 2024). A mobile version is not strictly required (Google says "very strongly recommended"), sites that don't work on mobile can still be indexed, but the real risk is content/parity loss , not hard exclusion.
Mobile/desktop content parity (highest value mobile check): equivalent primary content, matching robots meta tags, matching titles/descriptions, equivalent structured data, crawlable resources; avoid lazy loading primary content that requires user interaction.
Intrusive interstitials / ad density : flag full page interstitials, standalone consent redirect pages, persistent blocking dialogs, and excessive/distracting ad density (a named page experience aspect). Acceptable: small banners, standard CMS/legal dialogs.
"Read more" deep links : keep key content immediately visible on load (not behind tabs/accordions), don't hijack scroll on load, and preserve URL hash fragments, content hidden behind expandable sections is less likely to qualify.
Page experience is guidance, not a single ranking system. Only Core Web Vitals feeds ranking directly; HTTPS is a confirmed but lightweight signal (affects <~1% of queries). Relevance can still win even when page experience is sub par, so don't over weight security headers. Note: the standalone Page Experience report was removed from Search Console (monitor via the Core Web Vitals + HTTPS reports).
6. Core Web Vitals
LCP (Largest Contentful Paint): target <=2.5s
INP (Interaction to Next Paint): target <=200ms
INP replaced FID on March 12, 2024. FID was removed from Chrome's field data tools (CrUX API, PageSpeed Insights) on September 9, 2024 (Lighthouse is a lab tool that never reported FID). Do NOT reference FID anywhere.
CLS (Cumulative Layout Shift): target <=0.1
Evaluation uses 75th percentile of real user data
Use PageSpeed Insights API or CrUX data if MCP available
7. Structured Data
Detection: JSON LD (preferred), Microdata, RDFa
Validation against Google's supported types
See seo schema skill for full analysis
8. JavaScript Rendering
Check if content visible in initial HTML vs requires JS
Identify client side rendered (CSR) vs server side rendered (SSR)
Flag SPA frameworks (React, Vue, Angular) that may cause indexing issues
If dynamic rendering is detected, flag it as technical debt rather than a valid setup.
Google documents it as "a workaround and not a recommended solution" because of the added
complexity and resource cost.
See https://developers.google.com/search/docs/crawling indexing/javascript/dynamic rendering
Recommended rendering strategy:
Strategy Use Case
SSR Public SEO content, dynamic pages
SSG Static content, blogs, docs
CSR Authenticated / behind login content only
Preferred frameworks: Next.js, Astro, React Router v7 (Remix), SvelteKit
JavaScript SEO: Canonical & Indexing Guidance (December 2025)
Google updated its JavaScript SEO documentation in December 2025 with critical clarifications:
1. Canonical conflicts: If a canonical tag in raw HTML differs from one injected by JavaScript, Google may use EITHER one. Ensure canonical tags are identical between server rendered HTML and JS rendered output.
2. noindex with JavaScript: If raw HTML contains <meta name="robots" content="noindex" but JavaScript removes it, Google MAY still honor the noindex from raw HTML. Serve correct robots directives in the initial HTML response.
3. Non 200 status codes: Google does NOT render JavaScript on pages returning non 200 HTTP status codes. Any content or meta tags injected via JS on error pages will be invisible to Googlebot.
4. Structured data in JavaScript: Product, Article, and other structured data injected via JS may face delayed processing. For time sensitive structured data (especially e commerce Product markup), include it in the initial server rendered HTML.
Best practice: Serve critical SEO elements (canonical, meta robots, structured data, title, meta description) in the initial server rendered HTML rather than relying on JavaScript injection.
9. IndexNow Protocol
Check if site supports IndexNow for Bing, Yandex, Naver
Supported by search engines other than Google
Recommend implementation for faster indexing on non Google engines
Agent Friendly Pages & Agentic Browsing
AI agents (not just AI summarizers) increasingly read sites through three
channels: vision models on screenshots, raw HTML/DOM, and the accessibility
tree (the cleanest signal). Audit criteria: semantic HTML (real <button
and <a , not <div onclick ), label associations, interactive target sizing,
layout stability across templates, cursor: pointer correctness, live in
references/agent friendly pages.md .
Google now ships a Lighthouse Agentic Browsing category (default on since
Lighthouse 13.3.0, Chrome 150+; buckets: agent centric accessibility, CLS +
llms.txt, three WebMCP audits). It reports a fractional pass ratio (X of N),
not a 0 100 score , keep that distinct from this skill's own Agent UX 0 100
heuristic below. Lighthouse 13.4.1 re enabled the category through the PSI API.
It is also available through Lighthouse CLI with
only categories=agentic browsing , DevTools, and the PSI web UI. See
references/agent friendly pages.md .
Audit command
The scanner outputs an Agent UX score (0 100) plus itemized issues:
HTML findings: real