Audit and improve content for AI Overviews and answer engines, including citability, entity clarity, crawler access, brand signals, and passage structure.
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-seo-geo-8de3c81aacd4 ,按照其中的说明把「seo-geo」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
Google's official position, published under Search Central docs:
"From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO."
Read references/google-ai-optimization-guide.md for the full synthesis,
myth-busting list (llms.txt, chunking, AI-rephrasing, mention-farming,
all rejected by Google as ineffective), and the Who/How/Why test for
content quality.
Audits should frame GEO findings as SEO fundamentals applied to AI-search surfaces, not as a separate optimization discipline. When community recommendations contradict Google's primary source, defer to Google and note the contradiction in the report.
Third-party figures below were not re-verified on 2026-09-23. Quote them only with their source and date, or leave them out.
| Metric | Value | Source |
|---|---|---|
| AI Overviews reach | 2.5 billion+ monthly active users, reported from Google I/O 2026 keynote coverage; not confirmed on a Google-owned source; 200+ countries | Third-party I/O reporting |
| AI Overviews query coverage | ~50% of queries (third-party measurement; varies by country) | Industry data |
| AI Mode monthly users | 1B+, reported from Google I/O 2026 keynote coverage; not confirmed on a Google-owned source | Third-party I/O reporting |
| AI Mode model | Google upgrades it often (Gemini 3.5 Flash became the default on 2026-05-19, and newer Flash models have shipped since); never tie advice to a model | Google (blog.google) |
| AI-referred sessions growth | 527% (Jan-May 2025) | Third-party (attributed to SparkToro; not re-verified) |
| ChatGPT weekly active users | 900 million | OpenAI |
| Perplexity monthly queries | 500+ million | Perplexity |
Brand mentions correlate 3x more strongly with AI visibility than backlinks. (Ahrefs December 2025 study of 75,000 brands)
| Signal | Correlation with AI Citations |
|---|---|
| YouTube mentions | ~0.737 (strongest) |
| Reddit mentions | High |
| Wikipedia presence | High |
| LinkedIn presence | Moderate |
| Domain Rating (backlinks) | ~0.266 (weak) |
Only 11% of domains are cited by both ChatGPT and Google AI Overviews for the same query, so platform-specific optimization is essential.
Self-contained answer blocks are easy for AI systems to quote. Third-party studies suggest roughly 130-170 words; this is a readability heuristic, not a Google requirement (Google's AI optimization guide says you do not need to chunk content for AI). And ~44% of AI citations come from the first 30% of a page (SE Ranking study), front-load your most citable, self-contained answer rather than burying it below the fold.
Strong signals:
Weak signals:
92% of AI Overview citations come from top-10 ranking pages, but 47% come from pages ranking below position 5, demonstrating different selection logic.
Strong signals:
Weak signals:
Multi-modal content can support selection in AI answers (third-party claims only; no primary source gives a figure).
Check for:
Strong signals:
Weak signals:
Many AI crawlers fetch raw HTML without running JavaScript (for example GPTBot and PerplexityBot in public tests), while Googlebot renders JavaScript and feeds AI Overviews and AI Mode. Server-side rendering keeps content visible to all of them.
Check for:
references/llmstxt-evidence.md)Check robots.txt for these AI crawlers:
| Crawler | Owner | Purpose | Obeys robots.txt? |
|---|---|---|---|
| GPTBot | OpenAI | Model training only (NOT ChatGPT Search) | yes |
| OAI-SearchBot | OpenAI | ChatGPT Search citability (the crawler that decides it) | yes |
| ChatGPT-User | OpenAI | ChatGPT browsing (user-triggered) | "may not apply" per OpenAI (user-triggered) |
| ClaudeBot | Anthropic | Model training only (NOT Claude's search features) | yes |
| Claude-SearchBot | Anthropic | Claude/Claude.ai search-result citability (the crawler that decides it) | yes |
| Claude-User | Anthropic | Claude browsing on a user's behalf (user-triggered) | yes (Anthropic: all three bots honor robots.txt) |
| PerplexityBot | Perplexity | Perplexity AI search (not used to crawl for foundation-model training) | yes |
| Perplexity-User | Perplexity | Fetches for a user's question (user-triggered) | generally ignores |
| CCBot | Common Crawl | Training data (often blocked) | yes |
| Bytespider | ByteDance | TikTok/Douyin AI | yes |
| cohere-ai | Cohere | Cohere models | yes |
| Google-Extended | Gemini/Vertex training & grounding only (NOT Google Search) | yes | |
| Google-CloudVertexBot | Site-owner-requested Vertex AI Agent crawls | yes | |
| Google-Agent | User-triggered agent fetches (agentic browsing for a user) | no (user-triggered) | |
| Google-GeminiNotebook | Fetches individual user-added source URLs (replaced Google-NotebookLM, supported until August 2026) | no (user-triggered) | |
| Google Messages | User-triggered fetch |
Sources: OpenAI crawlers,
Google crawlers overview,
Anthropic crawler support article,
Apple Applebot-Extended support article.
Anthropic's current crawler support article documents only ClaudeBot, Claude-User,
and Claude-SearchBot; it does not list anthropic-ai, so the previously-unverified
anthropic-ai row has been removed rather than kept as a guess.
Recommendation: Allow OAI-SearchBot, Claude-SearchBot, and PerplexityBot for AI search visibility. GPTBot, ClaudeBot, CCBot, and Applebot-Extended are training-only signals -- allow or block them on licensing preference, not on search-visibility grounds.
Two pairs are routinely conflated. Each claim below may only be supported by its own bot's robots.txt status -- check them separately and report them separately.
| Claim you want to make | Bot to check | Bot that does NOT support this claim |
|---|---|---|
| "Content is citable in ChatGPT Search" | OAI-SearchBot | GPTBot |
| "Content is available for OpenAI model training" | GPTBot | OAI-SearchBot |
| "Content can be used for Gemini/Vertex training & grounding" | Google-Extended | Googlebot |
| "Content is eligible for Google Search / AI Overviews" | Googlebot | Google-Extended |
| "Content is citable in Claude's search features" | Claude-SearchBot | ClaudeBot |
| "Content is available for Anthropic model training" | ClaudeBot | Claude-SearchBot |
| "Content can be used for Apple Intelligence training" | Applebot-Extended | Applebot |
| "Content is discoverable via Siri, Spotlight, or Safari search" | Applebot | Applebot-Extended |
Google-Extended governs Gemini and Vertex AI training and grounding use only.
It does not affect inclusion in ordinary Google Search, or in AI Overviews and AI
Mode, both of which are served from the Googlebot index. Never score
Google-Extended as a "Google Search readiness" signal, and never cite a blocked
Google-Extended as evidence that a site is missing from Google Search.OAI-SearchBot is the crawler that determines ChatGPT Search citability.
GPTBot is OpenAI's separate training crawler. Checking GPTBot access tells
you nothing about whether ChatGPT Search can cite the page. A site that blocks
GPTBot and allows OAI-SearchBot is fully citable in ChatGPT Search.Claude-SearchBot is the crawler that determines citability in Claude's own
search features. ClaudeBot is Anthropic's separate training crawler (per
Anthropic's crawler support article). Checking ClaudeBot access tells you
nothing about Claude search citability, and vice versa; report each separately.Applebot-Extended is a training-data opt-out signal, not a crawler that
fetches pages itself. Per Apple's support article, disallowing
Applebot-Extended opts a site out of Apple Intelligence / generative-model
training use, but the page remains discoverable through Siri, Spotlight, and
Safari as long as Applebot itself is allowed. Never cite a blocked
Applebot-Extended as evidence a site is missing from Apple's search surfaces.Do not use these names interchangeably in report prose. When reporting crawler access, name the specific user-agent that was checked and the specific capability it governs.
Google's user-triggered fetchers generally ignore robots.txt rules (Google-Agent, Google-GeminiNotebook, Google Messages); OpenAI says robots.txt "may not apply" to ChatGPT-User, while Anthropic's Claude-User honors it. robots.txt cannot block them, use server-side access controls. Google's canonical crawling/robots reference moved to developers.google.com/crawling (migrated 2025-11-20); IP-range files now live at
/crawling/ipranges/andgooglebot.jsonwas renamedcommon-crawlers.json. Emerging: Web Bot Auth (RFC 9421) lets bots authenticate via aSignature-Agentheader + key directory (used by Google-Agent); reverse-DNS verification remains the fallback.
Read references/llmstxt-evidence.md for the primary-source evidence (Mueller, Illyes, SE Ranking 300k-domain study, OtterlyAI server-log audit) on why /llms.txt is not currently a citation lever for major AI search systems. claude-seo reports presence but assigns no citation-ranking weight.
Google now states this explicitly. Google's AI optimization guide, introduced 2026-05-15 and clarified 2026-06-15, says
llms.txtand other AI-text files are not needed for Google Search and do not help or hurt visibility or rankings. They may still serve non-Google systems. Never recommendllms.txtas a Google ranking or citation lever. Source: developers.google.com/search/docs/fundamentals/ai-optimization-guide
llms.txt is a community proposal for giving LLMs a curated map of a site; no major AI provider has confirmed using it.
Location: /llms.txt (root of domain)
Format:
# Title of site
> Brief description
## Main sections
- [Page title](url): Description
- [Another page](url): Description
## Optional: Key facts
- Fact 1
- Fact 2
Check for:
/llms.txtNew standard (December 2025) for machine-readable AI licensing terms.
Backed by: Reddit, Yahoo, Medium, Quora, Cloudflare, Akamai, Creative Commons
Check for: RSL implementation and appropriate licensing terms.
| Platform | Key Citation Sources | Optimization Focus |
|---|---|---|
| Google AI Overviews | Strongly ranking-correlated, cites pages that already rank well | Traditional SEO + passage optimization |
| Google AI Mode (Gemini models, upgraded often) | Weakly ranking-correlated; broader pool (~9 domains cited/query, Ahrefs) | Distinct surface: freshness, entity authority, citable passages beyond position 5 |
| ChatGPT | Wikipedia (47.9%), Reddit (11.3%) | Entity presence, authoritative sources |
| Perplexity | Reddit (46.7%), Wikipedia | Community validation, discussions |
| Bing Copilot | Bing index, authoritative sites | Bing SEO, IndexNow |
Two Google citation engines, not one. AI Mode and AI Overviews reach the same conclusion ~86% of the time but cite the same URLs only 13.7% of the time (Ahrefs study, 540K query pairs). Treat them as separate surfaces: ranking well in classic Search feeds AI Overviews, but AI Mode draws from a broader pool where freshness and entity authority outweigh raw position. Score both.
AI Mode is also a booking surface (2026-08-27). Flight price tracking with email alerts (180+ countries and territories), hotel booking through integrated partners, and fares shown in points or miles now happen inside AI Mode. Travel and hospitality clients should check partner eligibility; nothing here is a documented ranking change.
UX is now unified, surfaces still distinct. At Google I/O 2026 (2026-05-19) Google merged AI Overviews and AI Mode into "one seamless AI Search experience" (question → AI Overview → follow-up in AI Mode) with a new intelligent Search box. The experience is one flow, but the two citation engines remain technically distinct (different models/link sets), keep scoring both.
Google added many AI citation/source surfaces across AI Overviews and AI Mode (May 2026):
Controlling AI-feature appearance: there is no AI-specific opt-out file, but since 2026-08-31 every site has a Search Console control, "Search generative AI" (include by default, exclude, or inherit), that controls eligibility for AI Overviews and AI Mode; it is not a ranking signal or a training control. Beyond that, appearance is governed by standard preview/index directives, nosnippet, data-nosnippet, max-snippet, noindex (distinct from the third-party AI-crawler robots controls above). Source: developers.google.com/search/docs/appearance/ai-features
Search agents (live, not just WebMCP): Google's "Information Agents" run in the background to monitor topics, plus agentic booking/calling for select categories (rolling out to US users, summer 2026), so agent-friendly-page optimization (real interactive elements, accessibility tree, layout stability) now matters for actions, not only citations. Audit that with /seo agentic (the seo-agentic sub-skill), which also reads Lighthouse's Agentic Browsing fraction.
For AI Overviews or AI Mode visibility changes, check the dated product and
core-update entries first:
"${CLAUDE_PLUGIN_ROOT}/scripts/claude-seo" run seo_updates.py --kind product --kind core --json.
Treat a stale ledger (freshness.stale) as incomplete.
Generate GEO-ANALYSIS.md with:
GPTBot, Google-Extended, CCBot,
ClaudeBot, Applebot-Extended) and search citability (OAI-SearchBot,
Googlebot, PerplexityBot, Claude-SearchBot, Applebot) are distinct
findings and must never be merged into one line./llms.txt file (optional: ignored by Google Search; may help other AI crawlers)If DataForSEO MCP tools are available, use ai_optimization_chat_gpt_scraper to check what ChatGPT web search returns for target queries (real GEO visibility check) and ai_opt_llm_ment_search with ai_opt_llm_ment_top_domains for LLM mention tracking across AI platforms.
| Scenario | Action |
|---|---|
| URL unreachable (DNS failure, connection refused) | Report the error clearly. Do not guess site content. Suggest the user verify the URL and try again. |
| AI crawlers blocked by robots.txt | Report exactly which crawlers are blocked and which are allowed. Provide specific robots.txt directives to add for enabling AI search visibility. |
| No llms.txt found | Note the absence (optional file; Google Search ignores it) and provide a ready-to-use llms.txt template for non-Google AI crawlers. |
| No structured data detected | Report the gap and provide specific schema recommendations (Article, Organization, Person) for improving AI discoverability. |
For prompt-guided AI content optimization, use /seo flow optimize <url>, FLOW's 21 optimize-stage prompts complement GEO's citability and structure analysis with evidence-led AI prompts.
| no (user-triggered) |
| Applebot-Extended | Apple | Apple Intelligence / generative-AI training data opt-out only (NOT Siri, Spotlight, or Safari search; does not itself crawl, it labels content already fetched by Applebot) | yes |