--- name: ai-search-visibility-audit description: "Run an 8-stage AI search visibility audit for a brand, cluster by cluster, using the Ahrefs MCP connector plus web research, then write it up as a Word document with a prompt tracking spreadsheet. Use this skill whenever someone asks to audit AI search visibility, check how a brand shows up in AI answers, run an AEO or GEO audit, diagnose why a brand is not cited, run an AI bot accessibility audit, or says things like 'audit our AI search visibility', 'how do we show up in ChatGPT', 'run the AI search audit on brand X', 'why aren't we cited', 'are AI bots blocked on our site', or 'AI visibility baseline for Y'. Use it even when the request sounds quick, because the staged structure is the point and every stage that cannot be automated produces a ready-to-run template instead." --- # AI Search Visibility Audit You are running an 8-stage AI search visibility audit: Map, measure, diagnose, prioritise. The output is a structured report in the chat, the same report as a Word document, and a prompt tracking spreadsheet the user runs on a schedule. Everything is audited **cluster by cluster**. Platform citation overlap runs as low as 14%, so a single blended visibility number for a whole domain hides both the best result and the worst one. Be honest about the split between what you can measure and what needs a human run. You **can** automate scoping, the demand map, the accessibility checks, AI Overview presence, cited-domain analysis and the entity check. You **cannot** query ChatGPT, Gemini, Perplexity or Copilot yourself, so for the cross-platform baseline in stage 4 you produce a pre-filled run sheet and never invent numbers. ## What you need before you start 1. **Brand and domain.** Required. Ask and stop if missing. 2. **Clusters.** 3 to 5 commercial territories. If not given, propose them from the domain's top organic keywords and get a yes before continuing. 3. **Personas, markets, competitors.** Ask once. If the user says "just run it", default to 2 personas (technical evaluator, budget owner), the domain's biggest market, and the top 3 organic competitors. If the Ahrefs MCP connector is unavailable, say so plainly, skip the Ahrefs-backed steps with one line each and continue with the rest. Never substitute guessed numbers. ## Rules that apply throughout - **Consult the Ahrefs `doc` tool** before the first use of any endpoint - **Limit to 10 results per call**, batch fields into single calls, never repeat a call, derive rather than re-fetch - **Skip what the plan blocks.** Brand Radar endpoints are plan-gated. If blocked, note it in one line and fall back to the manual template - **Every claim carries its number**, and count every API call - **Every finding carries a cluster.** No domain-level findings - Flags: `[!!]` critical, `[!]` worth watching, `[+]` a strength --- ## Stage 1: Scope No API calls beyond one top-keywords pull if clusters were not supplied. Confirm clusters, personas, markets and both competitor lists (commercial competitors and answer competitors) in one short table. Keep a running list of anything unconfirmed for the Open Questions section. ## Stage 2: Build the prompt set 2 to 3 calls, Keywords Explorer (`matching-terms`, `search-suggestions`). Build a persona by journey matrix per cluster across 5 journey stages: Problem discovery, category education, brand and product comparison, objections and risk, then buying and post-purchase. Write 25 to 40 core prompts total, in the persona's language, with realistic constraints. Prompts are conversational questions with a constraint attached, never a keyword with "best" bolted on. Add a follow-up per prompt. Recommend FanoutFox (free Chrome extension, fanoutfox.com) for the user to run their 3 most commercial prompts through, since it reveals the sub-queries ChatGPT actually runs and whether a page was cited, mentioned or only fetched. Output: The prompt table, which becomes sheet 1 of the spreadsheet. ## Stage 3: AI bot accessibility audit No Ahrefs calls. Using web fetch: 1. **Robots policy.** Read robots.txt and table the rule per AI user agent: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, Bytespider, Amazonbot, meta-externalagent, Bingbot 2. **Reality check.** Fetch 3 to 5 key URLs per cluster and report the status returned. State clearly that a CDN or WAF can block bots that robots.txt welcomes, and that only per-user-agent testing or server logs prove it. Give the user the curl one-liner to run per agent 3. **Rendering.** Report whether the main content is present in the fetched HTML. Where it looks thin, flag possible client-side rendering `[!!]` and give the Screaming Frog method: crawl once in text-only mode, once with JavaScript rendering, compare word counts per template 4. **llms.txt.** Check existence, size and whether it reads as a concise guide. Report neutrally: Fetched is not read, and read is not visibility. Flag anything over roughly 50KB as likely to be truncated 5. **Crawl waste patterns.** If the user can supply a log export, look for the 4 classic patterns: tracking and affiliate redirect paths, discontinued product pages, dynamic search result pages with high fetch and near-zero indexing, and pagination duplicates splitting one signal. Without logs, say so and offer it as a follow-up Recommend the tools by name: Screaming Frog for the crawl and rendering comparison, Ahrefs Site Audit or Sitebulb for the wider technical sweep, and JetOctopus, Botify or Lumar for log analysis. ## Stage 4: Visibility baseline 2 to 4 calls. 1. Pull top organic keywords with `serp_features` included. Report AI Overview presence as a share of keywords pulled, table the affected keywords with volume and position, and say whether AI Overviews reach the money terms or only the informational tail. State the limit: This shows an AI Overview appears, not that the domain is cited inside it 2. If Brand Radar is available on the plan, pull brand mentions and share of voice per cluster. If gated, one line and move on 3. **Produce the manual run sheet.** Every core prompt as a row, with the 6 recording fields per platform: answer present, brands and order, description and accuracy, citations, source type, errors. State the cadence: Full set monthly, 10-prompt sentinel weekly 4. Your own contribution: For each cluster's top 3 comparison prompts, answer from your own knowledge which brands and sources are likely, labelled clearly as "one model, single run, [date]". A smoke test, never the baseline ## Stage 5: Citation source audit 2 to 3 calls where Brand Radar is available (`cited-domains`, `cited-pages` per cluster topic), otherwise derive from `serp-overview` on cluster head terms. Per cluster: - Top cited or top-ranking domains, classified as owned, earned, community, commercial or competitor-owned - The dominant winning format: comparison, research, docs, tool, video or third-party list - The shortlist check: Which third-party best-of pages dominate, whether the brand appears and whether the description is current. Fetch the top 1 to 2 and verify - Name the 3 third-party pages where inclusion or correction would move the most ## Stage 6: Entity check No Ahrefs calls. 1. Answer cold from your own knowledge: Who is [brand], what does it do, what is it known for, [brand] vs [main competitor]. Label as single-model evidence 2. Web-search the brand and check consistency across site, LinkedIn, Crunchbase, Wikidata and the main directories 3. Fetch the homepage and about page. Check Organization, Product and Person schema presence and whether name, description and positioning agree 4. Table: Claim, what the model says, correct or not, likely source, fix location ## Stage 7: Diagnosis No calls. The judgement stage. Map every gap from stages 3 to 6 to one of 7 causes, and state the evidence that confirms it: | Cause | Confirm it with | Owner | |---|---|---| | Access | A 403, 429 or redirect chain for an AI agent in stage 3, or content appearing only after JavaScript | Engineering | | Relevance | Cited on informational prompts in a cluster, absent from the commercial ones | Content, SEO | | Evidence | Cited pages carry original numbers, named sources, first-hand testing and dates where yours does not | Content, Product marketing | | Entity clarity | Stage 6 returned inconsistent or wrong answers, or schema and third-party descriptions disagree | SEO, Brand | | Format | Stage 5 shows one dominant format in the cluster and the brand's best page is a different one | Content, Design | | External context | Stage 5's cited domains are mostly third party, and the brand is missing or described wrongly on them | Partnerships, PR, Community | | Freshness | Numbers in the AI answer differ from current pricing or specification | Content ops | Every gap gets a cause and an owner. Note explicitly that one of the 7 is solved by publishing more content. If all 7 fire on every cluster, say the scope was too broad and recommend narrowing. ## Stage 8: Scorecard and roadmap No calls. Score each cluster out of 100: | Component | Weight | |---|---| | Presence in answers | 0.30 | | Prominence and description accuracy | 0.20 | | Owned citations | 0.20 | | Third-party shortlist strength | 0.20 | | Access | 0.10 | Cap each component at 100% of its own target, so a strong component cannot mask a failing one. Where the manual sheet has not been run yet, score what is measurable now, mark the rest "pending first manual run" and show the formula so the user can complete it. Close with the 90-day roadmap ordered by impact over effort, with 2 overrides: Access fixes first regardless of score, then freshness corrections on heavily-cited third-party pages. ## Summary section Snapshot table, one line per cluster: score or partial score, dominant cause, top action. Up to 5 priority actions with expected effect. Open questions, 4 to 6 bullets. Then the standing caveat, verbatim: Prompts are sampled, citations fluctuate and attribution is incomplete, so this is an honest directional view rather than a precise traffic number. ## Final line State the total number of API calls made. --- ## Then produce the files 1. **Word document** of the full report. Use the docx skill if available, otherwise python-docx. Body DM Sans 12 with Calibri fallback, headings in one accent colour, header rows bold with a light fill, flags as bold labels, a cover line carrying brand, date and the single most important finding, page numbers in the footer. Name it `-ai-search-visibility-audit-.docx` 2. **Prompt tracking spreadsheet.** Sheet 1 the core prompt set with all 6 recording fields per platform, sheet 2 the 10-prompt sentinel subset, sheet 3 the scorecard with live formulas. Name it `-ai-visibility-tracker-.xlsx` Tell the user where both files are. ## Where to take it next Mention once at the end, only if the user seems likely to want more: The monthly rerun against the same sheet, since the trend line is the product. A server log analysis to split training, indexing and real-time fetches. A rendering deep-dive per template. And the paid scale-up (Profound, Peec AI, AccuRanker AccuLLM, Rankscale), which changes cadence and coverage rather than method.