I often see marketing teams skip over the audit phase when starting to think about AIO. They read about how important AI Optimization is and immediately go ahead and create new content, add schema markup to their website, or optimize their service pages for AI citation. While some of what they do could possibly be correct, without an initial baseline there’s no possible way to know where the gaps exist in the five layers of AIO and therefore which layers are performing adequately so additional time spent on those areas would likely be wasted.
The Semrush AI Visibility Index 2026, which looked at 126 million actual user requests from ChatGPT, Gemini, Google AI Mode and Google AI Overview, found that only nine percent of marketing teams had access to all of the measurement data necessary to assess the level of AI search visibility they needed. Additionally, forty-five percent of these same marketing teams did not have sufficient means to measure AI search visibility at all.
You simply can’t fill a gap if you don’t identify one. This audit was developed to enable any marketing team to perform the audit with no special tools other than GA4, Google Search Console, and a web browser. It should take about ninety minutes to complete. It will cover all five layers of the Infidigit AI Discovery Framework (Foundation, Semantic Architecture, AI Extractability, Entity/Citation Authority, and Measurement). Each step will produce a clear finding and drive a corresponding action item. The purpose of this audit is to generate a specific list of actions your team can take within the next week of completion.
Prior to beginning this audit, please open a blank document to record your findings as you progress through each step. After completing each step use the framework provided to score yourself. Based on your overall scores, you will determine which layer to address first.
Layer 1: Foundation Audit
The foundation layer will determine whether AI systems are able to access, crawl, and view your site. The majority of content and schema work will be rendered useless should there be failure in this layer, as a crawler that cannot reach your content will never have the opportunity to either index or cite you, regardless of the quality of the structure of your content.
Step 1 AI Crawler Access Check | 15 minutes
Go to the robots.txt file on your website by simply entering /robots.txt after your domain into a browser. Then review each line with User-Agent entries and each Disallow directive.
You will need to look for the following: GPTBot (OpenAI), OAI-SearchBot (OpenAI), ClaudeBot (Anthropic), Claude-SearchBot (Anthropic), PerplexityBot, GoogleOther (Google uses this for AI overview training & retrieval), Googlebot-Extended.
A wildcard disallow, defined as User-agent:* followed by Disallow:/, would block any crawler that does not have an explicit allow directive. If you have a wildcard disallow set up and the AI crawlers mentioned above do not exist within explicit Allow directives above it, then they will ALL be blocked from accessing your site. A wildcard disallow is the most frequent and the most impactful Layer 1 failure we see.”
It’s also good to take a look at your CDN or hosting providers’ bot-management settings. Most CDNs such as Cloudflare and Fastly provide additional layers of protection against bots, including AI crawlers. A site may have a pristine robots.txt but still be blocking GPTBot via Cloudflare’s bot-blocking service without even being aware of it.
To finish, open 3 of your top pages in a browser while disabling JavaScript (Chrome: Settings > More Tools > Developer Tools > Rendering Tab > Check box for ‘Emulate CSS media type’, then Click down arrow in rendering section and choose ‘no JavaScript’). The content that loads will be accessible to most AI crawlers. If important sections of your pages disappear when you disable JavaScript, then those sections of content are likely inaccessible to AI crawlers.
| Score this step
3 points: All major AI crawlers are explicitly allowed, no infrastructure blocking, key content renders without JavaScript 2 points: Crawlers allowed but CDN not checked, or minor content rendering gaps 1 point: Some crawlers blocked but not all 0 points: Wildcard disallow without explicit crawler allowances, or infrastructure-level blocking confirmed |
Step 2 llms.txt and Organization Schema Check | 10 minutes
Go to yourdomain.com/llms.txt. Is there a file in the location you just visited? Read it. Can you confirm from this document how well it describes what your company does now? Are all of your key pages listed? If your company’s description is old or if the file doesn’t exist, then this is a gap.
Next, go to the source of your home page (right-click on the page > View Page Source > Search for “organization”). Do you see the Organization schema as a JSON-LD block? If so, do you see the Organization schema include: name, url, logo, description, foundingDate, and at least one sameAs reference referencing a verified external account, for example LinkedIn?
The sameAs reference is what provides the link between your Website Entity and your External Entity Footprint. In addition, without a sameAs reference, an AI System encountering both your LinkedIn page and your website will have no structured reference that confirms them to be part of the same organization.
| Score this step
3 points: llms.txt present and current, Organization schema with sameAs references confirmed 2 points: One present but not both 0 points: Neither present |
Layer 2: Semantic Architecture Audit
The semantic architecture layer determines whether your content is organized in such a way that AI systems can understand what subject domain (brand) of your brand you are covering, and which sub-topics you cover. If your site has strong crawlability but no semantic architecture, it will be crawled by AI crawlers; however, due to lack of topic relationships, it will not have consistent citations.
Step 3 Topical Map Existence Check | 10 minutes
Open a spreadsheet. List every service, product or topic your brand provides. Now map these into a hierarchy: broad parent topics (i.e., automotive), sub-topics (e.g., battery replacement) and micro-topics (e.g., alternator vs. Battery). Do you have at least one page for each level of that hierarchy?
The most common gap in this area is a content estate which consists entirely of commercial pages (service pages), location pages with no educational or informational content sitting under them to signal depth of topical authority. Example of a site with 30 service pages and zero insights/educational blog posts has flat architecture and will be interpreted by all AI systems as directory not an authority.
Check your blog/insight section. Count how many posts there are. Are they internally linking from service pages through contextual internal links or do they sit as stand-alone pages with no semantic relationship to any other commerical service pages on the website. Desired model of web site design is a hub & spoke structure where all informational content provides both links to and from the appropriate service pages.
| Score this step
3 points: Documented topical map exists with parent, sub-topic, and micro-topic levels covered; educational content links contextually to commercial pages 2 points: Some content depth exists but not systematically organized into a hub-and-spoke structure 1 point: Content exists but no internal linking between informational and commercial pages 0 points: Only commercial pages; no educational content; internal links are navigational only |
Step 4 H2 Heading Structure Check | 10 minutes
Open three of your most important pages (ideally your main service page, your most visited blog entry and your home page). Read each H2 heading on every page and count how many are question based and how many statement based. A statement title like ‘why llm SEO Matters’ is just a label. A question title like “what is llm SEO and how does it improve AI search visibility?” is both what buyers type into their AI assistant as well as a structural signal to AI retrieval systems that this section provides an answer to specific questions.
Then read the first two or three sentences under each H2. Do those sentences give a complete standalone answer to the H2 question? If you removed just that H2 and those first sentences from the webpage, would they make sense as an independent answer? If yes, then that section can be extracted. If no, then it cannot.
This is the principle of the extraction unit. It is the most consistently underimplemented element in content we audit. The global ecommerce platform that was responsible for 1,139 times more traffic generated by artificial intelligence than any other type of traffic, structured its content with this same principle: all sections must be able to provide an independent answer and all opening sentences must respond directly to the H2 question
| Score this step
3 points: Majority of H2s are question-format; first 2-3 sentences under each H2 give a standalone complete answer 2 points: Some question H2s exist but inconsistently applied; opening answers sometimes complete, sometimes not 1 point: Most H2s are statement-based; opening content does not give standalone answers 0 points: All H2s are statement headlines; no extractable unit structure anywhere on the audited pages |
Layer 3: AI Extractability Audit
AI extractability determines whether you can get clean, independently meaningful answers to your questions from a layer of structured data and content precision, not just volume. Extractability is also a key element of LLM SEO, which structures your content for easier processing by algorithms and retrieval and reference by these systems.
Step 5 Schema Markup Audit | 15 minutes
Go to Google’s Rich Results Test tool (search.google.com/test/rich-results). Test three URLs: your homepage, your primary service page, and your most-visited piece of content.
For each page, record which schema types are detected. The minimum viable schema set for AIO is:
- Organization schema on the homepage with name, url, logo, description, and sameAs references
- Article or BlogPosting schema on all authored content with author (linking to Person schema), datePublished, and dateModified
- Person schema for Kaushal on all authored content pages, linking to LinkedIn
- FAQ schema on any page that contains question-and-answer sections
- HowTo schema on any page that describes a step-by-step process
The B2B ecommerce platform that achieved 245x improvement in AI Overview keyword presence had schema implementation as one of the core solutions: Product, Category, and Blog schema implemented to increase eligibility for rich results. The ICICI Prudential engagement that produced 39x LLM traffic growth included Organization, Product, FAQ, and Dataset schema as part of the AIO framework. Schema is not a secondary consideration; it is what makes your content claims machine-readable.
Note every schema type that should exist on each page but does not. This becomes your schema implementation priority list.
| Score this step
3 points: Organization, Article/BlogPosting, Person, and FAQ schema all present where applicable; no validation errors 2 points: Some schema present but not the full set; or present with validation errors 1 point: Basic schema only (e.g. website schema) with no structured data on content pages 0 points: No schema beyond basic meta tags; FAQ sections exist in HTML but are not marked up |
Step 6 Content Precision Check | 10 minutes
Open your primary service page or homepage. Read the first 200 words carefully.
Ask yourself three questions about what you just read:
First: Is it a verifiable claim with a number that can be referenced? “We are able to provide exceptional results” is not, however, “Infidigit’s LLM SEO engagements have resulted in a 37x through 1139x increase in AI session referrals from our clients’ websites.” A claim of this nature may be referenceable; a general claim is not.
Second: Does it follow an Entity Attribute Value (EAV) format? Does it identify an entity (your company/brand or one of your clients), describe the characteristics or attributes of this entity (the services you provided or the outcomes you were able to achieve), and quantify these attributes (with numbers, timeframes, markets)? AI Systems retrieve and reference content that follows an EAV format. They tend to ignore generic marketing statements.
Third: Within the first 200 words of the piece, are there any elements that an AI retrieval system can utilize to generate a response to a potential customer’s inquiry? Would the first 200 words alone provide sufficient information for an AI like ChatGTP to create a suitable, fact-based response when responding to a question such as “What does Infidigit do for large enterprise companies”?
The reason this is important is because AI Systems will favor and prefer extracting dense, factual, quantified content over marketing content. The Dun & Bradstreet account that grew by 57X in LLM generated traffic and had over 1500+ AI Overview placements began with recognizing the same gap: Optimized SEO Content that was not optimized for AI Extraction.
| Score this step
3 points: First 200 words contain specific verifiable claims, EAV structure, and at least one extractable standalone answer 2 points: Some specific claims but mixed with generic marketing language 1 point: Mostly marketing language with occasional specific claims deeper in the page 0 points: First 200 words are entirely generic marketing language with no specific, verifiable claims |
Layer 4: Entity and Citation Authority Audit
What does the entity and citation authority layer do for your brand? Does it help prove that third parties confirm what you claim about your brand? Are they saying the same thing about you across all of their platforms? In a companion piece to this series, we have covered how to clearly define your online identity (off-site). For this article, we will take a closer look at what areas have the biggest opportunity for improvement.
Step 7 Off-Site Entity Consistency Check | 10 minutes
Open five tabs and navigate to: your LinkedIn company page, the top result when you search your brand name on Google (usually a Wikipedia page or Crunchbase if either exists), your Clutch profile, the most recent press article about your brand, and the top Reddit thread mentioning your brand name if one exists.
Read the description of your brand on each source. Answer these questions:
- Does each source describe the same core service or product offering in consistent language?
- Does each source use the same key terminology? If your website says AIO, does your LinkedIn say AI search optimization or digital marketing?
- Does each source describe your current positioning, or does any of them reflect an older version of what you do?
- Is your brand missing from any of these sources entirely?
Data for 2026 from Semrush shows a similarity among the top three brands to be visible through AI on all platforms; each of these brands has been described by third parties using the same language. A comparison of the scores of Patagonia’s AI Visibility showed an average of 79 and 80 during each of the four major AI platforms measured throughout the time frame that measurements were taken. Each of these outdoor specialty review websites use the same descriptions of Patagonia. Consistency is created above the level of AI in the ecosystem that provides the source material used as input to create an AI.
| Score this step
3 points: All five sources describe the brand consistently; terminology matches; no outdated descriptions 2 points: Mostly consistent but one or two sources use older language or different terminology 1 point: Significant inconsistencies across sources, or two or more sources are missing 0 points: Major inconsistencies across all sources, or three or more sources are missing |
Layer 5: Measurement Audit
Step 8 AI Traffic Measurement Check | 10 minutes
Open Google Analytics 4 go to Reports, then scroll down to “Life Cycle”, then find “Acquisition,” and then find “User Acquisition.” Find the “Default Channel Group” column. Is there an AI Assistant Channel listed? If so, select it and look at the session count in this Channel over the past 90 days. The native addition of this Channel to GA4 occurred on May 13, 2026. Therefore, the channel will be listed with no configuration required.
If you don’t see an AI Assistant Channel listed, your GA4 Property has likely not yet been upgraded to the latest default channel grouping configurations. Have your Sessions from chatgpt.com, perplexity.ai, claude.ai, and/or gemini.google.com been showing up under Referral Traffic or Direct Traffic?
Now go into Google Search Console and find the “Performance Tab”, then find “Search Results”. Are you seeing a Generative AI Tab (Introduced on June 3, 2026). If so, click the Tab and take note of the Total Impressions shown since April 5, 2026. These are your AI-Overview Impression Counts: The Number of Times Your Pages were Displayed within Google’s AI Features.
Last, Go into Search Console’s “Regular Performance Report” and Filter for your Brand Name as a Query. Take a look at the Trending Click Volume for your Brand Searches over the Last Three Months. The branded search lift is the downstream effect of an AI Citation Activity which does not produce a recorded Click.
| Score this step
3 points: GA4 AI Assistant channel present with data; Search Console Generative AI report accessible with impression data; branded search tracked as a standing metric 2 points: GA4 AI channel present but Search Console Generative AI report not yet accessible; or vice versa 1 point: Checking GA4 AI sessions manually through Referral reports; no dedicated AI tracking setup 0 points: No AI traffic measurement of any kind; no awareness of current AI session volume |
Step 9 Citation Rate Baseline Run | 10 minutes
This is the most important single step in the audit because it produces the only direct, empirical measure of your current AI visibility: how often does your brand actually appear in AI-generated answers?
Open ChatGPT, Perplexity, and Gemini in three browser tabs. Run these five prompts in each, typing them exactly as written:
- Who are the best [your service category] agencies or companies for [your target customer segment]?
- What are the top options for [the primary problem your brand solves]?
- What is [your brand name] known for?
- What do people say about [your brand name]?
- Which [your service category] company would you recommend for an enterprise brand?
Document your findings for each prompt and each platform: whether your company appears; what it says when it does; whether its description is accurate; and which other competitive companies’ descriptions show up next to or in place of yours.
AIO’s baseline for this kind of audit is based on these same prompt-by-prompt audits. Before their AIO engagement, ICICI Prudential had less than 200 AI Overview Keyword placements. They had a total of 1,093 AI Overview Keyword placements after the baseline measure was taken; entity optimizations were performed; content structure was optimized; and schemas were implemented. As well as the 39 times increase in LLM Platform Sessions, this baseline allows us to see (and attribute) the results of our optimization efforts. Without it, we could not tell which of the increases in AI Visibility came from the fact that you already had them prior to the start of the engagement; and which came from changes made through AIO’s optimization process.
| Score this step
3 points: Brand appears in 3 or more of the 5 prompts across 2 or more platforms; descriptions are accurate 2 points: Brand appears in 1-2 prompts; or appears but descriptions are partially inaccurate 1 point: Brand appears only when name is specifically mentioned in the prompt 0 points: Brand does not appear in any of the 5 prompts across any platform |
Scoring: What Your Result Means and What to Do Next
Total your scores across all nine steps. The maximum is 27 points.
| Score | Maturity Level | Immediate Priority |
| 22 to 27 | Strong foundation | Entity clarity and off-site citation pipeline. Your on-site infrastructure is solid. The next gains come from Layer 4. Run the full off-site entity audit from Article 6 in this series. |
| 16 to 21 | Good base, gaps to close | Focus on the layer where you scored lowest. If it was Layer 3, restructure H2s and add FAQ schema first. If it was Layer 2, build the topical map and informational content. Do not spread effort across all layers simultaneously. |
| 9 to 15 | Foundational work required | Start with Layer 1. If AI crawlers cannot access your content, nothing else matters. Fix robots.txt and CDN settings in week one. Then move to schema implementation in week two. |
| 0 to 8 | Not yet operational | Layer 1 is your only priority for the first 30 days. Confirm crawler access, implement Organization schema, verify llms.txt, and run the citation baseline. Do not invest in content or entity work until these are confirmed. |
What the Audit Tells You Beyond the Score
The numerical score is a useful triage tool, but the most valuable output of the audit is the specific list of findings from each step. Each gap you identified maps to a specific task:
- A Step 1 finding about blocked AI crawlers maps to a robots.txt or CDN configuration task assigned to a developer, with a target completion date of this week.
- A Step 2 finding about missing Organization schema maps to a JSON-LD implementation task, typically under an hour of dev time once the content is written.
- A Step 3 finding about no topical map maps to a content planning session that produces a documented hierarchy, followed by an editorial calendar that builds out the gaps.
- A Step 4 finding about statement-based H2s maps to a content editor pass across the three to five highest-priority pages, restructuring headings and adding direct opening answers.
- A Step 5 finding about missing FAQ schema maps to a structured data implementation across every page that has a Q-and-A section, which may be an Elementor or WordPress plugin task rather than a custom dev task.
- A Step 6 finding about generic content in the first 200 words maps to a rewriting task for the opening paragraphs of each priority page, replacing marketing language with specific, EAV-structured claims.
- A Step 7 finding about inconsistent external descriptions maps to the off-site entity audit and correction process described in Article 6 of this series.
- A Step 8 finding about missing AI traffic measurement maps to a GA4 channel group setup task and a Search Console check, both completable in under an hour.
- A Step 9 finding about low citation rate baseline maps to a prioritization decision: which of the upstream layers (1 through 4) is most likely to be the root cause, and which layer-specific task should come first.
There are three things common to all of our AIO engagement experiences at Infidigit. First, there are brands that do not start their AIO audits from the beginning and then go sequentially through each level (technical) and therefore have less success with respect to speed and durability of results. Second, the B2B ecommerce platform that increased its AI Overview keyword by 245x worked through a complete SEO and AIO audit prior to working on any of the other levels including content and schema. Third, an audit of a website should be considered the first deliverable, and not as preparation for any subsequent work.
The 10-Question AIO Maturity Self-Check
The Semrush 2026 AI Visibility Index includes a 10-question maturity self-assessment that I find useful as a complement to the layered audit above. It focuses on measurement and operational maturity rather than the technical and content specifics the layered audit covers. Score one point for each yes.
- Is your AI Visibility spread across your key platforms within 10 points (meaning you are not heavily concentrated on one platform)?
- Are at least 30 percent of your AI citations from your own domain, rather than entirely from third-party sources?
- Have your brand’s AI mentions grown in the past three months?
- Does your brand have a single named internal owner of AI search?
- Do you track mentions and citations as separate metrics?
- Have you mapped your competitive co-occurrence set, meaning the brands that appear alongside yours in AI responses?
- Do you measure AI visibility per platform separately, not as a single aggregated score?
- Do you appear among your industry’s top brands on multiple platforms, not just one?
- Are authoritative third-party sources cited alongside your own content in AI responses about your brand?
- Is your brand described with consistent language across the sources AI quotes from?
| Scoring this check
9 to 10: Category-leading AI visibility maturity. Focus on compounding what is working. 6 to 8: Middle tier, where most enterprise brands sit. Identify the three questions answered no and treat them as your next quarter priorities. 3 to 5: Behind the curve. Foundational measurement and operational infrastructure is the priority before any optimization. 0 to 2: AI search not yet operationalized. The layered audit above is your starting point. |
One More Thing Before You Start
The goal of this audit is not to provide a final grade or ranking. Rather, it is to create a definitive “to-do” list of all areas needing improvement along with a specific prioritization plan. By the time you complete the audit, you should be able to demonstrate that you have: Confirmed your robots.txt file, developed an actionable list for implementing schema markup, identified areas in need of rewriting to use H2 headings instead of other types of heading structures, completed a cross-reference check on the credibility of external sources, established Google Analytics 4 (GA4) and Search Console measurement settings, and have developed a baseline reference point for citations based upon the responses provided to each of the five questions.
This list will serve as the basis for the direction document sent to your internal team the very same week. Not six months later. According to Semrush data, approximately 81% of organizations reporting that they had fully integrated their SEO and AI search strategies reported either an increase in traffic or lead volume from AI search tools. That type of integration begins by identifying exactly where you are today, and that is what the audit provides.
If after conducting the audit, you find that the identified gaps exceed what your internal team can quickly resolve, then the referenced case studies below detail examples of what an outside AIO engagement can generate during the term of the engagement. During the engagement term, the ICICI Prudential engagement was able to go from less than 200 AI Overview placement occurrences per month to 1,093 monthly placements. Similarly, a global e-commerce platform went from essentially no AI-generated traffic to a 1139 times increase in LLM session activity during the engagement duration. These results represent typical outcomes rather than extraordinary ones.






