How we check a company's AI visibility: the method behind our audits
The five areas we examine on a company website before saying anything about AI visibility, why each matters, and what we do not claim. No pilot results yet: this is the method, stated in advance.
Founder & SEO Lead
We had planned a piece on what we found scanning fifty Gulf company websites. We have not run that pilot yet, so we will not write it as if we had. Instead, here is the method we will use, stated before any results exist, so that when we publish findings you can judge them against it.
Area 1: Access
If a crawler cannot get in, nothing else matters. We read the site's robots.txt and check, for each of a set of well-known crawlers, whether it is allowed. We separate three kinds: crawlers that power search and answers, crawlers that fetch a page when a user asks for it, and crawlers that collect data for training. Blocking the first kind is a problem for visibility. Blocking only the third is a legitimate choice, so we report it as a note, not an error.
We also request a page while identifying as some of these crawlers, to see whether the server or a security layer turns them away even though robots.txt allows them. A firewall can block a crawler that the file welcomes.
Area 2: Structure
Structured data, written as schema.org JSON-LD, tells machines who the company is and what a page is, without deduction from prose. We check that it exists, that it is valid, that it identifies the organisation, and that it links to external profiles.
Area 3: Metadata and indexing
Titles, descriptions, canonical addresses and language declarations decide how a page is presented. We look for the settings that override everything else: a noindex instruction, or a snippet restriction that lets a page be found but not quoted.
Area 4: Content
Many crawlers used by AI products do not run scripts. We measure how much readable text is in the page as first delivered, whether headings are present, whether some are phrased as questions, whether there are lists or tables, and whether the page shows who wrote it and when.
Area 5: The rest of the site
A polished home page does not help if the other pages cannot be read. We sample further pages to see whether problems repeat.
What this method does not measure
It does not measure whether an assistant names your company. That needs the separate question-based test we describe in our article on checking whether ChatGPT recommends your company. The technical check and the question-based test answer different questions: can the site be read, and is the company named.
It also does not weigh the reputation of your company or what third parties publish about you, which matter and which we review separately in an audit.
Why we publish the method first
Scores and rankings of companies are easy to misuse. By fixing the method and its limits in advance, and by committing to publish any pilot findings in aggregate and without naming companies, we aim to keep the work checkable. If you would like the method applied to your own site, ask for an audit below.
Frequently asked questions
Not yet. We planned to publish aggregated findings from a pilot scan of Gulf company websites, but that pilot has not been run, so there are no results to report. When it is done we will publish aggregated findings without naming companies.
No. The technical check tells you whether your site can be read and understood. Whether an assistant names your company also depends on what independent sources say about you and on choices the assistant makes that nobody outside can see.
That is the client's decision. Blocking crawlers that collect data for model training is a legitimate choice. Blocking the ones that power search and answers means those assistants may not be able to read your site. We report which is which and leave the decision to you.