How the GEO scanner scores
The scanner distributes 100 points across five areas. We publish every check, what it is worth and how it is measured, so you can verify the result and reproduce it. If we change a weight, the engine version changes, and it appears on every report.
Summary: 100 points across five areas
| Area | Points |
|---|---|
| Access for AI engines | 30 |
| Structured data | 20 |
| Metadata and indexing | 20 |
| Citable content | 20 |
| Rest of the site | 10 |
Grade: A from 85 (well prepared for the ai era); B from 65 (solid base with improvements pending); C from 40 (limited visibility in ai engines); D below that (practically invisible to ai).
Access for AI engines — 30 points
If an engine cannot get into your website, nothing else matters. This category measures whether crawlers are allowed in and whether, in practice, your server lets them through.
Readable robots.txt
2 ptsWe request /robots.txt at the root of the domain. It counts if it answers 200 with plain text (not with the HTML home page, as many single-page apps do).
Permissions for 15 crawlers
14 ptsWe evaluate each user-agent with Google's semantics: the most specific group wins, and within it the longest rule. Bots are weighted by function: search and answer bots (Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot) ×3; bots that visit on the user's request (ChatGPT-User, Claude-User, Perplexity-User) ×2; training bots (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot) ×1. Blocking a search bot is critical; blocking only training bots is a warning, because it is a legitimate decision.
Real access test
6 ptsWe request your page three more times, identifying ourselves as GPTBot, ClaudeBot and PerplexityBot (with our signature at the end of the user-agent). If the server answers 4xx, serves a CDN challenge page or returns less than 30% of the text it gave us, we flag it. If the test cannot be completed, we give half the points and say so. If your server rejects even the scanner, we retry once presenting ourselves as a browser so we can analyse the site, and we flag it as a warning.
llms.txt
4 ptsHalf the points for publishing it as plain text; the other half if it follows the proposal: it starts with a Markdown title and has at least 3 links.
XML sitemap
4 ptsThe one declared in robots.txt or, if there is none, /sitemap.xml and /sitemap_index.xml. It counts if it contains <urlset> or <sitemapindex>.
| User-agent | Platform | Type | Weight |
|---|---|---|---|
| Googlebot | Google Search, AI Overviews and AI Mode | Search and answers | ×3 |
| Bingbot | Bing and Microsoft Copilot | Search and answers | ×3 |
| OAI-SearchBot | ChatGPT Search | Search and answers | ×3 |
| Claude-SearchBot | Claude (search) | Search and answers | ×3 |
| PerplexityBot | Perplexity | Search and answers | ×3 |
| DuckAssistBot | DuckDuckGo (DuckAssist) | Search and answers | ×3 |
| ChatGPT-User | ChatGPT (visits on the user's request) | On-demand visits | ×2 |
| Claude-User | Claude (visits on the user's request) | On-demand visits | ×2 |
| Perplexity-User | Perplexity (visits on the user's request) | On-demand visits | ×2 |
| GPTBot | OpenAI (model training) | Training | ×1 |
| ClaudeBot | Anthropic (model training) | Training | ×1 |
| Google-Extended | Gemini (training and Gemini answers; does not affect AI Overviews) | Training | ×1 |
| Applebot-Extended | Apple Intelligence (training) | Training | ×1 |
| Meta-ExternalAgent | Meta AI (training) | Training | ×1 |
| CCBot | Common Crawl (datasets used to train many models) | Training | ×1 |
Structured data — 20 points
Structured data (schema.org in JSON-LD) tells search engines and models who you are and what each page is, without having to deduce it from the text.
Valid JSON-LD
6 ptsAt least one application/ld+json block that is valid JSON. A malformed block is critical: engines discard it entirely.
Brand entity
4 ptsA type that identifies who is behind the site: Organization, LocalBusiness, ProfessionalService, Person and similar.
Complete entity profile
4 ptsOne point for each piece of entity data: name, url (or @id), logo (or image) and at least one sameAs pointing to external profiles.
Content schema
6 ptsA type that describes the page: WebSite, WebPage, Article, BlogPosting, FAQPage, Service, Product, BreadcrumbList, HowTo, among others.
Metadata and indexing — 20 points
Metadata and indexing decide whether the page can appear and how it is presented. A noindex or a nosnippet overrides everything else.
<title>
3 ptsBetween 10 and 70 characters: full points. Out of range: 1 point. No title: critical.
Meta description
3 ptsBetween 50 and 170 characters: full points. Out of range: 1 point.
Canonical
2 ptsDeclared and pointing to the same site. A canonical to another domain hands over authorship and is flagged.
Open Graph
2 ptsog:title, og:description and og:image present.
Language
1 ptslang attribute on the <html> tag.
A single H1
2 ptsExactly one H1 outside scripts and templates.
Indexable
4 ptsNo noindex either in the robots/googlebot meta or in the X-Robots-Tag HTTP header (general or aimed at Googlebot or an AI bot).
Can be quoted
3 ptsNo nosnippet or max-snippet:0 (critical: you can be found but not quoted). A max-snippet below 50 characters leaves 1 point; data-nosnippet blocks, 2.
Citable content — 20 points
Most AI crawlers do not run JavaScript and quote specific fragments. This category measures whether your content is in the HTML and in an easy-to-extract format.
Text without JavaScript
8 ptsVisible words in the initial HTML, outside scripts, styles and <head>. 300 or more: full points; between 150 and 299: half; fewer: critical.
Subheadings
3 ptsAt least two H2/H3.
Question subheadings
2 ptsAt least one H2/H3/H4 phrased as a question (“how”, “what”, “why”, “how much”…): the format that best fits how people ask an assistant.
Lists or tables
2 ptsAt least one list (2+ items) or table (2+ rows) outside menus, header and footer.
Trust signals (E-E-A-T)
3 ptsOn regular pages: link to “about us” (2 points) and to contact (1). On articles (Article/BlogPosting or og:type article): a declared author, a date within the last 18 months and a link to “about us” or contact, one point each.
HTTPS
2 ptsThe final URL, after redirects, is served over https.
Rest of the site — 10 points
A flawless home page is no use if the rest of the site cannot be read. We check a sample to see whether the problems repeat.
Sample of up to 4 pages
10 ptsWe choose URLs from the sitemap and from the page's links, one per section (/services, /blog, /contact…). On each we check 5 things: that it responds, that it is indexable, that it has a title, JSON-LD and at least 150 words without JavaScript. The score is the share of checks passed. If there is no way to discover other pages, we give half and flag it.
How priorities are ordered
Each problem carries the points you would recover by fixing it and an indicative effort (low, medium or high). The report's executive summary picks the three with the best return: recoverable points, ×1.5 if critical, divided by the effort (×1; ×1.6; ×2.6). The “estimated score” adds those three fixes to your current score.
What the scanner does not measure
- Whether ChatGPT, Perplexity or Gemini cite you today, or for which questions.
- Your topical authority and your mentions in the sources the models consult.
- The quality or accuracy of your content: only its format and accessibility.
- Performance (Core Web Vitals) or user experience.
That is what the GEO audit covers.
Frequently asked questions
Why does the scanner weight search bots and training bots differently?
Because blocking them has different consequences. If you block OAI-SearchBot, PerplexityBot or Googlebot, those platforms cannot read your site when someone asks them a question and cannot cite you. If you block GPTBot or CCBot, your content does not enter model training: you lose some presence in what they “know”, but you can still appear in their answers through search. That is why the first group weighs three times as much and blocking it is critical, while blocking the second is a warning.
Does Google-Extended control whether I appear in AI Overviews?
No. According to Google's documentation, Google-Extended only decides whether your content is used to train Gemini and for its answers in the Gemini apps. AI Overviews and AI Mode are part of Search and depend on Googlebot. That is why the scanner evaluates the two separately.
Why can the real-access test give a false positive?
Because we request your page with the user-agent of GPTBot, ClaudeBot and PerplexityBot from our own servers, not from those of OpenAI, Anthropic or Perplexity. If your firewall verifies bots by their IP address, it may block us and let the real ones through. We say so in the report itself: if this finding appears, check it in your CDN dashboard.
Why doesn't the scanner measure whether AI already cites me?
Because that cannot be measured well with a single automated query: answers change with the question, the country, the history and the model. The scanner measures technical barriers, which are deterministic. Measuring real citations takes dozens of repeated questions across several models over time, and that is part of the GEO audit.