Skip to main content
ENGINE v2.0

How the GEO scanner scores

The scanner distributes 100 points across five areas. We publish every check, what it is worth and how it is measured, so you can verify the result and reproduce it. If we change a weight, the engine version changes, and it appears on every report.

Summary: 100 points across five areas

Grade: A from 85 (well prepared for the ai era); B from 65 (solid base with improvements pending); C from 40 (limited visibility in ai engines); D below that (practically invisible to ai).

Access for AI engines — 30 points

If an engine cannot get into your website, nothing else matters. This category measures whether crawlers are allowed in and whether, in practice, your server lets them through.

Readable robots.txt

2 pts

We request /robots.txt at the root of the domain. It counts if it answers 200 with plain text (not with the HTML home page, as many single-page apps do).

Permissions for 15 crawlers

14 pts

We evaluate each user-agent with Google's semantics: the most specific group wins, and within it the longest rule. Bots are weighted by function: search and answer bots (Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot) ×3; bots that visit on the user's request (ChatGPT-User, Claude-User, Perplexity-User) ×2; training bots (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot) ×1. Blocking a search bot is critical; blocking only training bots is a warning, because it is a legitimate decision.

Real access test

6 pts

We request your page three more times, identifying ourselves as GPTBot, ClaudeBot and PerplexityBot (with our signature at the end of the user-agent). If the server answers 4xx, serves a CDN challenge page or returns less than 30% of the text it gave us, we flag it. If the test cannot be completed, we give half the points and say so. If your server rejects even the scanner, we retry once presenting ourselves as a browser so we can analyse the site, and we flag it as a warning.

llms.txt

4 pts

Half the points for publishing it as plain text; the other half if it follows the proposal: it starts with a Markdown title and has at least 3 links.

XML sitemap

4 pts

The one declared in robots.txt or, if there is none, /sitemap.xml and /sitemap_index.xml. It counts if it contains <urlset> or <sitemapindex>.

The 15 crawlers evaluated
User-agentPlatformTypeWeight
GooglebotGoogle Search, AI Overviews and AI ModeSearch and answers×3
BingbotBing and Microsoft CopilotSearch and answers×3
OAI-SearchBotChatGPT SearchSearch and answers×3
Claude-SearchBotClaude (search)Search and answers×3
PerplexityBotPerplexitySearch and answers×3
DuckAssistBotDuckDuckGo (DuckAssist)Search and answers×3
ChatGPT-UserChatGPT (visits on the user's request)On-demand visits×2
Claude-UserClaude (visits on the user's request)On-demand visits×2
Perplexity-UserPerplexity (visits on the user's request)On-demand visits×2
GPTBotOpenAI (model training)Training×1
ClaudeBotAnthropic (model training)Training×1
Google-ExtendedGemini (training and Gemini answers; does not affect AI Overviews)Training×1
Applebot-ExtendedApple Intelligence (training)Training×1
Meta-ExternalAgentMeta AI (training)Training×1
CCBotCommon Crawl (datasets used to train many models)Training×1

Structured data — 20 points

Structured data (schema.org in JSON-LD) tells search engines and models who you are and what each page is, without having to deduce it from the text.

Valid JSON-LD

6 pts

At least one application/ld+json block that is valid JSON. A malformed block is critical: engines discard it entirely.

Brand entity

4 pts

A type that identifies who is behind the site: Organization, LocalBusiness, ProfessionalService, Person and similar.

Complete entity profile

4 pts

One point for each piece of entity data: name, url (or @id), logo (or image) and at least one sameAs pointing to external profiles.

Content schema

6 pts

A type that describes the page: WebSite, WebPage, Article, BlogPosting, FAQPage, Service, Product, BreadcrumbList, HowTo, among others.

Metadata and indexing — 20 points

Metadata and indexing decide whether the page can appear and how it is presented. A noindex or a nosnippet overrides everything else.

<title>

3 pts

Between 10 and 70 characters: full points. Out of range: 1 point. No title: critical.

Meta description

3 pts

Between 50 and 170 characters: full points. Out of range: 1 point.

Canonical

2 pts

Declared and pointing to the same site. A canonical to another domain hands over authorship and is flagged.

Open Graph

2 pts

og:title, og:description and og:image present.

Language

1 pts

lang attribute on the <html> tag.

A single H1

2 pts

Exactly one H1 outside scripts and templates.

Indexable

4 pts

No noindex either in the robots/googlebot meta or in the X-Robots-Tag HTTP header (general or aimed at Googlebot or an AI bot).

Can be quoted

3 pts

No nosnippet or max-snippet:0 (critical: you can be found but not quoted). A max-snippet below 50 characters leaves 1 point; data-nosnippet blocks, 2.

Citable content — 20 points

Most AI crawlers do not run JavaScript and quote specific fragments. This category measures whether your content is in the HTML and in an easy-to-extract format.

Text without JavaScript

8 pts

Visible words in the initial HTML, outside scripts, styles and <head>. 300 or more: full points; between 150 and 299: half; fewer: critical.

Subheadings

3 pts

At least two H2/H3.

Question subheadings

2 pts

At least one H2/H3/H4 phrased as a question (“how”, “what”, “why”, “how much”…): the format that best fits how people ask an assistant.

Lists or tables

2 pts

At least one list (2+ items) or table (2+ rows) outside menus, header and footer.

Trust signals (E-E-A-T)

3 pts

On regular pages: link to “about us” (2 points) and to contact (1). On articles (Article/BlogPosting or og:type article): a declared author, a date within the last 18 months and a link to “about us” or contact, one point each.

HTTPS

2 pts

The final URL, after redirects, is served over https.

Rest of the site — 10 points

A flawless home page is no use if the rest of the site cannot be read. We check a sample to see whether the problems repeat.

Sample of up to 4 pages

10 pts

We choose URLs from the sitemap and from the page's links, one per section (/services, /blog, /contact…). On each we check 5 things: that it responds, that it is indexable, that it has a title, JSON-LD and at least 150 words without JavaScript. The score is the share of checks passed. If there is no way to discover other pages, we give half and flag it.

How priorities are ordered

Each problem carries the points you would recover by fixing it and an indicative effort (low, medium or high). The report's executive summary picks the three with the best return: recoverable points, ×1.5 if critical, divided by the effort (×1; ×1.6; ×2.6). The “estimated score” adds those three fixes to your current score.

What the scanner does not measure

  • Whether ChatGPT, Perplexity or Gemini cite you today, or for which questions.
  • Your topical authority and your mentions in the sources the models consult.
  • The quality or accuracy of your content: only its format and accessibility.
  • Performance (Core Web Vitals) or user experience.

That is what the GEO audit covers.

Frequently asked questions

Why does the scanner weight search bots and training bots differently?

Because blocking them has different consequences. If you block OAI-SearchBot, PerplexityBot or Googlebot, those platforms cannot read your site when someone asks them a question and cannot cite you. If you block GPTBot or CCBot, your content does not enter model training: you lose some presence in what they “know”, but you can still appear in their answers through search. That is why the first group weighs three times as much and blocking it is critical, while blocking the second is a warning.

Does Google-Extended control whether I appear in AI Overviews?

No. According to Google's documentation, Google-Extended only decides whether your content is used to train Gemini and for its answers in the Gemini apps. AI Overviews and AI Mode are part of Search and depend on Googlebot. That is why the scanner evaluates the two separately.

Why can the real-access test give a false positive?

Because we request your page with the user-agent of GPTBot, ClaudeBot and PerplexityBot from our own servers, not from those of OpenAI, Anthropic or Perplexity. If your firewall verifies bots by their IP address, it may block us and let the real ones through. We say so in the report itself: if this finding appears, check it in your CDN dashboard.

Why doesn't the scanner measure whether AI already cites me?

Because that cannot be measured well with a single automated query: answers change with the question, the country, the history and the model. The scanner measures technical barriers, which are deterministic. Measuring real citations takes dozens of repeated questions across several models over time, and that is part of the GEO audit.