Public methodology · rule version geo-readiness-v1.0.0
How the Free AI Visibility Checker calculates a GEO Readiness Score
This page documents the crawl scope, category weights, scoring thresholds, coverage calculation, benchmark rules, and limits behind the MiklosKovacs.io Generative Engine Optimization (GEO) assessment.
What is the score designed to answer?
The score estimates whether observable public website signals create a usable foundation for AI search discovery, interpretation, verification, and citation. It is a readiness diagnostic—not a measurement of visibility inside every model, not a market percentile, and not a prediction of traffic or revenue.
What does the scanner inspect?
A run retrieves the submitted homepage, the standard robots.txt, the standard sitemap.xml, an optional llms.txt for context, and up to six representative internal pages. That means a normal completed scan reviews up to seven HTML pages.
| Stage | What is checked | Maximum scope |
|---|---|---|
| Homepage | HTTP response, title, description, canonical, robots meta, H1/H2 headings, schema types, links, visible text, and selected trust/content signals. | 1 page |
| Crawler controls | Root access rules for Googlebot, OAI-SearchBot, GPTBot, ChatGPT-User, Claude bots, PerplexityBot, and Google-Extended. | robots.txt |
| Discovery | Standard sitemap.xml; if it is a sitemap index, up to three child sitemaps are read. | Up to 80 discovered URLs per sitemap source |
| Representative pages | Service, product, pricing, about, contact, proof, FAQ, guide, and resource URLs are prioritized from the sitemap and homepage links. | Up to 6 pages |
| Optional context | llms.txt availability is recorded but does not add points and is not treated as a Google Search requirement. | 1 file |
What are the eight category weights?
Each category receives a 0–100 subscore. The final score is the weighted sum of those eight subscores, rounded to the nearest whole number.
1. Crawl access and indexability
- Homepage responds successfully: 4 of 9 raw points
- No homepage noindex signal: 2 points
- Googlebot, OAI-SearchBot, and PerplexityBot allowed at root: 1 point each
2. Site discovery and architecture
- Usable standard sitemap: 4 of 10 raw points
- One point for each representative internal page found, up to 6
3. Entity and offer clarity
- Homepage title: 2 of 10 raw points
- Homepage H1: 3 points
- Contact-language pages: up to 3 points
- Organization or LocalBusiness schema: 2 points
4. Answer-ready content
- Homepage H2 headings: up to 4 of 10 raw points
- Pages containing FAQ or direct-answer language: up to 3 points
- Reviewed text volume: 2 points from 900 words or 3 from 1,800 words
5. Evidence and citation quality
- Pages containing proof, client, result, case-study, or comparable language: up to 5 of 10 raw points
- Named external links: up to 3 points
- Review schema: 2 points
6. Author and ownership trust
- Person schema: 4 of 10 raw points
- Article or BlogPosting schema: 2 points
- Contact-language pages: up to 4 points
7. Structured data and consistency
- Unique schema types across readable pages: up to 6 of 10 raw points
- Organization schema: 2 points
- Service or Product schema: 2 points
8. Experience, freshness and measurement readiness
- Readable pages with canonical URLs: up to 4 of 10 raw points
- Readable pages with meta descriptions: up to 3 points
- Audit coverage of at least 70%: 2 points; at least 90%: 3 points
How are the final score and coverage calculated?
coverage = round(100 × readable attempted pages ÷ all attempted pages)
| Score range | Readiness label |
|---|---|
| 80–100 | Strong GEO foundation |
| 60–79 | Building GEO readiness |
| 40–59 | Early GEO foundation |
| 0–39 | Limited GEO readiness |
If coverage is below 80%, the label is marked provisional. The two lowest category scores become the priority signals. The recommendation attached to the weakest category becomes the quick win.
What are the important limitations?
- Server HTML, not a full rendered browser: content injected only after client-side JavaScript may not be present in the scan.
- Representative sample: up to seven pages cannot prove the condition of every URL on a large website.
- Standard discovery locations: a sitemap stored elsewhere may be missed unless linked from the pages the scanner reads.
- Root crawler interpretation: the current robots test focuses on whether root access is blocked; it is not a complete standards-compliant simulation of every URL pattern.
- Language signals are heuristics: headings, proof language, contact language, and text volume indicate structure but cannot prove editorial quality or factual accuracy.
- No live answer-engine rank check: the tool does not claim access to private ChatGPT, Gemini, Claude, Perplexity, or Google ranking systems.
- No outcome guarantee: readiness does not guarantee a mention, citation, recommendation, click, lead, sale, or revenue result.
How is benchmark data produced?
The annual benchmark uses the authenticated batch-scan endpoint and the same deterministic rule version as the public tool. Each listed public website is scanned once during the declared collection window. Only aggregate statistics are published; company names, emails, and individual scores are excluded.
- Sampling frame: the public benchmark page states the markets, business types, sample source, sample size, and whether the sample is random or convenience-based.
- Aggregation: median readiness score plus the proportion of sites meeting specific observed conditions.
- Missing checks: unavailable or incomplete scans remain visible through coverage and are not silently converted into confirmed failures.
- Version control: benchmark years with different scoring rules are not compared without a documented restatement.
How does optional email delivery work?
Company name and website URL are sufficient to display the score. If no email is entered, no assessment lead record is created for email delivery and no newsletter subscription is attempted; standard hosting security logs may still record the web request. If an email is supplied, it is used to save and send the requested assessment. Newsletter consent is a separate, optional checkbox.
