How Kuroma scores AI Readiness
AI Readiness is a score out of 100 for your own website. Enter a domain on the free AI Readiness audit and you have the number in about a minute, with no account and no card.
The score measures one thing: how easily an AI engine can reach your pages, take a clean answer out of them, and quote it. That matters because of where your buyers now ask their questions. An AI answer names a handful of sources and skips everything else. When it names your competitor and not you, the reason is usually duller than it feels: the engine could pull an answer out of their page and could not pull one out of yours. Nothing tells you it happened. This score points at the parts of your own site that are in the way.
One number, 19 factors, 4 categories. A factor is one thing we look at on the page, and every one of them is published below with the weight it carries, our verdict on the evidence, and links to the research we based it on.
What the letter beside your score means
That letter is your grade. It is decided entirely by the score out of 100, with nothing added or taken away, and every audit uses the same ranges.
| Grade | Score | What it means |
|---|---|---|
| A+ | 73 to 100 | Excellent! Your page is highly optimized for AI search visibility. |
| A | 58 to 72 | Great job! Your page is well-optimized with minor improvements possible. |
| B | 50 to 57 | Good foundation. Implementing recommendations could significantly boost visibility. |
| C | 43 to 49 | Room for improvement. Focus on high-priority recommendations. |
| D | 35 to 42 | Needs work. Several areas require attention for better AI visibility. |
| F | 0 to 34 | Significant improvements needed. Start with critical recommendations. |
How should you read the score?
The score measures citation readiness, not predicted traffic. AI answers are probabilistic: identical questions produce different citation sets run to run, so no single number can promise placement. Readiness stacks the odds; Kuroma's visibility scans then measure the outcome across engines, and the two together tell you what to fix and whether it worked.
Two ways to read this page. What you have read so far is the short version, and it is enough to act on. Everything below is the evidence behind it: what each factor is worth, and how well the score predicts real AI citations. Every section down there opens with one plain sentence, so you can stop at any depth. Nothing here is held back for a sales call.
Why you can check any of this. This page renders from the same definition the scoring engine runs, so what you read here is exactly what we ship. That is worth something to you. When another tool tells you your score fell from 71 to 64, you cannot tell whether your site changed or their model did. Here you can, and so can anyone you ask to check our work.
Rubric version 2026-08-13. Every earlier version is listed under what changed and when.
Validated against observed AI answers
In plain terms: we checked the score against reality. We took the scores brands already had, then looked at whether AI engines actually quoted them in the weeks that followed. The maths only ever saw the earlier weeks, so the later ones were a real test. Read the table on a scale where 0.50 means the score ranks brands no better than a coin toss and 1.00 means it ranks them perfectly. None of those numbers is a site score. One caveat about the table itself: these figures were measured on the version of our scoring we were running before 9 August 2026. On that date we corrected faults in eleven of the nineteen factors, so the scores either side of it are not comparable, and this test has not yet been re-run on the corrected version. It cannot be, quickly: the test compares what we predicted against what the engines did in the weeks that followed, so it needs several weeks of results under the corrected scoring before there is anything honest to measure. Those weeks are accruing now, and we will publish the new figures whichever way they move.
We regressed observed citation outcomes from our own visibility scans (201,695 AI answers collected over 22 weeks) on the audit factor scores each brand held BEFORE those answers were generated. Validation is out-of-sample on a strict time split: the model is fit on earlier weeks and scored only on later weeks it never saw. The metric is AUC, the probability the model ranks a cited case above an uncited one; 0.5 means no better than chance, 1.0 means perfect ranking.
| Engine | Held-out AUC (later weeks, never seen in training) |
|---|---|
| ChatGPT | 0.84 |
| Grok | 0.76 |
| Claude | 0.68 |
| Google AI Overviews | 0.63 |
| Perplexity | 0.61 |
| Google AI Mode | 0.61 |
| Engine | Held-out AUC (later weeks, never seen in training) |
|---|---|
| Claude | 0.83 |
| Google AI Overviews | 0.52 |
| Perplexity | 0.44 |
| ChatGPT | 0.38 |
| Grok | 0.33 |
| Google AI Mode | 0.30 |
Corpus: 201,695 AI answers, 869,783 extracted citations, 22 weeks, validated 2026-07-07. Fitted cohort: 17 brands.
Gemini is excluded from this table for an honest reason: across our whole corpus it almost never cites a brand’s own domain (it cites retailers, press, and review sites instead), so "was the brand’s domain cited" is the wrong outcome to grade it on. We are building a third-party-coverage label for it.
What this test does not prove.
- Read the two tables together, not separately. The same script over the same period produced both, and on the earlier cohort four of the six engines score below 0.5, which is worse than guessing. We publish that because a number you only show when it flatters you is not evidence.
- The held-out tests are five rows and four rows. That is small enough that neither table settles anything: an AUC from five observations is compatible with a very good model and with no model at all. Treat these as the first look, not as a validated result.
- AUC measures ranking power of the factor set as a whole. It does not license causal claims about any single factor, and we do not reweight the rubric from these coefficients.
- The validated audits come from the March to June 2026 engine, whose measurement core the current rubric inherits with evidence-based reweighting. The current rubric accrues its own validation cohort every scan week.
- The split validates prediction across time for brands the model has seen, not prediction for brand types it has never seen. The cohorts are 17 and 10 brands; we will publish updated numbers as they grow.
The 19 factors. Each category below carries a share of the overall score, and each factor inside it carries a share of its category. In these tables, Status is our verdict on the evidence for a factor. It is not a verdict on your site. One category is also called AI Readiness, the same name as the whole score. That category is one part of the score, not the score: it asks how quotable the page is once an engine has it, and its own section below lists what goes into it. The whole score is all 4 categories together.
Every factor below carries a verdict and the research behind it: controlled studies where they exist, large-scale observational data where they do not, and official platform guidance where vendors have spoken. When the evidence says a popular tactic does not work, we down-weight it and say so.
On-Page Content: 25% of the overall score
Whether the page goes into enough depth, answers what people actually ask, and is shaped so a machine can lift an answer out of it: tables and lists, plain writing, real headings, and words rather than facts locked inside an image.
| Factor | Weight in category | Status | Why (with evidence) |
|---|---|---|---|
| Content Depth & Comprehensiveness | 26% | Evidence-backed | Deep, comprehensive pages are far more likely to be cited than shallow ones, and missing information is a leading diagnosed cause of zero-citation failures. Coverage of the topic matters, not raw length. Cross-engine citation-driver study, deep vs shallow content and freshness gatekeeping (arXiv 2605.25517) · Zero-citation failure taxonomy: contextual gap, intent divergence, information scarcity (arXiv 2603.09296) · Google Search Central guidance on AI features (no special schema needed, no ideal page length) |
| Question Coverage & FAQ Presence | 26% | Evidence-backed | Covering the sub-questions users actually ask is the single strongest diagnosed lever: contextual gaps and intent divergence explain most zero-citation failures. FAQ formatting by itself is not the lever; answering the questions is. Zero-citation failure taxonomy: contextual gap, intent divergence, information scarcity (arXiv 2603.09296) · Google Search Central guidance on AI features (no special schema needed, no ideal page length) |
| Structured Data Quality | 18% | Rescoped to the evidence | Rescoped to extractable data structures in the HTML itself (tables and lists), which measurably improve machine extraction. This is distinct from schema.org markup, which shows no AI-citation lift in controlled tests. GEO: Generative Engine Optimization (arXiv 2311.09735) · Ahrefs controlled schema test (1,885 treated pages vs matched controls) |
| AI Readability Score | 12% | Evidence-backed | Engines preferentially select easier, clearer text, and fluency rewrites improve citation rates in controlled sandboxes. A moderate, causal signal. GEO: Generative Engine Optimization (arXiv 2311.09735) |
| Heading Hierarchy | 11% | Evidence-backed | Heading-scoped sections match how retrieval systems chunk pages, so clean hierarchy helps at the retrieval and chunking stage even though it is neutral at the answer-generation stage. Anthropic contextual retrieval (self-contained passages cut retrieval failures) · GEO: Generative Engine Optimization (arXiv 2311.09735) |
| Media Accessibility | 7% | Rescoped to the evidence | No study shows alt text lifts AI citations, and major AI fetchers request no images. Kept at reduced weight for one honest purpose: a text alternative must exist for content that is locked inside media. Vercel and MERJ AI crawler study (569M GPTBot requests: zero JavaScript execution, no image fetches) |
Technical Optimization: 20% of the overall score
Almost all of this is one question: can an AI engine fetch the page at all, or is it blocked, broken, or too slow to finish.
| Factor | Weight in category | Status | Why (with evidence) |
|---|---|---|---|
| Crawlability & Accessibility | 65% | Evidence-backed | The strongest technical factor: blocking a visibility crawler verifiably removes a site from that engine's answers, and OpenAI documents this directly. Scoring uses a bot-class matrix so only visibility-bot blocks count against the score; training-bot blocks are a policy choice reported neutrally. OpenAI crawler documentation (blocking OAI-SearchBot removes a site from ChatGPT search answers) · Vercel and MERJ AI crawler study (569M GPTBot requests: zero JavaScript execution, no image fetches) · Bing webmaster documentation: NOARCHIVE prevents content from being used in Copilot responses and grounding |
| Mobile Optimization | 20% | Deliberately down-weighted | No AI-engine evidence exists for mobile signals; this is classic search hygiene (Google page experience) retained at a modest weight. Google Search Central guidance on AI features (no special schema needed, no ideal page length) |
| Page Performance Indicators | 10% | Deliberately down-weighted | Speed correlations with AI visibility are weak and driven by extreme outliers; no platform names speed as a factor, so this stays a low weight and a coarse ladder. A fetch that fails scores zero. Otherwise the page is graded on what a crawler has to do to read it: how many scripts and stylesheets block the page appearing, and how much markup it spends per word of content. Cross-engine citation-driver study, deep vs shallow content and freshness gatekeeping (arXiv 2605.25517) · Vercel and MERJ AI crawler study (569M GPTBot requests: zero JavaScript execution, no image fetches) |
| Schema Markup Completeness | 5% | Near-zero weight for AI (kept for classic search) | Controlled tests show no AI-citation lift from schema.org markup, and Google states no special structured data is needed for AI features. Near-zero weight is kept only for classic rich results and Bing grounding. Ahrefs controlled schema test (1,885 treated pages vs matched controls) · Google Search Central guidance on AI features (no special schema needed, no ideal page length) |
Authority Signals: 25% of the overall score
What the page itself shows about who stands behind it: links out to real sources, the contact details and policies that make a site look legitimate, and a named author with a date.
| Factor | Weight in category | Status | Why (with evidence) |
|---|---|---|---|
| External Citations & References | 36% | Deliberately down-weighted | Citing sources helps in controlled sandboxes but the effect shrinks toward zero in competitive replications, so claims are capped and the weight stays modest within the category. GEO: Generative Engine Optimization (arXiv 2311.09735) |
| Trust & Brand Signals | 36% | Deliberately down-weighted | On-page trust badges are untested as a class; the evidenced trust channel is off-site reviews. Retained for publisher accountability (who publishes this, declared where a machine can read it) and for basic legitimacy signals such as contact routes and published policies. Cross-engine citation-driver study, deep vs shallow content and freshness gatekeeping (arXiv 2605.25517) |
| E-E-A-T Signals | 28% | Rescoped to the evidence | No study tests author bios as a causal AI-citation lever; the real authority channel is off-site brand footprint. On-page bylines and dates are kept as best-practice hygiene at reduced weight. Cross-engine citation-driver study, deep vs shallow content and freshness gatekeeping (arXiv 2605.25517) · Google Search Central guidance on AI features (no special schema needed, no ideal page length) |
AI Readiness: 30% of the overall score
How quotable the page is once an engine has it: facts and figures a reader can check, passages that stand on their own, a recent date, something the other results do not already say, code a machine can read, and text that is there before any JavaScript runs.
| Factor | Weight in category | Status | Why (with evidence) |
|---|---|---|---|
| Citation-Worthiness | 25% | Evidence-backed | Verifiable evidence units (statistics, quotes, sourced claims) causally improve citation in sandboxes, and confident wording is preferred over hedged wording. Effects shrink in competitive settings, so claims stay capped. GEO: Generative Engine Optimization (arXiv 2311.09735) · Cross-engine citation-driver study, deep vs shallow content and freshness gatekeeping (arXiv 2605.25517) |
| Answer Box Potential | 24% | Evidence-backed | The real mechanism is passage self-containment: heading-scoped sections that stand alone match retrieval chunking, and fixing context-severed passages measurably cuts retrieval failures. Anthropic contextual retrieval (self-contained passages cut retrieval failures) · GEO: Generative Engine Optimization (arXiv 2311.09735) |
| Content Freshness | 19% | Evidence-backed | Freshness is the only unanimous causal gatekeeper across engines in the 2026 evidence base, and AI-cited content skews measurably fresher at web scale. Google AI Overviews is the documented exception. Cross-engine citation-driver study, deep vs shallow content and freshness gatekeeping (arXiv 2605.25517) |
| Unique Value Proposition | 17% | Deliberately down-weighted | Unique-wording signals test near null; the real phenomenon is competitive redundancy, which a single-page crawl cannot measure deterministically. Weight halved accordingly. GEO: Generative Engine Optimization (arXiv 2311.09735) |
| AI Discoverability | 10% | Rescoped to the evidence | Rescoped to on-page machine-readability signals (semantic HTML, licensing clarity). The dead ai-plugin manifest check was removed and llms.txt stays informational only, because no major engine reads it. Vercel and MERJ AI crawler study (569M GPTBot requests: zero JavaScript execution, no image fetches) · Google Search Central guidance on AI features (no special schema needed, no ideal page length) |
| Raw HTML Availability | 5% | Evidence-backed | No major AI crawler executes JavaScript, so content that only appears after scripts run is invisible to AI search. This is among the best-evidenced technical findings in the audit. Vercel and MERJ AI crawler study (569M GPTBot requests: zero JavaScript execution, no image fetches) |
What we deliberately do not score
Many AI-visibility audits score these. The best available evidence says they do not move AI citations, so scoring them would inflate work that does not pay.
| Not scored | Why not (with evidence) |
|---|---|
| llms.txt | Triple-null: the vast majority of llms.txt files receive zero crawler requests, Google states it ignores them, and no major engine reads them for citation. Detected and reported as informational only. Google Search Central guidance on AI features (no special schema needed, no ideal page length) · Vercel and MERJ AI crawler study (569M GPTBot requests: zero JavaScript execution, no image fetches) |
| Schema markup as an AI-citation lever | A large controlled test found no AI-citation lift from adding schema.org markup, and Google states no special structured data is needed for AI features. Schema is scored near zero, for classic rich results and Bing grounding only. Ahrefs controlled schema test (1,885 treated pages vs matched controls) · Google Search Central guidance on AI features (no special schema needed, no ideal page length) |
| Per-run AI rank tracking | Single-run AI answers are statistically unstable (same-day identical queries agree on citations only a few percent of the time), so a per-run rank is noise. Kuroma surfaces mention rates over repeated runs instead. Cross-engine citation-driver study, deep vs shallow content and freshness gatekeeping (arXiv 2605.25517) |
| Keyword density optimization | Keyword stuffing tests causally negative for generative engine citation and has been dead in classic search for years. Nothing in this audit rewards it. GEO: Generative Engine Optimization (arXiv 2311.09735) |
Every version of this rubric, and what changed
A score is only worth tracking over time if you know when the ruler changed. Every version of this rubric is below, newest first, with what changed, why, and whether it moved anyone’s score. Where a version moved scores, two scores taken either side of it are not two measurements of the same thing, and we say so rather than leave you to assume otherwise.
This history was written on 2026-08-07 from our change records and our internal specification, rather than kept as each version shipped. The oldest versions here predate the shared definition this page renders from, so each of those says what it rests on.
Rubric version 2026-08-13, the one in use now
What changed. Answer Box Potential stopped giving you points for having a list, and started measuring the thing it is named after. It used to award a fifth of its score for "an ordered list, or two unordered lists, or three list items anywhere on the page". What it now looks for is passages: a heading that says what a span of your page is about, followed by writing that answers it and stops. Long enough to be an answer, short enough to be quoted whole, and made of writing rather than links. It asks that twice, because they are two different problems: whether a page of your length carries enough such passages, and whether your headings mostly open one or mostly label a link. The three other things it looked at are still there. Definitions are graded on a curve now rather than jumping at one and at three. A question-and-answer block is worth half what it was, because the Question Coverage factor already scores the same FAQ markup and the same question headings, and one fact should not be paid for twice at full price. Tables, definition lists and step-by-step markup are worth half what they were, because 89 percent of homepages do not have any, so at the old price they were a standing deduction rather than a description. Five more factors were rebuilt in the same batch, all for one reason: they had stopped being able to tell two sites apart. Trust and Brand Signals no longer awards points for having an HTTPS certificate, and no longer spends half its range asking whether you are a small shop. Star ratings, trust seals, testimonials, a postal address and a phone number were 52 of its 100 points. What it asks instead is whether a reader can tell who publishes this page: an about page, a legal or imprint notice, and a publisher declared in your markup with a name, a logo, a website and the accounts you post from. Ratings, seals and testimonials still count, at 15 points between them rather than 52. Structured Data Quality stopped rounding your lists and tables into three buckets. Two lists and a hundred lists used to be the same page. Lists also now carry more of the score than tables do, because nine pages in ten have no data table at all, so two fifths of this factor was unreachable for almost everybody. Question Coverage can now read a question written in Chinese, Japanese or Korean. A third of its score used to come from English phrase patterns, so a page answering questions in any other language scored nothing for it. It now counts sentences that end in a question mark, full-width or ASCII, which every language writes. AI Discoverability now looks at how much of your document sits inside a landmark, not only how many kinds of landmark you use, and pays separately for declaring a main region. Raw HTML Availability grades the text a JavaScript-free crawler receives on a smooth scale instead of four steps, so 299 words and 300 words are no longer opposite verdicts.
Why. We measured what the factor actually did across 200 real pages, and a third of them scored exactly the same number: 20. The reason was the list tier. The typical homepage carries 13 unordered lists and 86 list items, and almost all of them are the navigation menu, the footer and the card grid. So the largest single behaviour of a factor named Answer Box Potential was a bonus for having a menu, and 88.6 percent of homepages scored under 50 no matter what they had written. A factor where a third of the web lands on one number cannot tell a good page from a bad one, which is the whole job. The published mechanism behind this factor has always been passage self-containment, and nothing in it was measuring that. The list tier was removed rather than repaired. A list helps when its items say something, and a list whose items say something is already counted inside the section it sits in. Measured, list items carrying eight or more words of text that is not a link exist on 11 percent of homepages, so a repaired list tier would have been a fourth tier that almost nobody could reach, which is the defect rather than the fix. The other five were found the same way and the numbers were as stark. Question Coverage gave 110 of 200 pages the identical score of zero. Structured Data Quality had eleven possible answers in total and gave 42 percent of pages the same one. AI Discoverability had eleven, with a third of pages on a single number that means "a site template". Raw HTML Availability had four. Trust and Brand Signals failed differently and worse. Its floor was 15 points, because every page we audit is served over HTTPS and every page collected the same 15 for it, so no site could score zero however little it published about itself. Its ceiling was 50, because the best any page could do without being a shop with customer reviews was half the factor. We check that against a set of pages nobody disputes are excellent, and not one of them could be called good on a factor about trust. Crawlability had already stopped scoring HTTPS in August for the same reason; this finishes that. None of this is a matter of taste. A factor where a third of the web lands on one number cannot predict anything, and we are about to spend three months testing whether these scores predict whether AI assistants cite you.
Did scores move? Yes, and scores from before this date cannot be compared with scores from after it. On the 200 frozen pages, 181 move: 128 up and 53 down, by 9.1 points on this factor on average, which is about 0.66 of a point on the overall score. The direction of the big moves is the point. A news homepage carrying 142 headings and not one section you could quote falls from 63 to 33. cloudflare.com, which scored zero, rises to 48 on nine quotable sections it was never credited for. Across the world-class homepages we hold as a reference set the average rises from 24 to 51, and the number of them scoring under 50 falls from 10 of 11 to 5 of 11. The bottom did not move much, and that is deliberate. Of the pages we label as weak, 89 percent still score under 50 and 57 percent still score zero. A page with no heading, or with nothing but links under the headings it has, is not an answer box and is not scored as one. The other five moved too, and the same rule applies: do not compare a score from before this date with one from after it. On the 200 frozen pages, Trust and Brand Signals moves on 197 of them, by 15 points on average, and the reference set of excellent pages rises from an average of 29 to 44 with its best page going from 50 to 81. Question Coverage moves 125 pages, nearly all upward, and the pages scoring a flat zero fall from 110 to 65. AI Discoverability moves 160 pages by about 5 points. Raw HTML Availability moves 19 pages and leaves 181 exactly where they were. Structured Data Quality is the one that mostly goes DOWN, on 85 pages against 43 up, and that is the correction working. Under the old buckets, one table plus two lists plus one definition list scored 90 of 100 - the same as a reference page carrying eleven tables and forty lists. Pages that were being told they were nearly perfectly structured are now told where they actually sit. A page whose lists are all in its navigation menu still scores zero, which has been true since August and is not a change. Content Freshness changed too, and it moves more pages than any of the six above: 76 of the 200. Two separate repairs. It can now read a date written the Korean way. The pattern accepted the Japanese and Chinese characters for year and month but not the Korean ones, so a Korean page that printed its date in its own script was read as having no date at all. On one Korean news homepage the only thing the old pattern could find was the press-registration date in the masthead, from 2012, so we dated a page published that morning to fourteen years ago and scored it 41 instead of 79. Registration lines are now skipped, because a registration date is not a publication date in any language. The other repair sends more pages to zero, from 45 to 67, and that is the factor working rather than failing. Of those 67, forty-nine do not print a recent year anywhere a reader can see, and on the other eighteen every recent year sits inside the copyright line at the foot of the page. A copyright notice is generated by the template on every request; it says nothing about when the words above it were written. Paying a page thirty points for it meant a site that tells you nothing about its freshness scored the same as one that does. If your page states when it was written or updated, in any of the four languages, it now scores for that; if it does not, it scores zero and the reason is honest. E-E-A-T Signals gained two new ways to earn authority, and that means some pages score lower without having changed. Read this one carefully, because the direction is not obvious. The authority half of this factor used to be worth 10 points and had almost nothing to measure. It now recognises two things it could not see before: a machine-readable link from your organisation to its other official profiles, and a link to a public record of the organisation itself. Both are things a real company can point at and a fabricated one usually cannot, which is what this factor is for. That takes the authority half from 10 points to 25. Because the scale got longer, a page that has neither of those two signals is now measured against a larger total, so its score falls even though the page did not change and nothing about it got worse. It is the same answer measured on a more demanding scale. If your organisation publishes those links, you gain; if it does not, this is the clearest thing on the page to go and fix. One asymmetry we would rather state than have you find. The machine-readable profile link appears on 56 percent of the English pages we measure and 34 percent of the Chinese, Japanese and Korean ones, so more non-English pages take this reduction than English ones. That reflects a real difference in what sites publish today, not a judgement about the sites, and it is a gap that is cheap to close.
Rubric version 2026-08-11
What changed. Readability now finds your paragraphs by reading the page, not by looking for one particular HTML tag. It asks the same question it always asked - what share of your paragraphs are short enough for an assistant to quote whole - but a paragraph is now any block of writing that carries a sentence. Headings do not count, and neither does a block that is mostly a link, because neither is a paragraph. No weight moved: this factor is worth the same 35 of readability's 100 points it was worth yesterday.
Why. Readability divided your short paragraphs by your total paragraphs, and it counted paragraphs by looking for the <p> tag. A site that lays its writing out with <div> instead has no <p> at all, so that sum had nothing to divide by. An undefined answer was being read as zero, which is the worst possible result, and 12.8 percent of all the pages we had scored were being marked down for prose we never actually looked at. One clinic site was scoring 35 on 3,674 words across 324 sentences. The obvious repair - skip that part when we cannot run it and rescale the rest - was built, tested and thrown away, because it would have paid a site 19 points for DELETING its paragraph tags. Removing a measurement always flatters the pages that would have failed it. The only honest fix was to measure the thing the factor is named after.
Did scores move? Yes, and this is why scores from before today cannot be compared with scores from after it. Measured over 396 real pages, readability moves on 27.8 percent of them: 75 up and 35 down, by as much as 35 points. Pages that use <p> can move too, because they often carry writing in other blocks as well and that writing now counts. The effect on the overall score is small - about 0.04 of a point across the fleet - which is why the letter-grade boundaries set yesterday are left exactly where they are.
A fix inside version 2026-08-10, made on 2026-08-10
What changed. The score a letter grade starts at moved. A+ now begins at 73 rather than 75, A at 58 rather than 68, B at 50 rather than 62, C at 43 rather than 55, and D at 35 rather than 45. No score changed, and neither did any of the nineteen factors or their weights.
Why. The two factors rebuilt earlier the same day moved every audited site down by about 11 points. Left alone, the previous letter boundaries would have graded 70 percent of all published sites an F - which is the exact failure the boundaries were set to fix two days earlier, when an inherited school-grade ladder was failing 47 percent of the web and had never awarded a single A. What is held constant across both settings is the SHARE of sites in each grade, not the numbers. The numbers describe a measuring instrument; the shares describe the web. When the instrument changes, the numbers have to move so the description does not.
Did scores move? No. Not one score changed, and scores from before and after this entry are directly comparable - the letter is a label applied to a number that was already there. Many sites DO see a different letter today, but that is the earlier rebuild moving their score, not this. Measured over the 583 sites scored on the current generation, the grades now fall 0.5 percent A+, 3.4 percent A, 11.8 percent B, 23.5 percent C, 28.8 percent D and 31.9 percent F, within a point of the shape the previous setting produced.
Rubric version 2026-08-10
What changed. Two of the nineteen factors were rebuilt, and they shipped together because either one alone would have forced the same fleet-wide re-score. The factor that asks whether a page says something distinctive now reads a page in any language. It used to look for English phrases - "the only", "unlike", "case study", "our methodology" - and it refused to read any passage that was mostly not written in the Latin alphabet, so on a Chinese, Japanese or Korean page there was nothing left for it to grade and it was set aside entirely. It now looks for the things a page can only say if it knows its subject, and those look the same in every language: figures carrying a unit, a price or a percent sign; two such figures placed side by side in one passage, which is what a comparison a reader can check looks like; how many different kinds of figure the page uses; and the versions, models and standards the page names outright. How much of the page carries figures now matters more than how many figures it has, so a long page is no longer credited simply for being long. The factor that asks whether AI crawlers can reach and use a page was rebuilt around what a crawler can actually do with it. It used to award most of its points for things almost every site already has: a secure connection, a doctype, a character-set declaration, a robots.txt existing at all. Blocking an AI search engine was priced as a fixed deduction, so a site that shut out a third of the engines that could cite it lost a few points and usually stayed in the same grade. Permission is now a multiplier over the whole factor: shut out a third of those engines and a third of this score goes, and a page marked noindex scores zero however well built it is, because a page kept out of the indexes AI answers are grounded in cannot be found through them at all. The points that remain are spent on how many different pages on the same site it links onward to, and how many of its links are ones a crawler can follow rather than menu placeholders. A secure connection is no longer scored, because every audited site has one and a point everybody earns cannot tell anybody anything.
Why. Both factors had stopped separating one site from another, and we measure that directly. Across the same 200 real pages, the distinctiveness factor gave 91 percent of Traditional Chinese pages and 93 percent of Korean pages the same score, so it could not tell two sites apart in those markets at all. Setting it aside on those pages did not solve that; it just meant Chinese, Japanese and Korean sites got no feedback on it whatsoever. The crawler factor was worse, because it applied everywhere: it put 96 percent of all pages in one grade band, and every single one of the 30 Japanese pages in the same band. It is the largest single factor in the audit, so it was also the most expensive one that told you nothing. The same fault sat under both. Almost every point went to something virtually every site already had, or to English wording a page in another language could not produce, and a point like that cannot tell two sites apart however carefully it is measured.
Did scores move? Yes, in both directions, and scores from before this date cannot be compared with scores after it. Chinese, Japanese and Korean pages are now graded on the distinctiveness factor where it was previously left out of their score entirely. Every site is regraded on the crawler factor: one that lets the AI engines in and links onward through followable links scores higher than before, and one that blocks engines, marks pages noindex, or gives a crawler no path onward scores markedly lower. Measured across the same 200 pages, the share of pages sharing the single most common crawler score fell from 55 percent to 10 percent and the number of distinct scores rose from 27 to 59; on the distinctiveness factor the share scoring exactly zero fell from 63 percent to 34 percent, and the best score a well-built page could reach rose from 69 to 90 - which was itself a fault, because no site on earth could previously be called good on it.
Rubric version 2026-08-09
What changed. A full audit of the scoring engine found 28 faults, and fixing them changed what twelve of the nineteen factors measure. Two of those were not adjustments but corrections of things that had never worked. First, the text of a page was being read in a way that also swept up the page’s own program code and styling instructions and counted them as writing, at thirteen separate places in the engine. Second, the file a site uses to tell automated visitors which parts it may read was being run through a converter meant for web pages, which ran every line together, so a site blocking an AI search engine looked exactly like a site allowing it. Alongside those: how site access is weighted, how non-English pages are judged, what counts as a citation, what counts as fresh, and what counts as a distinctive page all changed. Page performance changed the most of the twelve. It had been graded on whether the page could be fetched and how quickly it answered, which almost every page passes, so it gave the same result to every site we measured it on in Traditional Chinese and to every site in Japanese. It now grades what a machine actually has to do to read the page: how many scripts and stylesheets have to load before the page appears, and how much markup the page spends per word of content.
Why. Every one of these was measuring something other than what it claimed to measure. The code reading was the clearest: a page with three hundred words of writing and a large program bundle was being credited with tens of thousands of words, and was then told to write tens of thousands more. The access file was the most consequential: across a sample of 47 real sites, 36 percent were blocking an AI search engine in a way we could not see at all, so they were scored as though they welcomed every one.
Did scores move? Yes, substantially, and in both directions. Scores from before and after this date are NOT comparable and must not be read as a change in a site’s performance. Sites carrying a lot of program code generally score lower now, because they are no longer credited for it. Sites that block AI search engines score lower on access, because that block is now visible. Some sites score higher where a measurement had been failing closed against them. On page performance specifically, nearly every site used to score full marks and most now score below that; measured across 200 real pages the effect on the overall score was under one point on average, and it was the same size in Traditional Chinese, Japanese, Korean and English. Our own published figure for how well the score predicts real citations was measured on the previous engine and does not describe this one; it is being re-measured and we will publish the new number whichever way it moves.
A fix inside version 2026-08-04, made on 2026-08-08
What changed. The letter grade attached to a score changed. A+ now starts at 75 instead of 90, A at 68 instead of 80, B at 62 instead of 70, C at 55 instead of 60, and D at 45 instead of 50. The score itself did not change, and neither did any of the nineteen factors or their weights.
Why. The old cutoffs were the school grading ladder and had never been checked against what real pages actually score. Across all 764 score cards published on this site they graded 47 percent of the web an F and had never once awarded an A. Running the same scoring over the home pages of Stripe, Cloudflare, Anthropic, OpenAI, Vercel, Notion and eight others put every one of them near 62, because a perfect 100 would require a single page to be a home page, a question and answer page, a cited research article and a bylined author page all at once. A ladder that calls Stripe a C is describing its own scale rather than the web. The new cutoffs are anchored to the measured distribution, where A+ is about the top 0.2 percent and A about the top 3 percent.
Did scores move? No. Every score is exactly what it was, and two audits either side of this change are directly comparable. Only the letter printed beside the score moved, and it can only have moved upward, because every cutoff was lowered. The published validation of how well the score predicts real citations is unaffected, because that measures how scores RANK against each other and a letter boundary is not part of that ranking. The anchor is fixed rather than recalculated as more sites are audited, so your grade will not change because somebody else ran an audit.
A fix inside version 2026-08-04, made on 2026-08-05
What changed. On a page that declares no canonical address, every outbound link was being counted as a link back to the page itself, so the page scored nothing at all for citing outside sources however many it cited.
Why. The rule that ignores links to your own site compared each link against the page’s canonical address. When a page declares no canonical address, that comparison ran against an empty piece of text, and every address contains an empty piece of text, so every link matched.
Did scores move? Yes, for one kind of page: one that declares no canonical address and does link out. Those pages gained up to 3.15 points of the overall 100 that they had not earned. Pages that declare a canonical address are untouched, and so are pages with no outbound links. The version stamp did NOT move, because the rubric did not change, only a fault in how it was applied. That means an audit does not record which side of this fix it was scored on. If you are comparing two audits of a page with no canonical tag across the start of August, read the difference with that in mind.
Rubric version 2026-08-04
What changed. Pages that are not in English stopped being measured with English machinery. Three things changed together. Words had been counted by splitting text on spaces, which Chinese and Japanese do not use, so a page of Chinese prose could be counted as almost empty and thrown out as a blank shell. Four factors had been switched off wholesale on Chinese and Japanese pages because they look for English words, when in fact much of what they look at is markup that belongs to no language. And the readability factor had been reading the whole page, navigation and footer included, instead of the main content.
Why. Each of those was a measurement fault rather than a change of opinion. A Chinese page was not being scored harshly, it was not being scored at all, and nothing on the page said so. Switching the four factors off was wrong in both directions at once: it over-credited Chinese and Japanese pages that carry no markup and under-credited the ones that carry plenty. The three shipped as one version because each on its own would have forced its own re-score of everything.
Did scores move? Yes, and this is the largest move in this list. Chinese and Japanese pages move by 9 to 16 points depending on how much markup they carry. English pages move slightly, because the readability fix reaches them too, which is why the re-score covered everything and not just the Chinese and Japanese part of it. English scoring is otherwise unchanged, checked factor by factor. Scores from this version and the one before it are not two measurements of the same thing, so every published score was re-run on this rubric before any of them was used in a benchmark again.
Rubric version 2026-08-02
What changed. Market adjustments began working. Kuroma adjusts a score for the market a brand sells into, for example crediting a Taiwanese site for being written in Traditional Chinese. That adjustment had never once been applied to anybody.
Why. The adjustment table is looked up by market code, and every part of the system that consulted it passed a market database identifier instead. The lookup missed silently, so every market scored identically and the feature existed only on paper for its whole life. A review the previous month had confirmed the table held the right 20 markets without ever asking whether anything read it.
Did scores move? Yes, for brands in the 20 markets that have adjustments configured. A brand can gain up to 8 points, or lose 1 where the site language does not match the market it sells into. A typical Taiwanese brand gains about 5. The move is not proportional, because a category already at 100 has nowhere to put its credit. Audits with no market attached are identical either side of this change, which is most of the public benchmark corpus.
Rubric version 2026-08-01
What changed. The Structured Data Quality factor started measuring what this page had already been telling visitors it measured: data structures inside the HTML itself, meaning tables with a header row, lists of three or more items, and definition lists. Until this version it was scoring schema.org markup instead.
Why. The factor was rescoped on the evidence in July and this page was updated with it, and the scoring code was not. For about three and a half weeks the published rubric described one thing and the engine measured another. Worth being precise about what that says: this page renders from the same definition the engine reads its weights from, so a weight here is always the weight the engine used, and that guarantee never covered what an individual factor’s code goes looking for. This is the gap it does not cover, and the reason we publish it. The scoring code was changed to match the published rubric rather than the other way round, because rewriting the page would have quietly reversed an evidence-based decision, and because schema.org is already scored on its own honest near-zero weight elsewhere in the rubric.
Did scores move? Yes, sharply, for one kind of page. A page rich in schema.org markup and empty of tables and lists went from near the top of this factor to near the bottom. The factor is worth 4.5 points of the overall 100. Scores either side of this version cannot be pooled: it would compare sites judged on schema.org against sites judged on tables.
Rubric version 2026-07-08
What changed. Chinese language patterns were added to three factors that until then only recognised English text: external citations, content freshness, and answer box potential.
Why. Those three factors looked for English words and phrases, so a Chinese page could not score on them however well it was written.
Did scores move? Chinese pages only. English scoring is identical, which is why this is the one version in this list whose scores can still be compared with the version before it.
Rubric version 2026-07-07
What changed. The whole rubric was reweighted against the evidence and this page was published for the first time. Six weights moved, listed below. Raw HTML Availability was added as a new factor, taking the count from 18 to 19, because no major AI crawler runs JavaScript and content that only appears after a script has run is invisible to them. Page speed also stopped being graded and became a pass or fail on whether the page can be fetched at all.
Why. The weights had been inherited from classic search practice, and where controlled studies existed they disagreed with those weights. Adding schema.org markup showed no lift in AI citations, so it fell to almost nothing. Blocking an AI crawler verifiably removes a site from that engine’s answers, so crawlability became the dominant technical factor. Author biographies have never been tested as a cause of AI citations, so they were cut back to a basic hygiene weight. This is also the version that made the rubric one shared definition read by both the scoring engine and this page, which is what makes everything above checkable.
Did scores move? Yes, for everyone. Most factor weights moved and a factor was added, so no score from before this point can be compared with one after it.
| Factor | Weight before | Weight after |
|---|---|---|
| Crawlability & Accessibility | 20% | 65% |
| Schema Markup Completeness | 35% | 5% |
| Page Performance Indicators | 25% | 10% |
| E-E-A-T Signals | 40% | 28% |
| Unique Value Proposition | 35% | 17% |
| Media Accessibility | 15% | 7% |
Rubric version 2026-06-28
What changed. Media Accessibility was added to the On-Page category as the 18th factor, entering at 15% of that category, and the five factors already there were rebalanced so the category still added up to 100%. Its weight was cut to 7% nine days later in the reweighting above.
Why. Text locked inside an image or a video is invisible to an AI engine, and nothing in the rubric was asking whether a text equivalent existed. The starting weight was set from the research rather than from our own data, and was described at the time as provisional.
Did scores move? Yes. A new factor entered a category and its neighbours’ weights moved to make room, so scores either side are not comparable.
How we know. Reconstructed from the change itself and from our internal specification. This version predates the shared rubric definition that this page renders from, so there was no published rubric at the time to check it against.
Rubric version 2026-06-06
What changed. The first version stamp, added on the day three scoring faults were fixed. An llms.txt file had been worth 30 points and dropped to 5. A site that blocked crawlers in its robots.txt file was having that counted against it twice, in two different factors. And a site with no robots.txt file at all was being awarded 15 points for its absence.
Why. Those three fixes moved scores, and nothing in a stored audit said which rubric had produced it. The stamp was added so that a later drop could be attributed to the rubric changing rather than to the site getting worse.
Did scores move? Yes. Audits from before this point carry no version stamp at all, so there is no way to tell which rubric produced them; the date they were taken is not evidence of a version. They are not comparable with anything in this list.
How we know. Reconstructed from the change that added the stamp and from our internal specification. The commit that made the three fixes carries no description of its own, and this version predates the shared rubric definition that this page renders from.
Now run it on your own site. The free AI Readiness audit scores any domain against this rubric in about a minute. Then run it on the competitor the AI keeps naming.