The Optimization Leaderboard
We make several changes to a page at once, so which of them is actually associated with that page getting found and cited? To answer it we threw away every untouched page and compared optimized pages against each other, matching on the site, on how many changes the page got, and on what kind of page it is.
On 130 sites out of 164 where the comparison is decided, pages that got an FAQ block earned more AI referral visits than same-page-type pages on the same site that did not. Every change we can test this way comes out ahead more often than not, and the FAQ block leads.
It says how often a change is ahead. It does not say by how much, and we are not publishing a magnitude, because the magnitude moves depending on how you measure it. The FAQ block reads as 2.5 times on pooled visits and the H1 tag as roughly half, while at site level both come out ahead. A few high-traffic pages inside a few sites drive the pooled version, which is exactly why the site is the unit here.
Share of sites where the change came out ahead, against a 50% no effect line. llms.txt is on a different footing: it is measured from crawler logs rather than referrals, because AI crawlers fetched it 31 times across 82 million requests in six months.
A question and answer section generated from the page's own content, rendered visibly and mirrored into FAQPage markup.
The title tag rewritten to carry the specifics a person would actually ask for.
Dense unstructured blocks rewritten to lead with the answer, in place, many times per page.
The page's single main heading rewritten to name the thing the page is about.
The file the field has spent two years recommending. Measured from crawler logs rather than from referrals, because the question is prior to that: across 82 million crawler requests over six months, AI crawlers fetched llms.txt 31 times in total. A file that is not read cannot be moving anything.
Finding 02 · what kind of page it is matters more than anything we do to it
The single biggest thing we learned building this is a warning about reading any table like it. Before page type was matched, heading structure looked like a +1.09 point winner. Adding page type to the match took it to +0.25, inside the noise. The first number was mostly telling us which templates our heading pass fires on.
So every comparison above is made between pages of the same type, on the same site, carrying the same number of changes. Only the composition of the bundle differs.
Two things that do not move it
Worth as much as the ranking. One of these is the single most widely recommended change in this field.
How much of the page there was to fix matters. How much better we made each piece of it does not, at least not measurably. Pages where the answerability score rose the most are not ahead of pages on the same site where it rose the least, once the number of rewritten blocks is matched: 0.14 points against a 2.68% baseline, inside the noise on 37,360 pages.
Not that it does not work, but that it cannot be measured here, and for an interesting reason. Every optimized page is served with an AI-readable HTML version, so there is no optimized page without it to compare against. And the problem it solves is rarer than the field assumes: on 4,499 audits in six months, 76.1% of sites already scored 9 or 10 out of 10 for serving their main content without JavaScript. Only 8 sites in the whole panel scored below 9 AND had enough pages in both arms to compare, which is no sample.
And the one that dwarfs them
5.2xNone of the changes above comes close to the effect of how much the page actually says. Pages carrying sixteen or more substantial content blocks earn an AI referral 5.2 times as often as pages carrying one, and the gradient is monotone all the way up.
Descriptive, and uncontrolled. A page with sixteen substantial content blocks differs from a one-block page in more ways than length, so this is not a claim that adding words earns citations. The evidence here is that the gradient is monotone across all five bands on 42,913 pages, and that it is several times larger than any per-change number on this page.
How often is the content not in the HTML at all?
An assistant that has to run your JavaScript to find your prices usually does not. Every page we optimise is served with an AI-readable HTML version for that reason, which also means there is no optimised page without one, so it cannot appear in the ranking above: it has no comparison group.
What we can report is how common the problem is, and it is less common than this field implies. On 4,499 audits over the last 180 days, 76.1% of sites already serve their main content without JavaScript. About one in nine does not, and 6.1% are severe.
One audit per site, on one page, so this describes homepages more than whole sites.
Changes held out of the ranking
6 of 10Listed rather than dropped, because a leaderboard that quietly omits what it could not measure reads as though it measured everything. The rule was fixed before the estimates were looked at: a change qualifies only if it is applied because of a content defect the change itself fixes, and if it existed for the whole window. A change applied because of what kind of page it is, or because of what markup the page already carries, cannot be pulled apart from that page characteristic by matching. No numbers are given for these, because the number would be a fact about page selection wearing the label of a change.
70.6% of these changes land on a page that already carried Article, NewsArticle or BlogPosting markup, usually from an SEO plugin. Pages that get it are therefore pages an SEO plugin already manages, which matching on page type does not fix. It is also a replacement rather than an addition, so even a clean estimate would answer a narrower question than whether Article schema helps.
Shipped 2026-07-01, so it exists only in the last six weeks of a five and a half month sample and its pages have a different calendar from every other row's. It is on 43.8% of optimizations now and is the first candidate for the randomized readout.
Applied to product pages by definition, so within any sample that contains other page types its association is largely the association of being a product page. After matching, only 11 sites still contribute.
The same problem, on 0.6% of optimizations. After matching, 1,063 pages and 8 referral events remain, which is no power at all.
On every optimization we shipped last month. A change with no untreated comparison pages has nothing to be compared against, at any sample size.
Also on every optimization, and it is the record of the other changes rather than a change to the page. After matching, 180 pages remain.
Method, and two estimates we threw away
The design above is the third one. The first two both compared optimized pages against untouched pages, and both produced large, significant, wrong answers.
Optimized pages against never-optimized pages, pooled
discardedOne site supplied 63,836 of the 68,300 control pages, at a 67.6% baseline crawl rate while most sites sat at 0.0%. This is one site's month.
Same, but within site and within baseline stratum, median across sites
discardedComparing treated pages against untreated pages compares groups that differ in importance, age and page type as well as in treatment. Nothing that keeps an untreated control group in this data survives those differences.
Optimized pages against each other, matched on site, bundle size and page type
standsEvery page is treated, so treatment status cannot confound anything. What is left is the composition of the bundle, matched on the two things that otherwise drive the outcome.
Across 89,231 pages optimized between March and August 2026, the median page was optimized 3.04 days after it was first discovered and 89.3% within four weeks. A pre-period drawn before the optimization therefore falls partly before the page existed, which is what inflated the two discarded estimates. Comparing optimized pages with each other removes the problem, because every page in the sample sits the same distance from its own optimization date.
At page level AI crawls are almost all zeros: of 2,343 active pages in a 28-day window, 8.9% received any AI crawl and the 90th percentile is 0. A referral visit is both further down the funnel and better behaved as an outcome: 2,088 of 75,496 pages received one. Crawls are a secondary outcome for the randomized study.
4 of 8 prechecks failed. Each failure changed the design rather than being written off. The earlier pooled estimate of +45.1pp is what the same data yields if you skip them.
Turning the top row into an effect
allocation lockedMatching gets an association. Only randomizing which pages get a change turns it into a causal effect, and that is the next study.
Randomize at the page level within a site, among pages that have all existed for at least eight weeks, so page age cannot enter the comparison at all.
Withhold one change type at random from half of the treated pages. That, and only that, turns a row in this table from an association into an effect.
Start with the FAQ block, since it is the row this table detects, and with the quick answer block, since it is the newest change and has never been measured.
Pre-register the window, the binary outcome and the site-level sign test before looking, since the outcome is sparse and the researcher degrees of freedom are large.
The order matters more than any single step: the allocation and the estimator were both finished and fixed before a single page was treated, and the estimator refuses to run against an allocation whose digest does not match. Nothing about the analysis can be chosen after the result is visible.
This is an association, not a randomized effect. We do not choose at random which changes a page gets, so a row describes the pages that received a change rather than what the change does to a page.
The sample is paying customers, overwhelmingly small and medium business sites on WordPress, so every figure describes those sites and not the web at large.
The outcome is an AI referral visit, meaning a person clicked through from an assistant. It is the end of the funnel rather than a crawl, which is deliberate, but it means a page can be read and cited without ever showing up here.
Referrals are attributed from tracking parameters and referrers. Assistants that strip both are invisible to it, and chatgpt.com is 84% of what remains, so the ranking is weighted toward how one assistant behaves.
trigger_source is NULL for every completed optimization before 2026-08-11, so we cannot check how many optimizations were themselves triggered by a content change, which is independently a reason a page might get read.
34.8% of pages are optimized more than once. Only the first optimization defines the bundle and the window.
That an FAQ block causes AI referrals. It is associated with them, which is the strongest claim an observational design supports. Only the randomized study can promote that to a cause.
That the three rows with no detected difference do nothing. Their intervals include zero, which means this design cannot separate them from zero, not that they are zero.
That the excluded changes do not work. They are excluded because the comparison would measure page selection, not because it came out badly.