Skip to main content
LovedByAIThe Lab
Live
11 Sep 19:01 UTC
ChatGPT-User · 0sPerplexityBot · 31sOAI-SearchBot · 3mClaude-User · 40sDuckAssistBot · 1mGrokBot · 2m
Study 001Optimization effects

The Optimization Leaderboard

We make several changes to a page at once, so which of them is actually associated with that page getting found and cited? To answer it we threw away every untouched page and compared optimized pages against each other, matching on the site, on how many changes the page got, and on what kind of page it is.

DesignMatched, within site
OutcomeAI referral visits, 90d
Unit of inferenceSite
Reported asShare of sites ahead
EstimandAssociation
Measured2026-09-09
Finding 01 · the ranking
79%of sites, the FAQ block comes out ahead

On 130 sites out of 164 where the comparison is decided, pages that got an FAQ block earned more AI referral visits than same-page-type pages on the same site that did not. Every change we can test this way comes out ahead more often than not, and the FAQ block leads.

What a win rate is not

It says how often a change is ahead. It does not say by how much, and we are not publishing a magnitude, because the magnitude moves depending on how you measure it. The FAQ block reads as 2.5 times on pooled visits and the H1 tag as roughly half, while at site level both come out ahead. A few high-traffic pages inside a few sites drive the pooled version, which is exactly why the site is the unit here.

ChangeSites won
FAQ block130 of 164 sites · p 6e-14
79%
Meta title119 of 154 sites · p 1e-11
77%
Heading and body structure126 of 184 sites · p 5e-7
69%
H1 tag103 of 159 sites · p 2e-4
65%
llms.txtmeasured from crawler logs, not referrals
none

Share of sites where the change came out ahead, against a 50% no effect line. llms.txt is on a different footing: it is measured from crawler logs rather than referrals, because AI crawlers fetched it 31 times across 82 million requests in six months.

What each change is, and what the number says
FAQ blockfaq-optimization

A question and answer section generated from the page's own content, rendered visibly and mirrored into FAQPage markup.

79% of sitesahead
Meta titlemeta-title

The title tag rewritten to carry the specifics a person would actually ask for.

77% of sitesahead
Heading and body structureheading-optimization

Dense unstructured blocks rewritten to lead with the answer, in place, many times per page.

69% of sitesahead
H1 tagh1-tag

The page's single main heading rewritten to name the thing the page is about.

65% of sitesahead
llms.txtllms-txt

The file the field has spent two years recommending. Measured from crawler logs rather than from referrals, because the question is prior to that: across 82 million crawler requests over six months, AI crawlers fetched llms.txt 31 times in total. A file that is not read cannot be moving anything.

no effect found

Finding 02 · what kind of page it is matters more than anything we do to it

The single biggest thing we learned building this is a warning about reading any table like it. Before page type was matched, heading structure looked like a +1.09 point winner. Adding page type to the match took it to +0.25, inside the noise. The first number was mostly telling us which templates our heading pass fires on.

5.29%of optimized standalone pages earned an AI referral in 28 days
2.54%of product pages
1.92%of blog posts, which is a 2.8x spread across templates
2.77%across the whole matched sample, which is why the ranking above reports how often a change is ahead rather than by how much

So every comparison above is made between pages of the same type, on the same site, carrying the same number of changes. Only the composition of the bundle differs.

Two things that do not move it

Worth as much as the ranking. One of these is the single most widely recommended change in this field.

The size of the rewriteno benefit found

How much of the page there was to fix matters. How much better we made each piece of it does not, at least not measurably. Pages where the answerability score rose the most are not ahead of pages on the same site where it rose the least, once the number of rewritten blocks is matched: 0.14 points against a 2.68% baseline, inside the noise on 37,360 pages.

JavaScript renderingno benefit found

Not that it does not work, but that it cannot be measured here, and for an interesting reason. Every optimized page is served with an AI-readable HTML version, so there is no optimized page without it to compare against. And the problem it solves is rarer than the field assumes: on 4,499 audits in six months, 76.1% of sites already scored 9 or 10 out of 10 for serving their main content without JavaScript. Only 8 sites in the whole panel scored below 9 AND had enough pages in both arms to compare, which is no sample.

And the one that dwarfs them

5.2x

None of the changes above comes close to the effect of how much the page actually says. Pages carrying sixteen or more substantial content blocks earn an AI referral 5.2 times as often as pages carrying one, and the gradient is monotone all the way up.

1 block2.43%
2 to 32.93%
4 to 73.86%
8 to 155.60%
16+12.63%

Descriptive, and uncontrolled. A page with sixteen substantial content blocks differs from a one-block page in more ways than length, so this is not a claim that adding words earns citations. The evidence here is that the gradient is monotone across all five bands on 42,913 pages, and that it is several times larger than any per-change number on this page.

IndexJavaScript readability

How often is the content not in the HTML at all?

An assistant that has to run your JavaScript to find your prices usually does not. Every page we optimise is served with an AI-readable HTML version for that reason, which also means there is no optimised page without one, so it cannot appear in the ranking above: it has no comparison group.

What we can report is how common the problem is, and it is less common than this field implies. On 4,499 audits over the last 180 days, 76.1% of sites already serve their main content without JavaScript. About one in nine does not, and 6.1% are severe.

Auditor readability scoremedian 9
Content is in the HTML76.1%
Mostly, with gaps12.6%
Significant parts hidden5.2%
Content needs JavaScript6.1%

One audit per site, on one page, so this describes homepages more than whole sites.

Changes held out of the ranking

6 of 10

Listed rather than dropped, because a leaderboard that quietly omits what it could not measure reads as though it measured everything. The rule was fixed before the estimates were looked at: a change qualifies only if it is applied because of a content defect the change itself fixes, and if it existed for the whole window. A change applied because of what kind of page it is, or because of what markup the page already carries, cannot be pulled apart from that page characteristic by matching. No numbers are given for these, because the number would be a fact about page selection wearing the label of a change.

Article schema rewriteSelects on existing markup
article

70.6% of these changes land on a page that already carried Article, NewsArticle or BlogPosting markup, usually from an SEO plugin. Pages that get it are therefore pages an SEO plugin already manages, which matching on page type does not fix. It is also a replacement rather than an addition, so even a clean estimate would answer a narrower question than whether Article schema helps.

Quick answer blockWindow too short
quick_answer

Shipped 2026-07-01, so it exists only in the last six weeks of a five and a half month sample and its pages have a different calendar from every other row's. It is on 43.8% of optimizations now and is the first candidate for the randomized readout.

Product schemaSelects on page type
json-ld-product

Applied to product pages by definition, so within any sample that contains other page types its association is largely the association of being a product page. After matching, only 11 sites still contribute.

Product category schemaSelects on page type
json-ld-product-category

The same problem, on 0.6% of optimizations. After matching, 1,063 pages and 8 referral events remain, which is no power at all.

Organization schemaNo comparison group
json-ld-organization

On every optimization we shipped last month. A change with no untreated comparison pages has nothing to be compared against, at any sample size.

Machine-readable change logNo comparison group
invisible-changes-bar

Also on every optimization, and it is the record of the other changes rather than a change to the page. After matching, 180 pages remain.

Method, and two estimates we threw away

The design above is the third one. The first two both compared optimized pages against untouched pages, and both produced large, significant, wrong answers.

01

Optimized pages against never-optimized pages, pooled

discarded

One site supplied 63,836 of the 68,300 control pages, at a 67.6% baseline crawl rate while most sites sat at 0.0%. This is one site's month.

+45.1pp
02

Same, but within site and within baseline stratum, median across sites

discarded

Comparing treated pages against untreated pages compares groups that differ in importance, age and page type as well as in treatment. Nothing that keeps an untreated control group in this data survives those differences.

+9.85pp, sign test p = 0.0037
03

Optimized pages against each other, matched on site, bundle size and page type

stands

Every page is treated, so treatment status cannot confound anything. What is left is the composition of the bundle, matched on the two things that otherwise drive the outcome.

Published above
Why there is no before and after here

Across 89,231 pages optimized between March and August 2026, the median page was optimized 3.04 days after it was first discovered and 89.3% within four weeks. A pre-period drawn before the optimization therefore falls partly before the page existed, which is what inflated the two discarded estimates. Comparing optimized pages with each other removes the problem, because every page in the sample sits the same distance from its own optimization date.

Why the outcome is a referral and not a crawl

At page level AI crawls are almost all zeros: of 2,343 active pages in a 28-day window, 8.9% received any AI crawl and the 90th percentile is 0. A referral visit is both further down the funnel and better behaved as an outcome: 2,088 of 75,496 pages received one. Crawls are a secondary outcome for the randomized study.

Prechecks · each one could have stopped publication
PASSReferral events join to pagesYes, on scheme-stripped path with the tracking query removed. 20,454 distinct paths receive a referral, 76.3% of them below the homepage.
PASSOutcome has supportYes at page level for referrals: 2,088 of 75,496 pages, 2.77%. Not at page level for crawls, where p90 = 0, so crawls are a secondary outcome for the randomized study.
FAILEstimate stable to page typeNo, and this is the important one. Adding page type to the match moved heading structure from +1.09pp to +0.25pp. Page type is now matched exactly.
FAILEstimate stable to bundle sizeYes once matched. Bundles run 4.4 changes without a given type and 7.1 with, so this has to be matched rather than adjusted.
PASSNo single site carries a rowYes. Between 271 and 310 sites contribute to each published row, and the site-level split is reported alongside every estimate.
PASSComparison is not treated against untreatedYes. Every page in the sample was optimized, which is what the two discarded estimates got wrong.
FAILChanges separable from each otherPartly. Matching on bundle size and page type isolates composition, but changes still co-occur, so a row is an association and not a randomized effect.
FAILTrigger labels available to audit selectionNo. trigger_source is NULL for all 144,183 completed optimizations before 2026-08-11, which covers most of this window.

4 of 8 prechecks failed. Each failure changed the design rather than being written off. The earlier pooled estimate of +45.1pp is what the same data yields if you skip them.

Turning the top row into an effect

allocation locked

Matching gets an association. Only randomizing which pages get a change turns it into a causal effect, and that is the next study.

01

Randomize at the page level within a site, among pages that have all existed for at least eight weeks, so page age cannot enter the comparison at all.

02

Withhold one change type at random from half of the treated pages. That, and only that, turns a row in this table from an association into an effect.

03

Start with the FAQ block, since it is the row this table detects, and with the quick answer block, since it is the newest change and has never been measured.

04

Pre-register the window, the binary outcome and the site-level sign test before looking, since the outcome is sparse and the researcher degrees of freedom are large.

Where that experiment stands
CompleteAllocation drawn and lockedBlock-randomized from a fixed seed against a pinned page-age cutoff, and reproducible: a second run has to reproduce the same SHA-256 digest or the analysis refuses to start.
CompleteEstimator written before treatmentThe readout script is finished and pre-registered, and checks the allocation digest before it will report anything. It cannot be tuned after seeing the result.
CompletePlacebo run passedRun on two windows that both precede treatment, the estimator returns +0.00pp with a 118/117 split of sites and a sign test of p = 1.0000. It finds nothing where nothing happened.
Not startedTreatmentNot started. Requires authorization to withhold one change from a random half of eligible pages for a single 28-day window on live customer sites.

The order matters more than any single step: the allocation and the estimator were both finished and fixed before a single page was treated, and the estimator refuses to run against an allocation whose digest does not match. Nothing about the analysis can be chosen after the result is visible.

Limits

This is an association, not a randomized effect. We do not choose at random which changes a page gets, so a row describes the pages that received a change rather than what the change does to a page.

The sample is paying customers, overwhelmingly small and medium business sites on WordPress, so every figure describes those sites and not the web at large.

The outcome is an AI referral visit, meaning a person clicked through from an assistant. It is the end of the funnel rather than a crawl, which is deliberate, but it means a page can be read and cited without ever showing up here.

Referrals are attributed from tracking parameters and referrers. Assistants that strip both are invisible to it, and chatgpt.com is 84% of what remains, so the ranking is weighted toward how one assistant behaves.

trigger_source is NULL for every completed optimization before 2026-08-11, so we cannot check how many optimizations were themselves triggered by a content change, which is independently a reason a page might get read.

34.8% of pages are optimized more than once. Only the first optimization defines the bundle and the window.

What this does not claim

That an FAQ block causes AI referrals. It is associated with them, which is the strongest claim an observational design supports. Only the randomized study can promote that to a cause.

That the three rows with no detected difference do nothing. Their intervals include zero, which means this design cannot separate them from zero, not that they are zero.

That the excluded changes do not work. They are excluded because the comparison would measure page selection, not because it came out badly.