Method note ·

What can an LLM SEO tool actually measure?

Two products carry the same name and observe different things. One watches the engines. The other reads the page. Which one you have decides what your numbers mean.

The term covers two jobs

A tracker sends prompts to AI engines on a schedule and counts how often a brand comes back. That is share-of-answer measurement. It requires live querying, it costs something per run, and its answer moves between runs because the engines do.

A readiness auditor reads a page and reports whether an engine could lift an answer out of it. No live querying, no per-check cost, and a stable answer, because the page is what is being measured.

Both are useful. They are not substitutes. A readiness score read as a citation count is a wrong reading of the number. WebmasterNinja is the second kind, which is worth stating plainly before any number appears.

What a page can be measured on

The readiness score is two layers multiplied:

finalScore = contentScore × eligibility

Content quality comes from five signals: evidence density, answer shape, trust markers appropriate to the page's entity type, freshness, and fundamentals. Each carries a weight anchored to a published effect size, and the weights shift with the page's detected query intent.

Eligibility is three gates, and they multiply rather than deduct: is the page indexed, does its content survive without JavaScript, does robots.txt admit the citation crawlers. The multiplication is the interesting part. A blocked crawler or a JavaScript-only page makes a page uncitable, and no amount of good writing multiplies past a zero. Treating those as small point deductions would let strong content outweigh being invisible, which is the wrong shape for the problem.

Three things that look measurable and are not

Several signals circulate as AI-visibility levers without evidence behind them. They are left out rather than scored:

There is a fourth case worth separating out. Some queries cannot be won on-page at all. Where a query reads as reputation or comparison intent, AI tends to cite third-party sites rather than a brand's own page. That is a ceiling, not a gap, and the honest move is to say so rather than propose a fix for it.

Where a language model sits in this

Not in the score. Every scored row is a rule with a stated threshold and a citation.

One surface is generative: drafting a title or meta description. Each candidate passes five checks before it is shown. It must contain the page's real primary topic, avoid superlatives and brand promises, invent no number absent from the page's own data, and invent no brand name. When nothing survives, the honest output is that no grounded draft could be produced, which is more useful than a plausible one.

The line worth holding

Whether ChatGPT, Perplexity or Google actually quote a page shifts by run and by day. Observing that needs repeated live checks against each engine, which is a tracker's job.

A readiness score measures the part that stays still and that you control. It is the more actionable of the two, and it is also the easier one to overclaim, which is why the panel carries the distinction in its own interface rather than only in documentation.

Every threshold behind the score is published on the methodology page, including the ones flagged as weakly sourced.

Read every threshold Add to Chrome · free