The AI Visibility Methodology Behind Every Number
This page is the AI visibility methodology in full, because every number on your panel should be something you can check rather than something you have to trust. You give us a question your buyers actually ask. We put it to every engine, every day, from the place you pick. Then established statistics decide what any of it means, because a number nobody tested is just a screenshot.
The answers are only half of the measurement. A citation is an engine telling you that a page held something it wanted in its answer, so every page it cites gets measured as closely as the answer does.
What We Ask and What We Count
A prompt is the question itself, in the words a buyer would use, not a keyword:
What is the best appointment scheduling software for dental practices?
That exact sentence runs daily on Google AI Overview, Google AI Mode, ChatGPT, Perplexity, Gemini and Copilot, from the location you set. Every answer is kept, along with who it cited and where you ranked that day, for as long as you hold the account.
Two rates come out of that. How often the engines answer your question at all, and how often they cite you when they do. A day with no AI answer is not a day you lost, so it never counts against you.
The Statistics Behind Every AI Visibility Number
Each number on your panel is guarded by an established test, chosen for the specific way that number could mislead you.
| The risk | What we use |
|---|---|
| A rate from a handful of checks looks like a fact | Wilson confidence interval on every rate, with its sample size |
| A month-to-month move is really just noise | Fisher's exact test |
| Run enough comparisons and chance hands you a winner | Benjamini-Hochberg correction across every test that period |
| Two rolling windows share most of their days, so comparing them is invalid | Periods are adjacent calendar months, disjoint by construction |
| Checking every day and firing on the first crossing is p-hacking | The change test runs once per period, and the run is recorded |
| Six engines get treated as six repeats of one measurement | Cochran-Mantel-Haenszel, engines kept as separate strata |
| The engines disagree and a combined number hides it | Disagreeing engines are counted, and a split verdict is reported per engine instead of pooled |
| A finding rests on one lucky day | The result is retested with its best day removed |
| The whole field moved and you got credit for it | Your move is measured against what the other cited domains did |
Nothing is called significant until it survives all of it. When the tests run and nothing survives, the report tells you how many ran and that none did.
Two consequences worth stating plainly. Descriptive numbers update continuously, so your rates and their intervals move every day. The test that can fire an alert does not: it waits for a month boundary and runs once. Which means your first rates appear immediately, and your first verdict on whether anything changed arrives at the end of your first full calendar month.
When the Data Cannot Support a Conclusion
Every rate is labelled with how much weight it can carry. Too few checks and the data is shown with no conclusion drawn from it, and the panel says so in those words.
Nineteen checks with no citation read as 0%, ranging 0% to 17%. Not proof of never, just evidence of not much, and the interval narrows in front of you as the checks accumulate.
Not significant is never reported as no change. The two are different statements and only one of them is something we measured.
Most months the honest answer is that nothing moved. You will get told that instead of a manufactured alert.
Who Else Got Cited in the Answer
You never type in a competitor list. Every domain the answers cite goes on it automatically, with the days it held and its best position.
That list routinely holds sites your rank tracking cannot see, because being cited and ranking are different things. On Google we keep the ordinary results from the same page on the same day, so you can also see who got cited without ranking at all.
How We Parse the Pages AI Engines Cite
Every page an answer cites is fetched and reduced to what it is actually about. Your page goes down the same path, so the two sides are measured identically.
| The problem | What we use |
|---|---|
| Nav, sidebar and footer text swamps the real content | Boilerplate removal, so only the body of the page is measured |
| Single words are too crude to name a topic | N-gram extraction at one, two and three words, so phrases count as phrases |
| Common filler outranks the words that matter | TF-IDF weighting, modified for this job, so a phrase counts for how distinctive it is rather than how often it appears |
What comes back is the phrases the cited pages have in common, held against your page: what is missing, what you use lightly, and which themes are gaining or fading.
The Thirty-Three Characteristics We Correlate
Every cited page is measured on thirty-three characteristics, and so is yours. Here they are, all of them:
- Writing - word count, Flesch-Kincaid grade, Flesch-Kincaid ease, SMOG, Coleman-Liau
- Headings - H1 count, H2 count, H3 count, total headings
- Links - internal links, external links
- Media - images, images carrying alt text, alt text ratio, whether the page carries video
- Structured data - whether schema is present, how many schema types are declared
- Page structure - semantic elements, DOM elements, HTML bytes, text bytes, link density, text density
- Metadata - title length, meta description length, H1 length
- Delivery - HTTPS
- Speed - Lighthouse score, LCP, CLS, TTI, TBT, FCP
We measure all of them, the ones the industry swears by and the ones that may well be myth. No metric is left out because it does not fit a theory, and none is included because we like it. Then we compare those characteristics across the pages winning a given answer, day after day, to see which ones the winners genuinely share.
We are not telling you what works. We are measuring what is happening and handing you the analysis. Some characteristics turn out to be table stakes, where every cited page agrees and there is nothing to tune. The rest are reported as distance from where the winners sit, and being unusual on the good side is called a difference rather than a fault.
Which Page of Yours Competes
We pick it out of your own site crawl. Nothing to configure, no page for you to nominate.
The page is chosen by a match score built from the prompt's words where they appear in your title, H1 and meta description, from your page's own distinctive phrases at one, two and three words, from the anchor text of links pointing at it, and from its PageRank inside your site. The floor is set deliberately above what one incidental word in a meta description would score, because measuring an off-topic page would manufacture a content gap that is really a matching failure.
When nothing on your site clears that floor, that is the finding, and it is the most useful one on the screen.
The Three Answers AI Visibility Tracking Gives You
Every prompt resolves to one of three, and they send you somewhere different.
Write a new page. The cited pages keep covering a topic your site never addresses. That is a gap in what you publish, not a flaw in a page you already have.
Improve the page you have. You have the right page and it differs from the ones the engines are choosing in ways we can name. You get one first thing to fix, not an audit dump.
Go get link authority. Your page carries every topic the winners share, at comparable weight. Nothing is left to add on the words, so more writing is not the answer and we will say so rather than keep you busy. This is the verdict that hands you off to the Backlink Prospector, because the work left is earning links rather than editing copy.
No Language Model Writes Any of This
No language model reads, scores, judges or writes anything in this analysis. Every number is arithmetic over what was measured, and AI answers are the thing being measured, never the thing doing the judging.
The report is assembled, not generated. Once the measurements exist they select from fixed sentence fragments and fill in their own values, so a readable report gets built without a word of it being written by a model. The same measurements always assemble the same report. That is a guarantee about the assembly and not about the day, because nobody can promise the engines will hand over the same measurements twice. What is promised is that nothing between the measurements and your screen adds anything of its own, and that every sentence is bolted to a number in a table on the same screen.
That matters because of what is normal elsewhere. It is common in this category for a language model to summarise, score or interpret the results before you see them, which makes the report itself something that cannot be reproduced. Here the engines are the only unpredictable thing in the building, and they are the thing being measured.
What This AI Visibility Methodology Does Not Claim
We work with the same AI engines and chat surfaces as everyone else. What differs is what happens to the observations after they land.
Correlation is not a promise. Cited pages tending to be longer is an observation, and we will not turn it into lengthen yours and you will be cited.
We cannot make an engine cite you. What this AI visibility methodology gives you is a measured account of who it cites, how your page differs from theirs, and whether anything actually changed. You can start tracking AI visibility on your own prompts and read every number back to the day it was captured.
Create your free LinkMap and put this into practice on your own pages.