How we measure.
KailxLabs measures AI visibility by asking each buyer question several times on the consumer AI apps buyers use, and counting the answers that name you. Every share carries its count and a likely range, and it becomes a percentage only once eight answers stand behind it. A change counts only when the before and after ranges separate, and a half of your questions is held back untouched to prove the work caused it.
Four rules, no exceptions.
Most AI visibility numbers flatter because of how they are collected. Ours follow four rules, in line with the IAB's August 2026 guidance that one response is not a measurement.
The apps your buyers use
ChatGPT on chatgpt.com, Google AI Overviews, Google AI Mode, Gemini and Microsoft Copilot, plus Perplexity through Sonar, the engine behind perplexity.ai. Developer APIs are labelled and never mixed into a headline.
Every number has a count
Answers change from run to run, so our agents ask each question several times. Every share shows the answers behind it and a likely range, and it becomes a percentage only once eight answers stand behind it.
One platform at a time
Platforms cite different sources, so each gets its own row. A change counts only when the before and after ranges stop overlapping.
Proof it was us
We hold back an untouched half of your questions. If the worked half moves and the untouched half does not, that is evidence. If both move, we tell you.
Everything below runs on our own measurement instrument, built by our founder. It stores every answer, every source and the exact question text sent, and every answer can be exported as CSV and JSON.
Which apps, and why.
Buyers ask the apps, not the developer APIs. The two give different answers from different sources, so a number read from an API can flatter or mislead.
| AI app | Where we read it |
|---|---|
| ChatGPT | chatgpt.com |
| Google AI Overviews | Google Search |
| Google AI Mode | Google Search |
| Gemini | gemini.google.com |
| Microsoft Copilot | copilot.microsoft.com |
| Perplexity | Sonar, the engine behind perplexity.ai |
of cited domains overlap between the ChatGPT API and the ChatGPT app on the same questions (Gemini: 14.8%). The ChatGPT and Gemini apps share 5.4%.
arXiv 2609.18729, academic audit · 1,536 answers · 16 Sep 2026
That is why every headline number we report comes from the consumer apps. Perplexity is read through Sonar, its own engine, and labelled that way. ChatGPT is read in its default mode. Claude has no consumer app anyone can measure, so any Claude figure we show is labelled as an API number.
Why one answer proves nothing.
Our agents check your priority questions every day on ChatGPT and on a fixed schedule on every other app, several runs each time, so one odd answer never decides anything.
of the 46,009 URLs AI engines cited were cited only once across 40 to 91 days of repeated runs. The brands named held far steadier than the sources behind them.
Otterly, AI search visibility stability · 252,407 answers, 520 prompts, 7 engines · 11 Sep 2026 · vendor data
response is not a measurement. The IAB framework asks for repeated runs, results per platform and ranges.
IAB, "Measuring Visibility in the AI Era" · Industry framework · 3 Aug 2026
So we ask each question several times on each app, keep every answer, and report counts over time. A single run that looks good or bad never decides anything.
What counts as named.
Whether an answer names you is decided by exact matching, never by a model's opinion.
| What we record | How it is decided |
|---|---|
| Answers that name you | Exact matching on your name and the aliases you declare, on whole words. Text inside links and citation pills does not count, and a match inside a longer rival name belongs to the rival. |
| Answers that link your site | One of your own domains appears among the sources the answer cites. |
| Answers that do both | Named and linked in the same answer. We report it on its own and never merge it with the other counts. |
| Answers that recommend you | A model grades how a naming answer presents you. The grade is kept only if it quotes a sentence that appears word for word in the answer. Otherwise the answer stays ungraded, and ungraded answers are left out of that count. |
| Failed answers | An error or an empty answer is a missing run, left out of every count. It never counts as the AI ignoring you. A search that shows no AI Overview is left out too, and we say so. |
| Questions with your name in them | Tracked apart and never used to lift your headline. |
How sure each number is.
Every share comes with its count and a likely range, so you can see how much weight it carries.
A count, never a percentage
With fewer than eight answers behind it, a share is shown as a count, such as 4 of 5 answers, with an early read label. A percentage would claim more certainty than the data has.
A share with a likely range
The range is a Wilson style 95% interval. For example, 8 of 10 answers is a likely range of 49 to 94%. More answers make the range narrower.
Widened, never narrowed
When a figure pools several questions, answers to the same question tend to agree with each other. The range is widened to allow for that, so it is never narrower than the data earned.
Only when ranges separate
A change is called real only when the before and after ranges stop overlapping, measured after each app has had time to show it. Smaller moves are labelled normal variation. Drops are reported as plainly as gains.
Blank, never zero
A question we have not measured on an app shows as not measured. A zero is written as none of the answers we read, with the number of answers behind it, never as the word never.
Prove it, or call it null.
Before any work begins, we split your questions into a worked half and an untouched half, and write the decision rule down. We only work on the first half. The second half shows what would have happened anyway: a model update, a season, a rival's launch.
Each month the two halves are compared, and the result is one of three verdicts. The rule is fixed in advance, so nobody can pick the verdict after seeing the numbers.
Each fix is also judged on the number it can move: work on your own pages on links to your site, earned mentions elsewhere on answers that name you.
What we never say.
Some numbers sound impressive and mean little. We leave them out.
AI answers have no stable position. We never give you a rank in ChatGPT or any other AI app, only how often you appear, with the count and range.
Each AI app is reported on its own row. We never average engines into one visibility score, and developer APIs are never blended with consumer apps.
No average uplift, no invented percentages and no results from other agencies' clients. KailxLabs publishes its own numbers on our numbers page.
Nobody controls what an AI says. We never promise a fixed spot or a rank.
Schema, llms.txt and FAQ markup are good hygiene, and we set them up properly. No good evidence shows they make AI cite you, and Google says it does not use llms.txt.
A recommendation grade stands only with the quoted sentence behind it. We do not yet publish an accuracy figure for that grade.
What measurement cannot see.
- Answers vary by location. Measuring for a city or a country changes the question text we send, not the physical place it is asked from. The report prints the market and the exact question text.
- Answers vary by account. We read the apps without a signed in user. Personal memory, chosen sources and past chats can change what one real buyer sees.
- Answers vary over time. Models and their sources change, sometimes weekly. That is why we keep measuring and keep an untouched half.
- Some modes are out of reach. ChatGPT is read in its default mode. Perplexity is read through Sonar, its own engine, rather than the perplexity.ai page.
- Correlation is not cause. Outside the worked and untouched comparison, a pattern in the answers is a pattern, not proof of why.
Method questions, answered.
Why not just take a screenshot of ChatGPT?
One answer is not a measurement. Answers change from run to run and from app to app, so we ask each question several times and report the count behind every figure. The IAB's August 2026 framework asks for the same: repeated runs, results per platform and ranges.
Why do you measure the apps and not the APIs?
Because buyers use the apps, and the API is a different surface. An academic audit found the ChatGPT API and the ChatGPT app share only 12.0% of cited domains on the same questions. Developer APIs are measured only when you ask, and labelled.
What is a likely range?
The span the true share probably sits in, given how many answers we read. Fewer answers mean a wider range. We show it so you can tell a real change from normal variation.
Do you give us a rank in ChatGPT?
No. AI answers have no stable position to rank. We report how often you are named, linked and recommended, each with its count and range. Classic Google positions are real ranks and are tracked separately when they matter.
Who built the measurement instrument?
Our founder built it, and we run it ourselves. It stores every answer, every source and the exact question text sent, and every answer can be exported as CSV and JSON.
What if nothing moves?
Then the verdict is null, and you see it. A null is a result. We say so plainly, and we say what we would change next.
- 26 Sep 2026: method rewritten around repeated runs on consumer apps, likely ranges and the untouched half.
Talk to the founder.
20 minutes with Kailesk Khumar. You leave knowing where AI names you today, and what it would take to become the answer.
- A live look. We ask five of your buyer questions on ChatGPT, Google and Perplexity while you watch.
- The reasons. The three biggest reasons AI names a rival instead of you.
- The plan. A clear next step and its price, whether or not you hire us.
Pick a time that suits you. The calendar invite arrives straight away. Prefer a link? Book on Cal.com.