Methodology

How we measure.

KailxLabs measures AI visibility by asking each buyer question several times on the consumer AI apps buyers use, and counting the answers that name you. Every share carries its count and a likely range, and it becomes a percentage only once eight answers stand behind it. A change counts only when the before and after ranges separate, and a half of your questions is held back untouched to prove the work caused it.

Last updated 26 Sep 2026

The rules

Four rules, no exceptions.

Most AI visibility numbers flatter because of how they are collected. Ours follow four rules, in line with the IAB's August 2026 guidance that one response is not a measurement.

Rule 1

The apps your buyers use

ChatGPT on chatgpt.com, Google AI Overviews, Google AI Mode, Gemini and Microsoft Copilot, plus Perplexity through Sonar, the engine behind perplexity.ai. Developer APIs are labelled and never mixed into a headline.

Rule 2

Every number has a count

Answers change from run to run, so our agents ask each question several times. Every share shows the answers behind it and a likely range, and it becomes a percentage only once eight answers stand behind it.

Rule 3

One platform at a time

Platforms cite different sources, so each gets its own row. A change counts only when the before and after ranges stop overlapping.

Rule 4

Proof it was us

We hold back an untouched half of your questions. If the worked half moves and the untouched half does not, that is evidence. If both move, we tell you.

Everything below runs on our own measurement instrument, built by our founder. It stores every answer, every source and the exact question text sent, and every answer can be exported as CSV and JSON.

Surfaces

Which apps, and why.

Buyers ask the apps, not the developer APIs. The two give different answers from different sources, so a number read from an API can flatter or mislead.

The AI apps KailxLabs measures and where each is read
AI appWhere we read it
ChatGPTchatgpt.com
Google AI OverviewsGoogle Search
Google AI ModeGoogle Search
Geminigemini.google.com
Microsoft Copilotcopilot.microsoft.com
PerplexitySonar, the engine behind perplexity.ai
The app is not the API
12.0%

of cited domains overlap between the ChatGPT API and the ChatGPT app on the same questions (Gemini: 14.8%). The ChatGPT and Gemini apps share 5.4%.

arXiv 2609.18729, academic audit · 1,536 answers · 16 Sep 2026

That is why every headline number we report comes from the consumer apps. Perplexity is read through Sonar, its own engine, and labelled that way. ChatGPT is read in its default mode. Claude has no consumer app anyone can measure, so any Claude figure we show is labelled as an API number.

Repeated runs

Why one answer proves nothing.

Our agents check your priority questions every day on ChatGPT and on a fixed schedule on every other app, several runs each time, so one odd answer never decides anything.

Sources move fast
48%

of the 46,009 URLs AI engines cited were cited only once across 40 to 91 days of repeated runs. The brands named held far steadier than the sources behind them.

Otterly, AI search visibility stability · 252,407 answers, 520 prompts, 7 engines · 11 Sep 2026 · vendor data

The industry standard
1

response is not a measurement. The IAB framework asks for repeated runs, results per platform and ranges.

IAB, "Measuring Visibility in the AI Era" · Industry framework · 3 Aug 2026

So we ask each question several times on each app, keep every answer, and report counts over time. A single run that looks good or bad never decides anything.

Counting

What counts as named.

Whether an answer names you is decided by exact matching, never by a model's opinion.

How KailxLabs decides what each answer counts as
What we recordHow it is decided
Answers that name youExact matching on your name and the aliases you declare, on whole words. Text inside links and citation pills does not count, and a match inside a longer rival name belongs to the rival.
Answers that link your siteOne of your own domains appears among the sources the answer cites.
Answers that do bothNamed and linked in the same answer. We report it on its own and never merge it with the other counts.
Answers that recommend youA model grades how a naming answer presents you. The grade is kept only if it quotes a sentence that appears word for word in the answer. Otherwise the answer stays ungraded, and ungraded answers are left out of that count.
Failed answersAn error or an empty answer is a missing run, left out of every count. It never counts as the AI ignoring you. A search that shows no AI Overview is left out too, and we say so.
Questions with your name in themTracked apart and never used to lift your headline.
Ranges

How sure each number is.

Every share comes with its count and a likely range, so you can see how much weight it carries.

Under eight answers

A count, never a percentage

With fewer than eight answers behind it, a share is shown as a count, such as 4 of 5 answers, with an early read label. A percentage would claim more certainty than the data has.

Eight answers or more

A share with a likely range

The range is a Wilson style 95% interval. For example, 8 of 10 answers is a likely range of 49 to 94%. More answers make the range narrower.

Pooled questions

Widened, never narrowed

When a figure pools several questions, answers to the same question tend to agree with each other. The range is widened to allow for that, so it is never narrower than the data earned.

A real change

Only when ranges separate

A change is called real only when the before and after ranges stop overlapping, measured after each app has had time to show it. Smaller moves are labelled normal variation. Drops are reported as plainly as gains.

Nothing measured

Blank, never zero

A question we have not measured on an app shows as not measured. A zero is written as none of the answers we read, with the number of answers behind it, never as the word never.

The control

Prove it, or call it null.

Before any work begins, we split your questions into a worked half and an untouched half, and write the decision rule down. We only work on the first half. The second half shows what would have happened anyway: a model update, a season, a rival's launch.

Each month the two halves are compared, and the result is one of three verdicts. The rule is fixed in advance, so nobody can pick the verdict after seeing the numbers.

Each fix is also judged on the number it can move: work on your own pages on links to your site, earned mentions elsewhere on answers that name you.

Illustration · fictional brand Verdict: positive
PositiveThe worked half moved beyond normal variation. The untouched half did not.
NullNeither half moved beyond normal variation. You see that too.
InconclusiveBoth halves moved, the untouched half moved more, or too few answers stand behind the change.
What we do not claim

What we never say.

Some numbers sound impressive and mean little. We leave them out.

No single rank

AI answers have no stable position. We never give you a rank in ChatGPT or any other AI app, only how often you appear, with the count and range.

No blended score

Each AI app is reported on its own row. We never average engines into one visibility score, and developer APIs are never blended with consumer apps.

No borrowed results

No average uplift, no invented percentages and no results from other agencies' clients. KailxLabs publishes its own numbers on our numbers page.

No promised spot

Nobody controls what an AI says. We never promise a fixed spot or a rank.

No levers that are not levers

Schema, llms.txt and FAQ markup are good hygiene, and we set them up properly. No good evidence shows they make AI cite you, and Google says it does not use llms.txt.

No grade without a quote

A recommendation grade stands only with the quoted sentence behind it. We do not yet publish an accuracy figure for that grade.

Limits

What measurement cannot see.

  • Answers vary by location. Measuring for a city or a country changes the question text we send, not the physical place it is asked from. The report prints the market and the exact question text.
  • Answers vary by account. We read the apps without a signed in user. Personal memory, chosen sources and past chats can change what one real buyer sees.
  • Answers vary over time. Models and their sources change, sometimes weekly. That is why we keep measuring and keep an untouched half.
  • Some modes are out of reach. ChatGPT is read in its default mode. Perplexity is read through Sonar, its own engine, rather than the perplexity.ai page.
  • Correlation is not cause. Outside the worked and untouched comparison, a pattern in the answers is a pattern, not proof of why.
See the method on your questions.A baseline measures 50 of your buyer questions this way, for $1,500.
See the baseline
FAQ

Method questions, answered.

Why not just take a screenshot of ChatGPT?

One answer is not a measurement. Answers change from run to run and from app to app, so we ask each question several times and report the count behind every figure. The IAB's August 2026 framework asks for the same: repeated runs, results per platform and ranges.

Why do you measure the apps and not the APIs?

Because buyers use the apps, and the API is a different surface. An academic audit found the ChatGPT API and the ChatGPT app share only 12.0% of cited domains on the same questions. Developer APIs are measured only when you ask, and labelled.

What is a likely range?

The span the true share probably sits in, given how many answers we read. Fewer answers mean a wider range. We show it so you can tell a real change from normal variation.

Do you give us a rank in ChatGPT?

No. AI answers have no stable position to rank. We report how often you are named, linked and recommended, each with its count and range. Classic Google positions are real ranks and are tracked separately when they matter.

Who built the measurement instrument?

Our founder built it, and we run it ourselves. It stores every answer, every source and the exact question text sent, and every answer can be exported as CSV and JSON.

What if nothing moves?

Then the verdict is null, and you see it. A null is a result. We say so plainly, and we say what we would change next.

Last updated 26 Sep 2026
  • 26 Sep 2026: method rewritten around repeated runs on consumer apps, likely ranges and the untouched half.
Book a call

Talk to the founder.

20 minutes with Kailesk Khumar. You leave knowing where AI names you today, and what it would take to become the answer.

  • A live look. We ask five of your buyer questions on ChatGPT, Google and Perplexity while you watch.
  • The reasons. The three biggest reasons AI names a rival instead of you.
  • The plan. A clear next step and its price, whether or not you hire us.

Pick a time that suits you. The calendar invite arrives straight away. Prefer a link? Book on Cal.com.

KK Kailesk Khumar AI visibility strategy call
20 min Video call Your time zone
Prefer email? Get a free written read instead
What you sell
How did you find us?

No newsletter. A reply from a person within two working days.

Got it. Kailesk will reply within two working days.
Talk to the founderBook a call