Glossary

Working definitions for AI search optimization.

The terms buyers meet when reading about AI search, ChatGPT citations and structured data, defined plainly. Where a term is folklore or only hygiene, the definition says so.

Discipline

Generative engine optimization (GEO)

Generative engine optimization (GEO) is the work of getting a brand named and cited inside answers written by AI apps such as ChatGPT, Google AI Mode, Gemini and Perplexity. It covers crawler access, clear facts on the brand's own pages, and presence in the third party sources those apps lean on when they answer.

Read the definition→
Discipline

Answer engine optimization (AEO)

Answer engine optimization (AEO) is the practice of making a brand the one AI apps name when buyers ask a question. It is judged by outcome: how often a brand is named across repeated runs of a fixed set of buyer questions, reported separately for each AI app, with the sample size and a likely range.

Read the definition→
Protocol

llms.txt

llms.txt is a proposed Markdown file at the root of a domain that summarizes a site for language models. It is optional hygiene, not a citation lever. Google has said its Search systems do not use it, and no major AI app has said it reads site llms.txt files when choosing what to cite.

Read the definition→
Content pattern

Answer Capsule

An Answer Capsule is a short paragraph near the top of a page or section that answers its question directly, names the subject in full and makes sense on its own. It is a house writing style for clarity. It is not a formula AI apps reward, and Google says rewriting or chunking pages for AI is not needed.

Read the definition→
Structured data

Schema.org @graph

A Schema.org @graph is a single JSON LD block that declares the entities on a page and links them through @id references, so the organization, its people, its services and the page itself form one connected description. It is good hygiene for search engines. Google says no special schema is needed for its AI features.

Read the definition→
AI crawler

GPTBot

GPTBot is the OpenAI crawler that collects public web content that may be used to train OpenAI models. It does not decide whether a site appears in ChatGPT search. That is the job of OAI-SearchBot, and fetches a user triggers inside ChatGPT come from ChatGPT-User. Blocking GPTBot is a training opt out, not a visibility choice.

Read the definition→
AI crawler

OAI-SearchBot

OAI-SearchBot is the OpenAI crawler that finds and indexes pages so they can appear, with links, in ChatGPT search answers. It is the OpenAI crawler that matters for being cited in ChatGPT. OpenAI says OAI-SearchBot is not used for training. ChatGPT-User is a separate agent that fetches a page when a user asks ChatGPT to open it.

Read the definition→
AI architecture

Retrieval Augmented Generation (RAG)

Retrieval Augmented Generation (RAG) is the pattern AI apps use when they search the web or an index before answering. The app retrieves candidate pages, reads the relevant passages and writes an answer that can cite them. RAG is why a brand can be named for recent facts without waiting for a model to be retrained.

Read the definition→
AI crawler

ClaudeBot

ClaudeBot is the Anthropic crawler that collects public web content that may be used to train Claude models. It is a training crawler. Search and citation in Claude come from separate agents: Claude-SearchBot indexes pages for Claude search results and Claude-User fetches a page when a person asks Claude to open it.

Read the definition→
AI crawler

PerplexityBot

PerplexityBot is the Perplexity crawler that finds and indexes pages so they can appear, with links, in Perplexity answers. Perplexity says it is not used to train foundation models. Perplexity-User is a separate agent that fetches a page in response to a user request.

Read the definition→
AI crawler

Google-Extended

Google-Extended is a robots.txt token, not a separate crawler. It lets a site choose whether Google may use its content to train Gemini models and to ground answers in the Gemini apps and Vertex AI. It does not affect Google Search, AI Overviews or AI Mode, which are governed by Googlebot.

Read the definition→
AI metric

Share of Model

Share of Model is an industry term for how often a brand is named in AI answers across a set of questions. A single blended score hides the fact that AI apps disagree, so report it per app: the named rate on each app, the number of answers behind it, and a likely range, never one figure across all apps.

Read the definition→
AI metric

Citation Frequency

Citation Frequency is the count of AI answers, in a measurement window, that name a brand or link to its domain. It is the raw count behind a named rate. It should always be reported with the number of answers checked and the AI app it came from.

Read the definition→
Content metric

Fact Density

Fact Density describes how many specific, checkable facts a page states, such as prices, locations, credentials, hours and policies, as opposed to general marketing claims. It is an editorial habit, not a scored metric. Pages that state the facts a buyer checks give AI apps and people something concrete to repeat.

Read the definition→
Protocol

ai.txt

ai.txt is an informal convention for stating, at the root of a domain, how a site permits AI training, retrieval and citation. It is not a standard, and no major AI app has said it reads it. Real controls live in robots.txt, per crawler, and in settings such as Google's Search generative AI control.

Read the definition→
Measurement protocol

Prompt Library

A Prompt Library is the fixed set of buyer questions run repeatedly on each AI app to measure whether a brand is named. It covers research, comparison, recommendation and buying intent, is agreed in writing before work starts, and does not change mid measurement, so before and after results are comparable.

Read the definition→
AI signal

Knowledge Graph Entity

A Knowledge Graph Entity is a business, person, place or concept that Google has recognized and given a unique identifier in its Knowledge Graph. A recognized entity is resolved consistently across queries, which helps Google Search and Gemini describe the right business when names are similar.

Read the definition→
Quality framework

E-E-A-T

E-E-A-T stands for Experience, Expertise, Authoritativeness and Trust. It is the framework in Google's Search Quality Rater Guidelines for judging content quality, weighted most heavily on Your Money or Your Life topics such as health, finance and law. Raters use it to assess results. It is not a single ranking score.

Read the definition→
Structured data

MedicalClinic Schema

MedicalClinic is a Schema.org type for a medical clinic. It sits under MedicalBusiness (a LocalBusiness type) and MedicalOrganization, so it inherits properties such as address, telephone and openingHoursSpecification, and adds clinic properties such as medicalSpecialty and availableService. It helps machines read what a clinic is. It is hygiene, not a citation lever.

Read the definition→
Structured data

LegalService Schema

LegalService is a Schema.org type for a law firm or legal practice. It extends LocalBusiness, and firms often pair it with a Person entry for each lawyer that lists bar admissions. It helps machines read who the firm is and who practices there. It is hygiene, not a citation lever.

Read the definition→
Technical metric

Time to First Byte (TTFB)

Time to First Byte (TTFB) is the time between a crawler or browser sending a request and receiving the first byte of the response. It measures how fast the server answers, separate from how long the page takes to finish loading. A fast, reliable response is basic hygiene for every crawler, search and AI alike.

Read the definition→
Schema

Speakable specification

SpeakableSpecification is a Schema.org property that points, by CSS selector or XPath, to the parts of a page best suited to be read aloud. Google supports it only as a beta for news publishers. There is no evidence that ChatGPT, Gemini or Perplexity use it when choosing what to cite.

Read the definition→
Schema

Dataset schema

Dataset is the Schema.org type for describing a published dataset. It carries properties such as variableMeasured, distribution, license, temporalCoverage, spatialCoverage and creator. Google Dataset Search uses it to find and list datasets. What earns trust is the data itself being published and checkable.

Read the definition→
Schema

reviewedBy attestation

reviewedBy is a Schema.org property on WebPage that names the person or organization who reviewed the content for accuracy. On health, legal or financial pages it records who checked the claims. It should always match a reviewer named in the visible page.

Read the definition→
Schema

ImageObject schema

ImageObject is the Schema.org type for describing an image with structured metadata such as contentUrl, caption, creditText, dateCreated and license. For visual businesses such as plastic surgery, cosmetic dentistry and home building, it turns a gallery into described, attributable content.

Read the definition→