What is LLM visibility?
LLM visibility is a measurement: how often, and how accurately, a person or brand is named when someone asks an AI assistant a question in their category. It is the closest thing this field has to a rank tracker, with one important difference. There is no fixed results page to count and no published ranking report, so every number comes from sampling answers and adding them up. That makes method everything. This page sets out what is actually being counted, how people measure it, where each method breaks, and how to read a vendor's dashboard number.
What LLM visibility actually measures
Visibility is the measurement layer. The practice of trying to raise it is generative engine optimization, and keeping the two apart saves a lot of confusion: one is the thermometer, the other is the heating.
The thing being measured is presence in a generated answer, and presence has more than one form. Being mentioned in passing is not the same as being cited as a source, and neither is the same as being described correctly.
| What is counted | What it tells you | How it is captured |
|---|---|---|
| Mention | Whether you are named at all in the answer | Search the answer text for your name |
| Citation | Whether a page of yours was used as a source | Read the links the assistant attaches |
| Position | Whether you are named first, mid list, or last | Note where the name falls in the answer |
| Accuracy | Whether the description of you is correct | Read the claim and check it |
| Sentiment | Whether the framing is positive, neutral, or negative | Read it; automated scoring here is crude |
Most commercial tooling reports the first two. Accuracy is the one that matters most for reputation and the one least often measured, because it cannot be counted automatically.
Why AI answer share is hard to pin down
Five things make this genuinely harder than rank tracking, and any honest measurement acknowledges all five.
Answers are non-deterministic: the same prompt run twice can produce different text, different examples, and a different list of names. Personalisation matters, because a logged-in session with saved memory and chat history is not the same test as a clean one. Model versions change without notice, and a number collected before a release is not comparable to one collected after. Retrieval varies by moment, since the live web moves. And there is no public ranking API, so every measurement is a sample of a distribution rather than a reading off a scoreboard.
None of that makes measurement pointless. It makes a single measurement pointless. A method run repeatedly, on a fixed prompt set, over time, produces something you can compare against itself, which is what you actually need.
Brand mentions in LLMs, citations, and accuracy are three different counts
It is common to be highly visible and consistently misdescribed. A firm can be named in most answers about its category while the same answers get its location, its services, or its pricing wrong. It is equally common to be accurate and invisible: everything an assistant says about you is right, and it never brings you up unprompted.
Those two problems have different fixes. Invisibility is a publishing and citation problem. Misdescription is a source problem, and where the wrong description came from a specific page there may also be a removal or correction route. Where it crosses into false statements of fact published by someone else, Google's legal removal request process shows the shape of the formal channel, which is separate from anything a visibility tool measures.
Measurement methods that hold up
- Fix a prompt set of ten to thirty questions a real buyer would type, written once and then left alone.
- Run each prompt several times in fresh sessions, logged out, with memory and chat history off.
- Test more than one assistant, and record which model version answered where the product tells you.
- Record the full answer verbatim, not just whether your name appeared, plus every cited link.
- Score each answer for mention, citation, position, and accuracy, using the same rubric each time.
- Repeat monthly on the same day, and read the trend rather than any single run.
That is a spreadsheet, not a platform. It is also the method every credible tool is automating, which is why knowing it lets you judge one.
Reading a visibility dashboard
A visibility percentage means nothing on its own. Four questions make it legible: which prompts, how many runs per prompt, which assistants and versions, and on what date. A vendor who can answer all four is measuring something. A vendor who cannot is reporting a number with no denominator.
Two further checks are worth making. Ask whether the runs were logged out, because logged-in tests inherit history and inflate familiar names. And ask whether accuracy was assessed at all, or only presence, because a rising visibility line while the descriptions stay wrong is not progress.
Improving visibility, and what it will not fix
The levers are the ordinary ones: publish clear, self-contained, well-sourced material; be consistent about your name and details everywhere they appear; earn mentions on sources the retrieval layer already reads; and keep listings correct. OpenAI's documentation on web search in its models is a useful reminder of why that works: when the retrieval step runs, it is reading the live web, so the live web is the surface you are actually working on.
Those levers are where the published evidence points. The research paper that introduced Generative Engine Optimization found gains for content that added citations, quotations, and sourced statements rather than for content that added keywords, which is a useful corrective to the idea that this is a volume game.
What it will not do is control an answer. There is no bid, no placement, and no method that promises a citation. And visibility work does not remove a true negative statement about you, because that was never a visibility problem in the first place. Sorting which of your problems is which is the useful first move, and it is one of the things a reputation audit records.
Questions about llm visibility: being found in ai answers
What is LLM visibility?
How often, and how accurately, a person or brand is named when someone asks an AI assistant a question in their category. It is a sampled measurement rather than a published ranking.
How do I measure LLM visibility?
Fix a prompt set, run each prompt several times in fresh logged-out sessions across more than one assistant, record answers verbatim with their citations, and repeat on the same schedule each month.
Is LLM visibility the same as generative engine optimization?
No. Visibility is the measurement. Generative engine optimization is the practice of trying to raise it. The two get sold together, which is why the words blur.
Why does my visibility number change between tools?
Because each tool uses a different prompt set, a different number of runs, different assistants, and a different date. Without those four details the numbers are not comparable.