Conversation-level AI visibility metrics score your brand across every turn of a chat instead of only the opening answer. Buyers rarely stop at one question — they narrow, they object, they compare, and only then do they pick.
A brand that wins turn one and vanishes by turn five looks healthy on a standard dashboard and still loses the deal.
This article defines three metrics built for that reality — first-mention turn, turn survival rate, and final-recommendation share — then tries to break each one. Every metric here fails badly when used alone. The failure analysis is the useful part, because most teams adopt exactly one of these and start reporting it as truth.
What are conversation-level AI visibility metrics?
Conversation-level AI visibility metrics measure a brand's presence across an entire multi-turn chat rather than a single response. Instead of asking "were we mentioned?", they ask when you first appeared, whether you stayed, and whether you survived to the final recommendation.
The distinction is structural. Prompt-level metrics — mention rate, citation share, AI share of voice, sentiment — are scalars. Each summarises a single answer. Conversation-level metrics are sequences. They only exist if you preserve turn order.
That sequence carries information the scalar throws away. Two brands can post identical 40% mention rates across a prompt set while one is a fixture in every closing shortlist and the other is a warm-up act, named in broad answers and dropped the moment the buyer says "for a five-person team." Same number. Opposite commercial position.
The three metrics in one line each
| Metric | Definition | Range |
|---|---|---|
| First-mention turn (FMT) | Index of the earliest turn you appear in | 1–N, undefined if absent |
| Turn survival rate (TSR) | Share of turns after your first mention where you remain present | 0.00–1.00 |
| Final-recommendation share (FRS) | % of conversations where you appear in the closing pick | 0–100% |
Why single-prompt tracking overstates and understates at the same time
Single-prompt tracking is a weak predictor of whether you get recommended. Across the 214 brands in our test set, the correlation between turn-1 mention rate and final-recommendation share was r = 0.34 — opening-answer visibility explains roughly 12% of the variance in who actually gets picked.
It errs in both directions:
- Overstates: 41% of brands present in the turn-1 answer were gone by turn 5. They register as visible on a prompt-level dashboard and contribute nothing to the decision.
- Understates: 27% of final-recommendation appearances came from brands absent at turn 1. These are the specialists — strong fit for a narrow constraint, invisible in the generic category answer.
Neither group is measurable with a one-shot prompt. Google's own documentation points at why the chat surface behaves this way: Google Search Central describes a "query fan-out" technique in AI Overviews and AI Mode, issuing multiple related searches across subtopics to build a response. Each follow-up turn re-fans the query against a different constraint set, so the candidate pool is rebuilt, not filtered.
If you are still running a prompt-level programme, our weekly AI search metrics scorecard is the layer this article sits on top of — conversation metrics extend it, they do not replace it.
How we tested this: 1,286 conversations across six engines
We ran the study in-house on MaxAEO's monitoring panel. Method first, so you can judge the numbers:
- Window: 3 March – 30 May 2026
- Engines: ChatGPT, Gemini, Perplexity, Claude, Copilot, Google AI Mode
- Categories: 9 B2B software categories (CRM, help desk, project management, HR, applicant tracking, expense management, e-signature, call tracking, LMS)
- Volume: 1,286 scripted conversations × 5 turns = 6,430 logged responses
- Brands observed: 214 with at least one mention
- Controls: fresh session per conversation, memory and personalization disabled, US English, logged out where the engine permitted it
Every response was labelled by rule first, then reviewed by two annotators, with disagreements resolved by a third. This is first-party panel data, not a survey — it reflects our prompt set and these nine categories, not the entire web.
The five-turn buyer script
Each conversation followed the same refinement ladder, because a comparable metric needs a comparable path:
- Broad category — "best help desk software"
- Constraint narrowing — "for a 12-person support team under $60 per seat"
- Objection — "what are the downsides of your top pick?"
- Head-to-head — "compare your top two for this use case"
- Decision — "which should I pick?"
What counted as a mention, and what counted as a recommendation
A mention meant the brand name appeared in the answer body. Footnote-only links with no in-body name were logged separately as AI citations, since they behave differently. A recommendation meant the brand was presented as an option to choose, not merely described. A negative mention — "that one is overkill for a team your size" — counted as present but was flagged, and that flag turns out to matter enormously below.
Result: where brands actually fall out
| Turn | Prompt type | Share of turn-1 brands still present | Drop vs. prior turn |
|---|---|---|---|
| 1 | Broad category | 100% | — |
| 2 | Constraint narrowing | 84% | −16 pp |
| 3 | Objection | 66% | −18 pp |
| 4 | Head-to-head | 61% | −5 pp |
| 5 | Decision | 59% | −2 pp |
The shape is the finding. The objection turn alone accounts for 44% of all disappearances — more than the constraint turn, and nearly nine times the closing turn. Attrition is front-loaded and it peaks where the model goes looking for criticism.

Metric 1: First-mention turn
First-mention turn (FMT) is the index of the earliest turn in which your brand appears. A brand named in the opening answer scores 1. A brand that only surfaces once the buyer adds constraints scores 3. Lower means earlier — which is not the same as better.
FMT tells you whether you own the category frame or a niche inside it. In our set the median was turn 2, distributed across brand-conversation pairs with any mention: 38% first appeared at turn 1, 24% at turn 2, 19% at turn 3, 12% at turn 4, 7% at turn 5.
How it fails alone:
- It is right-censored, and the censoring flatters you. Conversations where you never appear have no FMT at all. Drop those rows and your average improves as your visibility collapses. Worked example: you appear in 20 of 100 conversations at an average FMT of 1.5. Visibility craters to 5 of 100 — but those five are all turn-1 hits. Your FMT "improves" to 1.0 while reach fell 75%.
- Early is often cheap. Turn-1 answers in broad categories are name-dumps of eight to twelve vendors. Presence there costs little and closes less.
- It moves when your script moves. Rewrite the opening prompt and every FMT in the dataset shifts. It measures your prompt set as much as your brand.
The rule: never report first-mention turn without reach rate — the share of conversations containing at least one mention — on the same row. One number without the other is meaningless.
A low FMT that you cannot explain is usually a category problem, not a copy problem: the model does not file you where buyers are looking. That is a category design fix, and it moves FMT more reliably than another feature page.
Metric 2: Turn survival rate
Turn survival rate (TSR) is the share of turns after your first mention in which you remain present. A brand first named at turn 2 and present at turns 3, 4 and 5 scores 3/3 = 1.00. Present at 3 and 5 but not 4 scores 2/3 = 0.67. Mean TSR across our set was 0.58.
Survival varied by engine, in this order: Claude 0.66, Gemini 0.62, ChatGPT 0.59, Google AI Mode 0.57, Copilot 0.54, Perplexity 0.49. Read that as a property of our prompt set, not a league table — engines differ in how aggressively they prune a list, which is also why single-engine coverage misleads, as we break down in the multi-engine coverage comparison.
How it fails alone:
- It counts being trashed as surviving. In our data, 11% of "surviving" turns were negative-frame mentions — the brand named as the wrong fit. Naive TSR rewards being the cautionary example. Split it into positive-survival and negative-survival or the metric lies to you.
- The denominator collapses at the end. A brand first mentioned at turn 5 has zero remaining turns and scores 1.00 by default. Late entrants look invincible. Only compute TSR when at least three turns remain, or report a curve instead of a ratio.
- Presence is not position. A parenthetical aside at turn 4 scores the same as being the lead recommendation. Pair TSR with rank position or you will confuse mention with momentum.
- Errors look like drops. A refusal or timeout at turn 3 is missing data, not attrition — and treating it as a drop biases the whole curve downward. This is the same failure mode covered in our breakdown of how refusals, timeouts and tool failures bias visibility metrics.
The cleanest fix: stop treating survival as a percentage and treat it as survival analysis. Plot the drop-off curve, right-censor conversations that end early or error out, and report median turns survived. The practical playbook for holding position past turn one is in our guide to getting cited in the 2nd and 3rd reply.
Metric 3: Final-recommendation share
Final-recommendation share (FRS) is the percentage of tracked conversations in which your brand appears in the closing recommendation — the answer that follows an explicit "which should I pick?" It is the closest single number to a win rate in AI-assisted buying. Median brand FRS in our set was 9%; category leaders ranged from 31% to 58%.
It is also the metric most likely to be reported irresponsibly.
How it fails alone:
- It is brutally noisy at small n. Each conversation is one binary observation. Using the Wilson score interval for binomial proportions, an observed 30% FRS carries these 95% confidence bands:
| Conversations tracked (n) | Observed FRS | 95% confidence interval | Half-width |
|---|---|---|---|
| 40 | 30% | 18% – 45% | ±14 pp |
| 100 | 30% | 22% – 40% | ±9 pp |
| 200 | 30% | 24% – 37% | ±6 pp |
| 500 | 30% | 26% – 34% | ±4 pp |
At n = 40, a move from 30% to 38% is inside the noise floor. Most weekly AI visibility reports are built on far fewer than 40 conversations per category and present the delta as a result.
- "Final" is an artifact of your script. End on "which should I pick?" and you measure decisiveness. End on "anyone else I should consider?" and you measure long-tail recall. Two different metrics wearing the same label.
- It hides the path entirely. Two brands both post 30% FRS. One is present in all five turns; the other appears only at turn 5. The first is defensible. The second is one model update from zero.
- It inherits your prompt set's bias. If your scripts skew toward queries you already win, FRS measures your prompt selection. Build the set deliberately — our guide to creating a prompt set for AI brand monitoring covers the sampling logic.
Sample-size floor before you report anything
| Reporting cadence | Minimum conversations per category per engine | What you can honestly claim |
|---|---|---|
| Weekly | 40 | Direction only, never a delta |
| Monthly | 200 | ±6 pp movement is real |
| Quarterly | 500 | ±4 pp; safe for board reporting |
If you cannot hit the weekly floor, report monthly. A number with no interval attached is a decoration.
The three metrics stress-tested side by side
| Metric | What it answers | Why it fails alone | Pair it with |
|---|---|---|---|
| First-mention turn | How early do we enter consideration? | Undefined for never-mentioned chats; improves as reach collapses | Reach rate (% of conversations with ≥1 mention) |
| Turn survival rate | Do we stay once we're in? | Scores negative mentions and late entrants as wins | Sentiment split + a minimum-remaining-turns floor |
| Final-recommendation share | Do we close? | Wide confidence bands at low n; shaped by how the script ends | Wilson interval + the full conversation trace |
Read the right-hand column as the actual deliverable. The triad is not three KPIs on three tiles — it is one vector, and any component reported in isolation is a defensible-looking number with a known failure mode attached.
The conversation trace: store the sequence, derive metrics later
A conversation trace is a per-turn presence string for one brand in one conversation — 1-1-0-0-1 for a five-turn chat. All three metrics above are derivable from it: FMT is the index of the first 1, TSR is the share of 1s after that index, FRS is the final digit averaged across conversations.
This is the practical recommendation of the whole article. Store traces, not aggregates.
Answer engine optimization is roughly two years old as a discipline. Metric definitions will change again — probably twice in the next year. If your LLM brand tracking stores only "FRS = 28% in week 24," you cannot answer next quarter's question. If it stores the trace, plus per-turn sentiment and position, you can recompute any metric anyone invents, retroactively, without re-running a single conversation.
Eight trace patterns and what each one tells you to fix
The traces cluster. These eight shapes covered 81% of our labelled brand-conversation pairs, and each implies a different fix:
| Trace | Name | What it means | First fix |
|---|---|---|---|
1-1-1-1-1 |
Anchor | Category-defining presence | Defend it — monitor drift weekly |
1-1-0-0-0 |
Early fade | Named in broad answers, dropped once constraints appear | Constraint-specific fit content: team size, budget, stack |
1-1-1-1-0 |
Last-turn miss | In the running throughout, absent from the close | Comparison clarity, pricing transparency, third-party proof |
0-0-1-1-1 |
Late climber | Surfaces only when the buyer gets specific | Usually healthy economics — strengthen category association |
0-1-0-1-0 |
Flicker | Unstable entity; the model isn't sure what you are | Entity consolidation: consistent naming, category, schema |
1-0-0-0-1 |
Bookend | Recalled at open and close, missing in the middle | Weak mid-funnel evidence — reviews, comparisons, docs |
0-0-0-0-1 |
Lucky close | Appears only in the final list | Re-run before celebrating; often a sampling artifact |
0-0-0-0-0 |
Absent | No presence at all | A category and source problem, not a copy problem |
Early fade was the single most common non-absent shape, at 23% of traces with at least one mention. Last-turn miss came in at 11% — the smallest population with the highest expected return, because those brands are already in the consideration set and losing on a specific, fixable comparison.
If your executive dashboard demands one composite number, weight the triad rather than averaging it — and calibrate the weights against your own pipeline data instead of inventing them, using the same logic as weighting AI share of voice by revenue and intent. Always show the three components underneath. A composite that hides an Early fade pattern behind a decent average is worse than no composite.
How to instrument conversation-level AI visibility metrics in five steps
- Write conversation scripts, not prompts. Five turns — broad, constraint, objection, comparison, decision. Budget 20–40 scripts per category.
- Run each script in a fresh session per engine, with memory and personalization off, and log the full text of every turn, not just a match flag.
- Label each turn on four axes: presence, position, sentiment, and whether the mention was linked (an AI citation) or unlinked.
- Store the trace string per brand per conversation, alongside engine, date, and script ID. This is your durable asset; the metrics are just views over it.
- Compute the triad weekly with reach rate and confidence intervals attached, and alert on pattern shifts — Anchor degrading to Early fade — rather than on single-point movement.
Step 5 is where most AI search monitoring programmes go wrong. A pattern change is a signal. A four-point weekly wobble on n = 50 is weather. Wiring the triad into an existing reporting rhythm is covered in our AEO dashboard metrics scorecard.
What this costs to run
Rough budget from our own panel, per category per engine, at 40 scripted conversations of 5 turns:
- 200 model responses per category-engine cell per cycle
- Six engines × 9 categories = 10,800 responses per full sweep
- Labelling: rule-based first pass caught 88% of mentions cleanly; the remaining 12% needed human review, at roughly 20 seconds per ambiguous response
The implication: full-sweep conversation tracking is monthly work, not daily. Teams that try to run it daily end up sampling too thin per cell and reporting noise. Pick fewer categories and go deeper.
The objection turn is where most brands leak
Turn 3 destroys more visibility than any other turn in the ladder — 44% of all disappearances in our data. The mechanism is straightforward. When a buyer asks "what are the downsides?", the model stops summarising vendor positioning and goes looking for critical, comparative, third-party material.
If the only critical writing about your product lives on a competitor's comparison page, that is the source that gets summarised. You do not get a rebuttal; you get replaced.
The fix is not more marketing copy. It is publishing honest limitation content yourself — who you are not for, where you are weaker, which use cases warrant a different tool — so there is a first-party answer in the retrieval pool.
Three things that measurably moved the objection turn for brands in our panel, ranked by how often we saw them precede a survival improvement:
- A "who this is not for" section on the pricing or comparison page. Short, specific, named use cases. Models quote it verbatim.
- First-party comparison pages that concede real losses. A page claiming to win every row does not get cited on an objection turn; a page conceding two rows does.
- Dated changelog entries against known criticisms. "Previously limited to X; shipped in March 2026." Gives the model a rebuttal with a date attached.
Does location change the answer?
Yes, and it breaks cross-market comparison. In our panel the same five-turn script run against different city contexts ("best help desk software for a Berlin team") produced different turn-2 survivors than the US-English default — the constraint turn is where the model applies locale, so brands with thin regional evidence fade one turn earlier.
If you sell in multiple markets, compute the triad per market, not blended. A blended TSR hides the market where you fade at turn 2. The monitoring pattern is the same one described in multi-location AI visibility.
Where conversation-level metrics still fall short
Honest limits, because generative engine optimization has enough overclaiming already:
- The scripts are synthetic. Real buyers wander, backtrack, and paste in vendor pages. A five-turn ladder models the median path, not the actual path.
- Logged-out is a cold start. Users with memory and history get different answers. A clean panel measures the no-context case, which is the floor, not the average.
- Turn count is a choice. Five turns is arbitrary. Run seven and survival rates drop mechanically. Never compare TSR across differently-lengthed scripts.
- A trace shows what happened, not why. Pair it with a dated change log and controlled tests, or you will attribute a model update to your content refresh.
- Missing data still masquerades as attrition. Refusals and timeouts must be censored explicitly, not counted as zeros.
- Nine categories is not the web. Everything here is B2B software. Consumer categories with heavier review-site retrieval likely show a different objection-turn shape; we have not tested them.
None of that invalidates the approach. It means the triad belongs next to your prompt-level metrics, not instead of them — and anyone selling a single conversation score with no confidence interval attached is selling a decoration.
Frequently asked questions
How many conversations do I need before these metrics are reliable?
Roughly 200 per category per engine gets you to ±6 percentage points on final-recommendation share at a 30% base rate. At 40 conversations the band is ±14 points — directional at best. If you cannot reach 200, report the interval alongside the number and resist calling weekly movement a trend.
Should conversation-level metrics replace prompt-level ones?
No. Layer them. Prompt-level mention rate and citation share still tell you whether you exist in the retrieval pool at all. Conversation-level metrics tell you what happens to you after that. Fixing a zero mention rate is a sourcing problem; fixing an Early fade trace is a content-fit problem.
Which of the three metrics matters most?
Final-recommendation share is the outcome; the other two are diagnostics that explain it. But FRS on its own cannot tell you whether you are fragile. Read the trace shape first, then decide which number you are trying to move.
How do I track this across engines that behave differently?
Compute the triad per engine and never average raw turn survival rates across them — pruning behaviour differs enough that a blended figure is meaningless. Compare each engine against its own prior period instead.
Does this apply to brand mentions in ChatGPT specifically?
Yes, and ChatGPT is where multi-turn behaviour matters most, since conversational follow-ups are the dominant usage pattern rather than an edge case. Its mid-pack survival rate of 0.59 in our set means roughly two in five brands named early do not make it to the close.
How long before content changes show up in these metrics?
In our panel, changes that affected the objection turn — limitation pages, conceding comparison pages — showed movement in TSR within two to six weeks, gated by recrawl and index refresh rather than by the model. Turn-1 presence moved slower. Do not expect a week-over-week read on either.
Can I compute these metrics from an existing prompt-level tool?
Only if it stores full response text with turn order preserved. Most do not — they store a match flag per prompt, which discards the sequence. If your tool logs only "mentioned: yes/no," the traces are unrecoverable and you have to re-run the conversations.