Does ChatGPT Personalize Recommendations? How It Works

by

·

Does ChatGPT personalize recommendations test matrix comparing clean, logged-in, memory-enabled, and history-conditioned sessions

By maxaeo

Yes. ChatGPT can personalize recommendations when it can use information from your current conversation, saved memories, referenced chat history, or custom instructions. Personalization may change which options appear, their order, and the explanation. However, different answers alone do not prove personalization because ChatGPT recommendations also vary naturally.

The important question is not simply whether two users received different lists. It is whether a known personal preference caused a repeatable, preference-aligned change beyond the variation seen in a neutral control.

This guide explains what ChatGPT can use, how to limit personalization, and how to test its effect with the maxaeo Personalization Attribution Ladder and a variance-adjusted measurement method.

Methodology note: Product-behavior claims below are tied to OpenAI’s published documentation. The testing framework is an original maxaeo methodology. The worked dataset is explicitly simulated and is not presented as a benchmark of current ChatGPT recommendations.

ChatGPT recommendation personalization at a glance

Potential input Can it affect recommendations? What it may change
Current prompt Yes Product fit, filters, ranking, rationale
Earlier messages in the same chat Yes Selection, order, exclusions, explanation
Saved memories Yes, when enabled and available Recommendations across new conversations
Referenced chat history Yes, when enabled and available Recommendations based on prior conversations
Custom instructions Yes Persistent preferences, format, priorities
Files or connected services Potentially, when shared or authorized Recommendations based on supplied account data
Login status alone Not proven Available models, features, settings, or search access
Identical prompts with no personal context Answers may still differ Random selection, ranking, wording, or citations

What does “personalized” mean in a ChatGPT recommendation?

ChatGPT personalization is a change in an answer that follows from user-specific information. It should be distinguished from ordinary contextualization and general output variation.

There are three useful categories:

  1. In-thread contextualization: ChatGPT uses information stated in the current prompt or earlier in the same conversation.
  2. Persistent personalization: A new conversation uses saved memory, referenced past chats, or custom instructions.
  3. Surface variation: Answers differ because of model access, web results, location, language, account experiments, or normal model variability—not because of a personal preference.

Measure personalization across three output dimensions:

  • Selection: Did a brand enter or leave the shortlist?
  • Position: Did a brand move higher or lower?
  • Framing: Did the explanation change to reflect the user’s requirements?

For example, a neutral answer might rank Jira, Asana, and Linear in that order. A profile containing “I prefer lightweight tools and minimal administration” might rank Linear first and omit Jira. That repeated, preference-aligned change is stronger evidence than a single swap between two otherwise similar lists.

What information can ChatGPT use to personalize recommendations?

Current prompts and conversation context

Anything a user states in the active conversation can shape the answer. Budget, location, team size, integrations, technical ability, accessibility needs, and previous product experience can all change what “best” means.

This is personalization in an everyday sense, but it does not require a saved profile. A logged-out user can receive a tailored answer by providing detailed criteria in the prompt.

Saved memories

Saved memories allow ChatGPT to retain selected details for future conversations. OpenAI’s Memory FAQ distinguishes saved memories from information referenced from past chat history.

A memory such as “I run a bootstrapped company and avoid annual contracts” could affect a later software recommendation even when that preference is absent from the new prompt.

Deleting the original chat may not delete a separately saved memory. OpenAI says that fully removing remembered information may require deleting both the saved memory and the chat where the information was shared.

Referenced chat history

When the relevant feature is available and enabled, ChatGPT may use information from previous conversations without storing every detail as an explicit saved memory.

This mechanism is harder to test than same-thread context because the exact prior detail selected for use may not be visible in the new conversation. A credible experiment therefore needs repeated runs and a scripted set of precursor chats.

Custom instructions

Custom instructions provide standing guidance about the user or the desired response style. OpenAI’s Custom Instructions documentation confirms that these instructions are considered when ChatGPT responds.

A custom instruction such as “Prioritize tools with EU data residency” can influence recommendations even in a new chat. Custom instructions must therefore be disabled in any supposedly clean control.

Files and connected services

ChatGPT can also use information supplied through uploaded files or authorized connected services when those features are available. This is not unrestricted access to the user’s private data: the relevant information must be provided, connected, or otherwise made available through the product.

A user should not assume ChatGPT automatically sees general browser history, private documents, email, purchases, or activity in unrelated services.

For a deeper treatment of persistent context, see maxaeo’s analysis of how memory and logged-in profiles change AI answers.

Does being logged in automatically change recommendations?

No automatic personalization effect should be assumed from login alone. Logging in can make memory, custom instructions, account history, different models, and other account-level features available, but it does not reveal which factor influenced an answer.

A logged-out and logged-in comparison may also be confounded by:

  • Different model access
  • Different product plans
  • Search or browsing availability
  • Regional feature availability
  • Interface language and locale
  • Account-level experiments
  • Previously configured instructions

The defensible conclusion is that login can enable personalization mechanisms; it is not proof that a recommendation used them.

To isolate saved memory, compare two states within the same account: one with memory disabled and one with a standardized memory enabled. Keep the model, search state, locale, prompt, and custom instructions constant.

Does Temporary Chat remove personalization?

OpenAI’s Temporary Chat FAQ says Temporary Chat starts with a blank slate, does not access or create memories, does not appear in chat history, and is not used to improve OpenAI’s models. OpenAI may retain a copy for up to 30 days for safety purposes.

Temporary Chat is useful for control testing, but it is not automatically free of every influence. OpenAI notes that enabled custom instructions may still apply. The current prompt and messages inside the Temporary Chat also remain part of that conversation’s context.

For the cleanest practical session:

  1. Disable custom instructions.
  2. Start a Temporary Chat or a new logged-out session.
  3. Avoid files, connectors, and prior turns.
  4. Record the visible model, locale, and search state.
  5. Use the exact same prompt as the comparison condition.

How can you tell whether your recommendation was personalized?

Use this attribution sequence:

  1. Check the current prompt. Did it contain budget, location, preferences, exclusions, or prior product experience?
  2. Review earlier messages. Was the relevant preference stated earlier in the same conversation?
  3. Inspect saved memory. Ask ChatGPT what it remembers, then verify the memory settings directly.
  4. Check referenced chat history and custom instructions.
  5. Check files and connected services. Was account-specific information supplied?
  6. Repeat the prompt in a clean control. Disable the relevant context and compare several responses.
  7. Look for semantic alignment. Did the changed recommendation explicitly match the personal preference?

Asking ChatGPT “Why did you recommend this?” can identify the criteria used in the answer, but its explanation is not a reliable audit log of hidden system behavior. Treat the response as a clue, then verify it with observable settings and a controlled comparison.

Why can identical prompts produce different recommendations?

ChatGPT can generate different answers even when no personal information changes. Brand inclusion, order, wording, and citations may vary because the system is probabilistic and because models, search results, product data, or system configurations can change.

Common false positives include:

  • Two brands exchange positions while the shortlist remains nearly identical.
  • A product disappears in only one response.
  • The rationale changes without reflecting any personal preference.
  • Search-enabled answers cite different newly published pages.
  • Every test condition changes during the same collection window.
  • Logged-out and logged-in sessions use different models or features.

One screenshot measures difference, not cause. Run repeated neutral controls alongside the personalized condition. If both move by similar amounts, the result is more likely ordinary recommendation volatility.

See maxaeo’s study of AI answer volatility for the broader measurement problem.

How to test whether ChatGPT personalizes recommendations

The maxaeo Personalization Attribution Ladder uses four core conditions. Each step adds or isolates one source of context.

Condition Configuration Primary comparison
A. Clean control Logged out or Temporary Chat; no prior turns; custom instructions off Establishes a generic baseline
B. Logged-in neutral New chat; saved memory, referenced history, and custom instructions off Compare with A to detect account or surface differences
C. Saved-memory only Same account and settings as B; standardized saved memory enabled; referenced history and custom instructions off Compare with B to isolate saved memory
D. Same-thread context Same settings as B; standardized preference stated in an earlier turn Compare with B to isolate active conversation context

These comparisons have different meanings:

  • B minus A measures an account or product-surface difference. It does not isolate login as a causal personalization signal.
  • C minus B estimates the saved-memory effect.
  • D minus B estimates the effect of visible conversational context.

Optional extensions can isolate other mechanisms:

  • Custom-instruction condition: Enable only a standardized custom instruction and compare it with B.
  • Referenced-history condition: Create scripted precursor chats, enable only referenced chat history, and compare new-chat answers with B.
  • Connected-data condition: Provide the same authorized file or service data in every replicate and compare it with B.
Does ChatGPT personalize recommendations test matrix comparing clean, logged-in, memory-enabled, and history-conditioned sessions

What variables must remain constant?

Lock every variable except the one being tested:

  • Exact prompt wording and punctuation
  • Requested number of recommendations
  • Model and product plan
  • Search or browsing state
  • Country, language, and interface locale
  • Custom instructions
  • Saved-memory and referenced-history settings
  • Files and connected services
  • Previous turns in the test conversation
  • Collection window
  • Condition order
  • Requested citations and output format

Use a fresh conversation for every replicate. Rotate the order of conditions—for example, A-B-C-D, then B-D-A-C—so a model-side change during collection does not affect only one condition.

For the memory condition, use a dedicated test account where possible. Add the standardized memory before collection, verify it, and check it again after the final run. Do not mention the target brand in the memory profile.

What prompt should the test use?

Choose a prompt specific enough to create a stable comparison but neutral enough that it does not contain the preference being tested.

A suitable neutral prompt is:

Recommend five project-management platforms for a 25-person B2B SaaS product team. Rank them and give one concise reason for each recommendation.

Use the following standardized profile only in the conditioned cells:

The team values fast setup, lightweight product workflows, and strong Slack integration. It wants to minimize administration and avoid tools that require heavy configuration.

Do not name a target brand. Naming it would test brand priming, not whether ChatGPT independently matches products to the profile.

The requested shortlist length also matters. A brand has fewer opportunities to appear in a three-product answer than in a ten-product answer. Keep the number fixed and interpret visibility within that ceiling. Maxaeo’s analysis of how many brands an AI answer recommends explains this constraint in detail.

How many responses should you collect?

Use 20 responses per condition as a practical audit floor, not as a universal statistical threshold. Four conditions produce 80 records.

Increase the sample to 30–50 responses per condition when:

  • Neutral shortlists change frequently
  • The category contains many similar products
  • Inclusion differences are small
  • Web search introduces changing evidence
  • The result will support a material business decision

Repeat the experiment on at least one additional day. A pattern that appears once and vanishes in the next collection window may reflect model drift rather than durable personalization.

Where statistical inference is required, report confidence intervals or use an appropriate test for the metric. Do not convert 20 UI runs into a universal claim about all users.

What data should be captured?

Store each response as one structured record.

Field Example
Condition Saved-memory only
Run ID M-014
Timestamp 2026-07-21T10:15:00+08:00
Model Exact label displayed in the interface
Prompt version PM-TEST-01
Search state On
Brand list Linear, Asana, Notion, ClickUp, Jira
Rank positions Linear 1; Asana 2; Notion 3
Top recommendation Linear
Recommendation reasons “Lightweight product workflow”
Profile criterion referenced Low administration
Citations Domains and URLs shown
Raw response Complete answer text

Normalize aliases before counting. “monday.com” and “Monday” should map to one entity, while distinct products with similar names must remain separate.

Preserve raw answers next to coded data. This allows another analyst to audit subjective judgments about product names, recommendation reasons, and preference alignment.

Which metrics identify a personalization effect?

Shortlist overlap

Use Jaccard similarity:

Shortlist overlap = brands in both lists ÷ brands in either list

If the control list is {Jira, Asana, Linear, ClickUp, Monday} and the conditioned list is {Linear, Notion, Asana, Height, ClickUp}, the overlap is 3 ÷ 7, or 0.43.

A lower score means the two shortlists differ more.

Inclusion lift

Measure how often each brand appears:

Inclusion lift = conditioned inclusion rate − neutral inclusion rate

If a brand appears in 19 of 20 conditioned answers and 14 of 20 neutral answers, its inclusion lift is +25 percentage points.

Report both raw counts and percentages.

Rank movement

For brands present in both conditions, calculate the change in median rank. Track missing brands separately rather than assigning them an arbitrary rank without disclosure.

Top-recommendation switch rate

Pair each conditioned run with the nearest neutral run in the same collection block. Count how often the first recommendation changes.

Rationale alignment lift

Code whether the recommendation explanation reflects a criterion in the standardized profile.

Rationale alignment lift = conditioned alignment rate − neutral alignment rate

Two reviewers should independently code ambiguous rationales when the result will inform a significant decision. Report the coding rules and reviewer agreement.

Citation overlap

For search-enabled answers, compare the cited domains and pages. A user profile may change the evidence selected even when the product shortlist remains stable.

The maxaeo Personalization-Adjusted Distance metric

Raw shortlist difference overstates personalization because neutral answers also vary. The Personalization-Adjusted Distance (PAD) subtracts the normal control drift from the conditioned difference.

Calculate it in three steps:

  1. Convert shortlist overlap into distance:
    Distance = 1 − Jaccard similarity
  2. Calculate the median distance between repeated neutral-control answers collected in the same window.
  3. Subtract that neutral distance from the median distance between the conditioned and time-matched neutral answers.

PAD = median conditioned distance − median neutral-control distance

Interpretation:

  • PAD near zero: The conditioned answer changes no more than the neutral control.
  • Positive PAD: The condition creates additional shortlist movement beyond normal volatility.
  • Negative PAD: Neutral answers varied more than the conditioned answers.
  • Positive PAD plus rationale alignment lift: Stronger evidence that the movement reflects the supplied preference.

PAD has no universal “good” threshold. Category size, requested shortlist length, and baseline volatility all affect its scale. Publish the components rather than reporting the final score alone.

This control-adjusted approach follows the same logic as an AEO holdout: estimate what would have happened without the treatment before attributing the change. See maxaeo’s guide to holdout testing for AEO.

Worked example: what would personalization look like?

The following table is a simulated 80-response audit fixture, with 20 responses per condition. It demonstrates the calculations; it is not evidence of current behavior for these brands.

The neutral-control median overlap between repeated clean answers is 0.82, producing a baseline distance of 0.18.

Condition Median overlap with matched neutral Median distance PAD Top recommendation changed Rationale alignment
Clean control 0.82 within control 0.18 Baseline Baseline 10%
Logged-in neutral 0.80 0.20 +0.02 3/20 10%
Saved-memory only 0.55 0.45 +0.27 11/20 70%
Same-thread context 0.60 0.40 +0.22 10/20 85%

The logged-in neutral result is close to ordinary control variation. The saved-memory and same-thread conditions show larger adjusted distances and much higher rationale alignment.

In the simulated records, the saved-memory condition also changes brand inclusion:

Brand Neutral inclusion Saved-memory inclusion Inclusion lift
Linear 14/20 19/20 +25 percentage points
Jira 17/20 5/20 −60 percentage points

This pattern would be moderate-to-strong evidence because the change is repeated, exceeds neutral drift, and aligns with the injected preference for lightweight workflows. It would still apply only to the tested model, category, profile, and collection window.

What counts as convincing evidence?

Evidence level Observable pattern Defensible conclusion
Weak One answer changes Difference observed; cause unknown
Suggestive Repeated changes but PAD is small or rationales are unaligned Personalization is possible
Moderate Positive PAD, material inclusion or rank changes, aligned rationales The tested context likely influenced recommendations
Strong Replicated positive PAD across collection windows, low neutral drift, aligned rationales, consistent citation changes The tested context reliably influenced this query
Inconclusive High volatility in every condition or uncontrolled settings No attribution should be made

Publish the prompt, profile, settings, collection dates, run count, raw counts, coding rules, and limitations. Without those details, readers cannot distinguish a controlled result from an anecdote.

How can users reduce or reset personalization?

Users who want a more neutral recommendation can take the following steps:

  1. Start a new conversation to remove the active thread context.
  2. Review and delete relevant saved memories.
  3. Disable saved-memory and referenced-history settings where available.
  4. Disable custom instructions.
  5. Use Temporary Chat.
  6. Remove uploaded files and disconnect unnecessary services.
  7. Ask a neutral prompt that does not reveal prior preferences.
  8. Compare several answers rather than treating one response as definitive.

A useful neutral prompt is:

Recommend five options based only on the criteria in this message. Do not use previous conversations, remembered preferences, or custom profile information. State the criteria used and identify any missing information.

That instruction can improve transparency, but it is not a substitute for changing the product settings. A model’s statement that it ignored prior context is not an independent verification.

What does personalization mean for consumers?

Personalization can make recommendations more relevant, but it can also preserve outdated assumptions.

Check whether ChatGPT is relying on:

  • An old budget or location
  • A former employer or team size
  • A product preference that has changed
  • A previous technical skill level
  • An unverified constraint inferred from earlier chats

Ask for alternatives outside the inferred profile:

Give me three recommendations that fit my stated preferences and three credible alternatives that challenge them. Explain the trade-offs.

This reduces the risk of receiving a narrow shortlist based on stale or overly restrictive context. Recommendations should still be checked against current pricing, availability, documentation, security requirements, and independent evidence.

What does personalization mean for AI share of voice?

Personalization turns AI share of voice from one universal percentage into a distribution across buyer contexts.

A brand may appear frequently in generic recommendations but disappear when the user specifies:

  • Enterprise security requirements
  • A startup budget
  • A preferred integration
  • A particular country
  • Low administrative overhead
  • On-premises deployment
  • Accessibility requirements

Measure three layers:

  • Generic share of voice: Visibility without an audience profile
  • Segment-adjusted share of voice: Visibility for defined buyer profiles
  • Preference resilience: How often the brand remains recommended as relevant constraints accumulate

Every tracked prompt should retain its audience profile. Averaging unrelated profiles together can hide the segments where a brand is genuinely strong or consistently excluded.

How should brands respond to personalized recommendation gaps?

Brands cannot control a user’s memory or guarantee inclusion. They can improve the evidence answer engines use to evaluate fit.

Start with the conditions that cause exclusion:

  • Publish substantive pages for relevant industries, team sizes, and maturity levels.
  • Document integrations with setup requirements and limitations.
  • State security certifications, hosting regions, and deployment options precisely.
  • Make pricing structure and contract requirements clear.
  • Explain genuine product trade-offs instead of claiming to suit every buyer.
  • Publish comparison pages supported by verifiable capabilities.
  • Add attributable case studies with customer context and measurable outcomes.
  • Keep product facts consistent across documentation, owned pages, and credible third-party sources.

Avoid mass-producing thin persona pages. One detailed page with requirements, limitations, screenshots, documentation, and a relevant case study provides more evidence than dozens of near-duplicate landing pages.

How can teams operationalize the test?

A monitoring workflow should preserve more than brand mentions. Whether a team uses maxaeo or another platform, it should retain:

  • Prompt and audience profile
  • Answer engine and model
  • Memory, history, and custom-instruction state
  • Brand inclusion and rank
  • Recommendation rationale
  • Competitor co-occurrence
  • Citations
  • Raw response
  • Collection date
  • Neutral-control volatility

Run the test as a recurring panel rather than a collection of unrelated prompts. Keep a stable control group, add profiles as explicit experimental conditions, and version every prompt.

When evaluating platforms for this workflow, compare their handling of repeat runs, raw-answer retention, citation capture, model coverage, and prompt segmentation. Maxaeo’s tested guide to AI search visibility tools provides a practical evaluation framework.

Limitations of a controlled ChatGPT test

A controlled interface test can show that outputs change reliably after an observable context change. It cannot reveal internal model weights, undisclosed experiments, or every account-level signal.

Every report should disclose that:

  • ChatGPT outputs are not deterministic.
  • Behavior may differ by model, plan, country, language, and date.
  • Search-enabled answers depend on changing web content and retrieval results.
  • Temporary Chat may still follow enabled custom instructions.
  • A single product category cannot represent all recommendation queries.
  • UI labels and feature availability may change.
  • Twenty runs per cell provide a directional audit, not a universal rate.
  • Simulated examples must not be reported as observed market data.
  • Recommendation inclusion does not prove endorsement, accuracy, or commercial intent.

Frequently asked questions

Does ChatGPT personalize product recommendations?

Yes. ChatGPT can tailor product recommendations using the current conversation, saved memories, referenced chat history, custom instructions, and information supplied through files or connected services. Personalization may change product inclusion, ranking, rationale, or citations.

A different answer alone is not proof. The change should be repeated, larger than neutral variation, and aligned with a known preference.

Does logging in change the brands ChatGPT recommends?

Logging in can make personalization features and account-specific settings available, but login alone does not prove that personal information influenced an answer. Model access, search availability, plan differences, locale, and experiments may also differ.

Compare a logged-in neutral state with the relevant feature disabled against the same account with one personalization source enabled.

Does Temporary Chat remove all personalization?

Temporary Chat does not access or create memories, according to OpenAI. However, enabled custom instructions may still apply, and messages inside the active Temporary Chat still provide context.

Disable custom instructions and avoid prior turns, files, or connectors when creating a clean control.

Can ChatGPT see my browser or purchase history?

ChatGPT does not inherently have unrestricted access to general browser history, purchases, email, private documents, or unrelated accounts. It can use information that you share, upload, save as memory, or authorize through an available connected service.

Review connected services and data settings rather than assuming that every tailored answer came from hidden personal data.

Can I turn off personalized recommendations?

You can reduce personalization by starting a new conversation, deleting relevant saved memories, disabling memory and referenced chat history, turning off custom instructions, disconnecting unnecessary services, and using Temporary Chat.

The names and availability of these controls may vary by account and product version.

How many responses are needed to test personalization?

Twenty responses per condition are a practical starting point for a directional audit. Increase the sample when neutral recommendations are volatile, the product category is crowded, or the observed differences are small.

Always report raw counts, collection dates, settings, and the neutral-control variation.

Can a company influence personalized ChatGPT recommendations?

A company cannot control a user’s profile or guarantee that ChatGPT will recommend its product. It can publish accurate, retrievable evidence about product capabilities, integrations, pricing, security, use cases, and trade-offs.

Measure visibility by buyer profile to identify the contexts in which the product is a strong match or is consistently excluded.

Final answer

Does ChatGPT personalize recommendations? Yes—when relevant context is available through the current conversation, saved memory, referenced chat history, custom instructions, or authorized data. The visible effect may be a different shortlist, a new ranking, or a rationale tailored to the user’s preferences.

Do not treat a changed answer as proof. Compare one personalization source at a time against a repeated neutral control, then measure shortlist distance, inclusion lift, rank movement, rationale alignment, and citation overlap.

The strongest evidence is a replicated change that exceeds ordinary answer volatility and makes semantic sense given the preference introduced. That standard separates genuine personalization from random variation—and produces AI visibility findings that users, marketers, and decision-makers can audit.


Written by

Founder of MaxAEO. Helping brands get found in AI search across ChatGPT, Perplexity, Google AI Overviews, and more.

Run a free AI visibility audit →