
{"id":1852,"date":"2026-08-06T08:06:47","date_gmt":"2026-08-06T08:06:47","guid":{"rendered":"https:\/\/maxaeo.ai\/blog\/?p=1852"},"modified":"2026-08-07T03:44:40","modified_gmt":"2026-08-07T03:44:40","slug":"how-to-run-a-same-prompt-benchmark-across-chatgpt-perplexity-gemini-and-grok-yourself-no-tool-required","status":"publish","type":"post","link":"https:\/\/maxaeo.ai\/blog\/how-to-run-a-same-prompt-benchmark-across-chatgpt-perplexity-gemini-and-grok-yourself-no-tool-required\/","title":{"rendered":"How to Run a Same-Prompt Benchmark Across ChatGPT, Perplexity, Gemini and Grok Yourself (No Tool Required)"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Someone in the other thread asked whether anyone had compared the same prompt set across platforms. That is the useful question, and it never really got answered.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I have not completed a clean four-platform run that I can honestly present as benchmark results. What I can share is the manual protocol I built to make such a run comparable. It needs no monitoring product or browser extension, and it avoids the biggest trap in these comparisons: turning unlike answers into one precise-looking score.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The unit is simple:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><code>one prompt \u00d7 one platform \u00d7 one pass<\/code><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Keep every observation at that level until collection is finished.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">1. Build a prompt pack that can survive comparison<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Do not start by copying 50 loosely related prompts from a keyword tool. A smaller pack with explicit intent is easier to rerun and audit.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I would start with 18 prompts in three groups:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Category discovery: six prompts such as \u201cWhat are the best tools for [job]?\u201d<\/li>\n\n\n\n<li>Problem diagnosis: six prompts such as \u201cWhy is [type of company] missing from AI recommendations?\u201d<\/li>\n\n\n\n<li>Direct comparison: six prompts such as \u201cWhich is better for [use case], Brand A or Brand B?\u201d<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Before opening any platform, give every prompt an ID and freeze the exact wording. If one prompt turns out to be ambiguous, do not quietly edit it halfway through. Create a new ID and rerun the revised version everywhere.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Also freeze four conditions in a small run header:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>language;<\/li>\n\n\n\n<li>region or account location;<\/li>\n\n\n\n<li>signed-in versus signed-out state;<\/li>\n\n\n\n<li>whether follow-up questions are allowed.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Those details matter because a prompt is not the only input. Account context and follow-ups can change the answer shape enough to ruin a direct comparison.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">2. Run the four platforms in a controlled sequence<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Use ChatGPT, Perplexity, Gemini and Grok in the same 48-hour window. Run the complete prompt pack once on every platform, wait several hours, then repeat the full pack for pass two.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The sequence should be interleaved by prompt, not completed platform by platform over several days:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Run P01 on all four platforms.<\/li>\n\n\n\n<li>Run P02 on all four platforms.<\/li>\n\n\n\n<li>Continue through P18.<\/li>\n\n\n\n<li>Repeat the same order for pass two.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">This reduces the chance that a news event or index update affects only one platform&#8217;s batch. It also makes drift visible: if P07 changes between passes on Gemini but remains stable elsewhere, that disagreement stays attached to P07 instead of disappearing inside an average.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For 18 prompts, four platforms and two passes, the design produces 144 possible observations. That is the planned sample size, not a claim that those observations have already been collected.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">3. Capture the answer before interpreting it<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For the first pass through the sheet, record only observable fields:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Mentioned: whether the target brand appears, <code>yes<\/code> or <code>no<\/code>;<\/li>\n\n\n\n<li>Cited domains: every visible source domain, or <code>none<\/code>;<\/li>\n\n\n\n<li>Answer saved: a screenshot or pasted answer reference;<\/li>\n\n\n\n<li>Collection time: timestamp for that observation.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Do not score sentiment, rank position or source quality while collecting. Those judgments are useful later, but mixing collection and interpretation creates a moving standard. The tenth answer gets judged differently from the first because you have already seen a pattern.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">There are two platform-specific edge cases worth recording explicitly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">First, if a platform shows no citations, write <code>none<\/code>; do not infer sources from phrases in the answer. Second, if a platform cites the same domain multiple times, retain the domain once for source-mix analysis but keep the raw answer so citation frequency can be reviewed separately.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">4. Turn cited domains into a source map<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">After collection, classify each visible domain into one of five broad buckets:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>competitor website;<\/li>\n\n\n\n<li>directory or software marketplace;<\/li>\n\n\n\n<li>community;<\/li>\n\n\n\n<li>media or editorial site;<\/li>\n\n\n\n<li>video platform.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Keep the buckets broad on purpose. The first question is not whether a source is a review site or an affiliate review site. It is whether one answer engine repeatedly depends on communities while another prefers vendor pages or editorial lists.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For each prompt, build a four-column source map:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Prompt<\/th><th>ChatGPT<\/th><th>Perplexity<\/th><th>Gemini<\/th><th>Grok<\/th><\/tr><\/thead><tbody><tr><td>P01<\/td><td>Mention + source buckets<\/td><td>Mention + source buckets<\/td><td>Mention + source buckets<\/td><td>Mention + source buckets<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">That row supports concrete questions:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Which platforms mention the brand in both passes?<\/li>\n\n\n\n<li>Which platforms omit it in both passes?<\/li>\n\n\n\n<li>Which source bucket dominates each platform?<\/li>\n\n\n\n<li>Which domains recur for the same prompt?<\/li>\n\n\n\n<li>Where does the platform expose no visible source at all?<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This is already more actionable than a vendor ranking. If community sources recur only on Grok, the publishing opportunity is different from a case where editorial lists recur on Perplexity.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">5. Separate stable differences from answer drift<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">One answer is an observation, not a finding. Use pass two as a minimum stability check.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Mark a pattern <code>stable<\/code> only when it repeats under the same prompt and platform. Examples:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>the brand appears in both ChatGPT passes and neither Gemini pass;<\/li>\n\n\n\n<li>community sources dominate both Grok answers for P04;<\/li>\n\n\n\n<li>the same editorial domain is cited in both Perplexity passes for P11.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Mark it <code>unstable<\/code> when the two passes disagree. Do not break the tie by choosing the answer that supports your preferred story.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A compact decision table helps:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Pass 1<\/th><th>Pass 2<\/th><th>Label<\/th><th>What you may say<\/th><\/tr><\/thead><tbody><tr><td>Mentioned<\/td><td>Mentioned<\/td><td>Stable mention<\/td><td>Repeatedly mentioned in this test window<\/td><\/tr><tr><td>Not mentioned<\/td><td>Not mentioned<\/td><td>Stable absence<\/td><td>Repeatedly absent in this test window<\/td><\/tr><tr><td>Mentioned<\/td><td>Not mentioned<\/td><td>Unstable<\/td><td>Answer drift; rerun before acting<\/td><\/tr><tr><td>Source bucket A<\/td><td>Source bucket B<\/td><td>Unstable source mix<\/td><td>Retrieval pattern is not yet repeatable<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Even two passes are only a screening rule, not proof of a permanent platform behavior. The defensible conclusion is limited to this prompt pack, account context and 48-hour window.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">6. Keep platform results separate<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Do not add ChatGPT, Perplexity, Gemini and Grok into one visibility score. They expose different numbers of sources, use different retrieval systems and return different answer formats. A blended score hides the exact difference the benchmark is meant to reveal.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Report four platform-level outputs instead:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>stable mention rate by intent group;<\/li>\n\n\n\n<li>stable absence rate by intent group;<\/li>\n\n\n\n<li>recurring source buckets;<\/li>\n\n\n\n<li>recurring domains for the same prompt.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Only compare like with like. A platform with no visible citations should not receive a lower \u201csource quality\u201d score than one that exposes citations; the fields are not comparable.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Copyable collection sheet<\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Prompt ID<\/th><th>Intent<\/th><th>Prompt<\/th><th>Platform<\/th><th>Pass<\/th><th>Mentioned<\/th><th>Cited domains<\/th><th>Source buckets<\/th><th>Timestamp<\/th><th>Stability<\/th><\/tr><\/thead><tbody><tr><td>P01<\/td><td>Discovery<\/td><td>What are the best tools for [job]?<\/td><td>ChatGPT<\/td><td>1<\/td><td>Yes\/No<\/td><td>domain.com or none<\/td><td>Community<\/td><td>YYYY-MM-DD HH:MM<\/td><td>Pending<\/td><\/tr><tr><td>P01<\/td><td>Discovery<\/td><td>What are the best tools for [job]?<\/td><td>ChatGPT<\/td><td>2<\/td><td>Yes\/No<\/td><td>domain.com or none<\/td><td>Community<\/td><td>YYYY-MM-DD HH:MM<\/td><td>Stable\/Unstable<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">If anyone here runs this exact four-platform, two-pass format, I would be interested in comparing one thing: do your source buckets remain stable for the same prompt, or does the source mix move as much as the brand mentions do?<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A reproducible manual protocol for comparing the same prompts across four AI platforms without inventing a blended visibility score.<\/p>\n","protected":false},"author":1,"featured_media":1765,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-1852","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/posts\/1852","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/comments?post=1852"}],"version-history":[{"count":1,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/posts\/1852\/revisions"}],"predecessor-version":[{"id":1853,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/posts\/1852\/revisions\/1853"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/media\/1765"}],"wp:attachment":[{"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/media?parent=1852"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/categories?post=1852"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/tags?post=1852"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}