simplex

Blogs · 13 Aug 2026 · 14 min

We asked ChatGPT, Perplexity, Gemini, and Google AI Mode the same 12 buyer questions

An August 2026 snapshot of twelve buyer prompts across ChatGPT, Perplexity, Gemini, and Google AI Mode, plus a method you can rerun next week.

Josip Tomić

Partner

Partner at Simplex. He leads the AI agent work on YourGPT — the knowledge, the functions, and the site the agent sits on — and writes the studio reviews.

In August 2026, public answers from ChatGPT, Perplexity, Gemini, and Google AI Mode agree on category names more than on a winner. They miss ninety-day cost, dirty knowledge, and who owns a failure. This is the method for rerunning twelve buyer prompts, and the gaps the models still leave.

A buyer does not type “GEO strategy.” They type “n8n vs Zapier,” or “best Shopify agent that can refund,” then paste the shortlist into a second model.

G2’s March 2026 survey of 1,076 B2B software buyers found that 51% now start research in an AI chatbot more often than in Google, up from 29% in 2025. 71% use a chatbot at some point. 61% still use Google alongside it.

We did not pretend to sit inside four logged-in products. We documented what the public record currently says those products recommend.

What is this snapshot, and what is it not?

It is not a bake-off of four chat windows. OpenAI does not give this desk a lab seat. AI Mode is signed-in Search. Gemini changes with account, region, and grounding.

What we can reconstruct, in August 2026, from pages a reader can open without a login:

  • Official comparison and pricing pages (Zapier, n8n, Gorgias, Fin, Google Cloud, Search Central)
  • Click and citation studies (Ahrefs, Pew, SparkToro, Semrush, Princeton GEO)
  • Public Reddit threads where buyers name tools
  • Perplexity-style answer pages that already cite those primaries

The method is the product. The named winners will move. The gaps will not.

The jobs behind the answers are SEO, AEO, and GEO. The citation craft is how to get cited. The journey is buyers start in ChatGPT.

How did we construct the twelve prompts?

Questions this studio hears, checked against Semrush’s 200,000-AIO study: how / what / is, mostly informational, often under 1,000 monthly searches.

The twelve:

  1. What is the best no-code AI agent builder right now?
  2. What is the best AI agent for a Shopify store?
  3. n8n vs Zapier vs Make — which should a marketing team pick?
  4. What should a marketing website cost in 2026?
  5. What is the difference between SEO, AEO, and GEO?
  6. What is the difference between a chatbot and an AI agent?
  7. How do I get my brand cited in ChatGPT and Google AI Overviews?
  8. Next.js vs Webflow vs Framer for a brand site in 2026?
  9. Which support agent can look up an order, issue a refund, and close the ticket?
  10. Should I add an llms.txt file to my website?
  11. What is the best AI agent builder for real estate?
  12. How do I choose a web design studio when AI can generate a site for $20?

How would a reader rerun this in one afternoon?

Do not trust our table in November. Rerun it.

  1. Open four surfaces: ChatGPT with search on, Perplexity, Gemini with Search grounding, Google AI Mode (or the Overview if Mode is off).
  2. Paste one prompt. Do not add “be unbiased.” Buyers do not.
  3. Record named vendors, first three cited domains, any number, and the first caveat — or its absence.
  4. Rerun after 48 hours. Ahrefs has shown Overviews rewrite often without changing their mind.
  5. Save the sheet with a date. That file is the audit.

G2 interviewed 39 B2B marketers for the same report. The ones who were not guessing already type the buyer query every week.

What do the four systems actually do differently?

They do not share a brain. They share a habit of sounding sure.

ChatGPT, with search on, writes a confident shortlist and leans on encyclopedic and review sources. G2’s 2026 work is blunt: review-site citations are the trust signal buyers say they want.

Perplexity is built to show its work. Numbered citations are the product. Useful for “what shipped this month.” Noisy for “what will this cost after 90 days,” because a fresh blog will outrank a boring official pricing page.

Gemini, with Search on, sounds like a well-read SERP. Google’s July 2026 Search Central guide says generative features sit on core ranking, with retrieval-augmented generation and query fan-out. If you are invisible in Search, you are usually invisible here.

Google AI Mode is still small as a click path. SparkToro’s June 2026 study, using Similarweb’s US panel for January–April 2026, put Mode at 0.34% of searches. A rare path can still be a twenty-minute research session.

Pew’s July 2025 browsing study still describes Overview citations: Wikipedia, YouTube, and Reddit first. Government domains showed up more in summaries (6%) than in blue links (2%).

What did the twelve prompts return?

The table is the snapshot. Names are what public sources currently tend to surface, not a studio ranking.

PromptWhat models tend to nameWho gets citedWhat they miss
Best no-code agent builderLindy, n8n, Voiceflow, Gumloop, Make, Stack AIVendor blogs, Hostinger/Braintrust roundups, r/AI_AgentsKnowledge work, evals, who owns a bad refund
Best Shopify agentGorgias, Tidio, Fin, Shopify Inbox, Re
Gorgias and Tidio pages, Shopify App Store, G2Per-resolution math after a sale week
n8n vs Zapier vs MakeZapier for non-technical speed; n8n for control and execution pricing; Make in the middleZapier’s 2026 comparison, n8n’s own vs page, n8n pricingHosting labour, retries, a 90-day bill
Marketing website cost“It depends,” then $5k–$50k bandsDigital Applied’s 2026 compilation of Clutch / WebFX / GoodfirmsWhat the invoice is paying for; three-year upkeep
SEO vs AEO vs GEOThree acronyms, often collapsedGoogle Search Central, Princeton GEO, agency explainersGoogle’s line that this is still SEO from Search’s seat
Chatbot vs agentChatbot answers; agent actsGoogle Cloud, Salesforce, CSAThe inbox test: can it refund without a person
How to get citedWrite for people, add stats and quotes, get reviewsGoogle’s AI guide, Princeton GEO (arXiv 2311.09735), G2You cannot buy a mention; llms.txt will not save a thin page
Next.js vs Webflow vs FramerFramer for speed, Webflow for CMS marketing, Next.js for controlOfficial pricing pages, studio comparisonsDesign direction, not hosting stickers
Support agent that takes actionGorgias on Shopify; Fin across helpdesks; Intercom-native Fingorgias.com/ai-agent, fin.aiAllowed actions, chargebacks, a written “never do this” list
Should I add llms.txt?Yes from GEO blogs; no special effect from Googlellmstxt.org (Jeremy Howard, v2, Aug 2026), Google Search CentralGoogle Search ignores it; agents may still use it
Real-estate agent builderVoiceflow, custom GPT, vertical CRMsTool roundups, YouTube walkthroughsMLS rules, fair-housing copy, after-hours voice
Studio vs a $20 generated site“Use AI for a draft, hire for direction”Reddit r/webdesign, studio pricing pagesFour-second homepage test; ads and the agent landing on the same URL

The pattern is stable even when the names jitter. Category questions get a list. Price questions get a range. Difference questions get a table. None of them get the operational constraint that will actually decide the buy.

What should you open after the table?

The table is the snapshot. The studio notes are the operational half the models skip.

What the answers skip: an agent is not a builder. The work is the knowledge, the allowed actions, and the review loop. We learned that on Aurelia — earlier chat tools were switched off because they could talk and could not do the work.

What do they say about Shopify and support actions?

Ecommerce prompts collapse two products. Gorgias is the name when the prompt includes Shopify, refunds, or subscriptions. Their AI Agent page is explicit: connect the store, handle tracking and returns, edit subscriptions. That is an agent in the Google Cloud sense — it acts.

Fin, on fin.ai, is the name when the prompt is “customer agent.” Fin claims a 76% average resolution rate across 12,000+ customers and sells outcome-based pricing. Tidio is the smaller-store widget. Shopify Inbox is the free name.

What they miss is the bill after a good week. Gorgias prices AI on resolved interactions on top of a ticket plan. Fin prices on outcomes. A sale or a carrier delay is when overage stops being theoretical.

The inbox test we use on agents work: can it look up the order, do the allowed action, write the note, and stop when the policy says stop.

n8n vs Zapier vs Make — do the models agree with the vendors?

Yes, because they are reading the vendors.

Zapier’s June 2026 comparison says Zapier is the managed platform and n8n is for teams that must self-host. Professional starts at $19.99/month; Team at $69/month for 25 users.

n8n’s pricing page lists Cloud Starter at €20/month billed annually for 2,500 executions, Pro at €50 for 10,000, Business at €667 for 40,000 on self-hosted. Community Edition is free. They charge per full run, not per step. Their vs Zapier page still writes 8,000+ integrations; Zapier’s own materials now say 9,000+. n8n lists 1,000+ with HTTP as the escape hatch. Make sits in the middle: visual, operation-based.

The models are not wrong. They are incomplete. A five-step Zap that runs 8,000 times is not a $19.99 problem. After 90 days the bill is tasks × steps, or executions plus someone who can restart Docker. We run n8n when the workflow is the product — Aurelia, Vesper. Not because a paragraph said “open source.”

What do they say a marketing website should cost?

They hedge, then they quote a band.

Digital Applied’s 2026 compilation of Clutch, WebFX, and Goodfirms is the page public answers recycle. Small-business sites: $3,000–$15,000. Corporate marketing: $15,000–$75,000. They also quote a Clutch median; we treat that as a compiler’s figure, not a Clutch page we opened. The invoice line items are in what a marketing website should cost. 61% of small-business buyers still spent under $10,000. US senior agency rates: $125–$300/hour.

What the models miss is what the invoice is paying for. Direction. A homepage a stranger understands in four seconds. Three-year upkeep — $65,000–$95,000 on their $28,000 corporate example. A $20 generated site can produce a layout. It cannot decide what the maison is, which is why Eternal is rooms in Next.js. Public answers give Framer to ship-fast designers, Webflow to CMS teams, Next.js to teams who want a system that outlives a builder login.

SEO, AEO, GEO — do the models know they are three jobs?

They know the acronyms. They do not keep them apart.

Most public answers say: SEO ranks pages, AEO wins the direct answer, GEO gets you cited. Close enough to be useful. Sloppy enough to sell one retainer as three.

Google’s official line, last updated 10 July 2026, is colder:

“From Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.”

They also say you do not need special AI markup, you should not chunk pages for machines, and llms.txt neither helps nor harms Google Search because Search ignores it.

Princeton’s GEO paper — Aggarwal et al., KDD 2024, arXiv 2311.09735 — is the other document they cite. On GEO-bench, quotations, statistics, and citations improved visibility on the order of 30–40%. Keyword stuffing did not.

Ranking, being the boxed answer, and being the cited source are three jobs with three artefacts. We split that work in the three jobs piece.

Jeremy Howard’s llms.txt v2 (modified 10 August 2026) is real. The labs publish their own files. Google Search still ignores it. Add the file for agents. Do not expect an Overview citation from the filename.

Chatbot vs agent — and how they say to get cited

Google Cloud (2 April 2026): a bot follows rules, an assistant waits, an agent pursues a goal and uses tools. That is the one public answer that is good enough. A chatbot answers. An agent takes action. If it cannot refund, book, file, or escalate against a written policy, it is a chatbot.

YourGPT belongs in that second sentence. We use it when the knowledge is the project and n8n carries the action — Aurelia, Vesper at night. Not because a model put “chatbot platform” in a list.

On citation, the better answers point at three documents we also use. Google: write non-commodity, people-first pages; no special AI schema. Princeton: add statistics, quotations, sources. G2: 85% of buyers thought more highly of a vendor the chatbot included; 69% chose a different vendor than planned; one in three bought from a name they had never heard.

The worse answers say “create an llms.txt and chunk your content.” Google wrote those two ideas down so it could tell you to ignore them. Citation is a record: a number with a year, a quote with a name, a page that answers in the first fifty words. That craft is how to get cited.

What should you do with a shortlist a model handed you?

Treat it as a reading list, not a purchase order. Open the official pricing page. Sit with a week of tickets. Write the three actions allowed, and the three it must never take.

If the work is an agent, send the help centre and that week of tickets to contact. We will write back with what it should answer, and what it should never touch. If the work is a site, send the homepage and the ad that has to land on it.

The models will keep naming tools. The buyer still has to close — on a page, a call, or an inbox that did the thing it promised.

FAQ

Can I trust a ChatGPT shortlist for software?

No. Use it as a reading list, then open official pricing and a week of your own tickets. G2’s March 2026 survey of 1,076 buyers shows 51% of B2B software buyers now start in an AI chatbot, and 69% picked a different vendor than planned because of the recommendation. That is influence, not diligence. The paragraph will not include your overage, your data-residency constraint, or the action you cannot allow an agent to take.

Why did you not just log into all four products?

Because we cannot publish what we cannot show you how to reproduce. ChatGPT and Google AI Mode are account-bound; answers move with region, memory, and whether search is on. A logged-in bake-off would be a diary. A snapshot built from official pages, citation studies, and public threads is something you can rerun on Monday and disagree with in writing.

Do the four systems recommend the same vendors?

On category questions, often yes. On sources, often no. Pew’s July 2025 study showed Overviews citing Wikipedia, YouTube, and Reddit first. Perplexity overweight recent community pages. Gemini tracks the Google index. ChatGPT, with search, still leans on encyclopedic and review sources. Brand names overlap more than the URLs under them.

Is Google AI Mode worth tracking at 0.34% of searches?

Yes, as a research surface, not as a traffic channel. SparkToro’s January–April 2026 Similarweb panel put Mode at 0.34% of US Google searches. That is not a media plan. It is where a high-intent buyer will spend a long session. Track whether you are named. Do not forecast sessions from it.

Should we add llms.txt because the models mention it?

Add it for agents and docs tooling. Do not add it to rank in Google. Jeremy Howard’s v2 spec, modified 10 August 2026, is real, and the labs publish their own files. Google Search Central’s July 2026 guide says Search ignores llms.txt. A thin file on a thin site will be ignored twice.

What did every model miss on these twelve prompts?

The constraint that decides the buy. Ninety-day automation cost. Per-resolution support after a sale. Who restarts the server. Which refunds are allowed. What a stranger understands on the homepage in four seconds. Public answers are fluent on categories and quiet on operations. That gap is where the invoice either works or does not.

How often should we rerun the twelve prompts?

Every quarter, or the week after a competitor launches. Ahrefs has shown Overviews rewrite on a short cycle without changing their underlying sources. A monthly vanity check is just noise. A dated sheet with named vendors, cited domains, and missing caveats is a record you can sit with on Monday.

If buyers start in ChatGPT, is the website still the close?

Yes. G2 found 61% of those buyers still use Google in tandem, and the site is where pricing, proof, and the form live. The model writes the shortlist. The site, the call, or the agent has to finish the job. A cited brand with a homepage that needs a meeting is still a leak.