All postsAI GTM

What the AI engines tell your buyers before your first call

We ran 36 real buyer questions through ChatGPT, Claude and Perplexity. 108 answers, 703 citations, no incumbent, and all three engines hand the same variable back to the buyer undefined.

Juan Felipe Campos11 min read
ShareLinkedIn

[ key takeaways ]

  • Across 108 responses, Perplexity cited sources on every answer while ChatGPT and Claude cited on 31%, so the engine matters more than the question.
  • Inside ChatGPT the split is sharp: comparison questions earned citations 75% of the time and definition questions never did, while Claude inverted the pattern.
  • 82% of the 489 cited domains appeared exactly once, so no incumbent owns this category in the engines' eyes.
  • All three engines recommended agency first, then hire, and then handed the decision back with a qualifier none of them defines. Our hypothesis is that a published standard is what is missing; the 60-day re-run tests it.

What do the AI engines tell your buyers before the first meeting?

By the time a buyer books a call, they have usually already asked an AI engine the question they are about to ask you. We took 36 questions drawn from real sales conversations, the questions buyers ask when an operator leaves or a system stalls, and ran each one through ChatGPT (gpt-5.5), Claude (claude-sonnet-5) and Perplexity (sonar-pro) with web search on. The run produced 108 usable responses carrying 703 citations, counting each cited domain once per answer, across 489 unique domains, and it cost $5.04. Gemini was dropped after three consecutive upstream rate limits, so this is a three-engine result.

Which engine actually cites its sources?

Perplexity cited sources on all 36 of the 36 questions we asked it. ChatGPT cited on 31% of questions, and Claude on 31%. Our reading is that Perplexity is a search product with a language model attached, while the other two are language models with a search tool attached, deciding question by question whether to browse. The run records citations, not searches, so that is a reading rather than a measurement. If it holds, a ChatGPT user mostly hears the category as it stood when the training data was cut.

[ fig. 01 · citation rate by engine ]

perplexity, sonar-pro100%, 36 of 36
chatgpt, gpt-5.531%
claude, claude-sonnet-531%
Share of our 36 buyer questions on which each engine cited at least one source. Perplexity cited on all of them. ChatGPT and Claude cited on 31% each, and the next figure shows they disagree about which questions deserve one.

Which questions make an engine go looking?

We banded the 36 questions into four types before the run: comparison questions, where a buyer is choosing between two options; diagnostic questions, where something is underperforming; how-to questions; and definition questions, where the buyer wants a term explained. Pooled across all three engines, the citation rate was 72% on comparison questions, 41% on diagnostics and on how-tos, and 56% on definitions.

ChatGPT cited on 75% of comparison questions and 0% of definitions in our run. Claude ran the other way, citing on 42% of comparisons and 67% of definitions. Perplexity cited on 100% in all four bands.

[ fig. 02 · citation rate by engine and question type ]

comparison
diagnostic
how-to
definition
Perplexity
100%
100%
100%
100%
ChatGPT
75%
11%
11%
0%
Claude
42%
11%
11%
67%
all three engines
72%
41%
41%
56%
Share of questions earning at least one citation in our 36-question run, split by question type. ChatGPT cites when a buyer is choosing between options and cited no definition question in our run. Claude runs the other way.

A page that answers a comparison question can earn a citation in all three engines. A page that answers a definition can earn one in Perplexity and Claude and, in our run, not in ChatGPT, which cited on 0% of definition questions.

Across the four bands, answers drew on 5.9 to 7.3 distinct domains on average, averaged over every answer in a band, cited or not. A citation puts you on a panel of voices rather than making you the answer.

Who do the engines cite in this category?

Almost nobody, repeatedly. The 703 citations in our audit spread across 489 unique domains, and 400 of those domains, 82% of them, were cited exactly once. The top 20 domains together account for 22% of all citations. By type, 82.6% of citations went to long-tail agencies and independent sites, 9.4% to vendor blogs, 6.8% to community threads, and 1.1% to analyst firms, with Gartner's domain appearing three times in the 703.

[ fig. 03 · who gets cited, by source type ]

long-tail agencies and independent sites82.6%
vendor blogs9.4%
community threads6.8%
analyst firms1.1%

The long tail in one number: 400 of the 489 cited domains appeared exactly once, and the single most-cited domain, linkedin.com at 28 citations, is individual operators' posts rather than any company's page.

Share of the 703 citations in our audit by the type of source cited. The engines assemble answers from a long tail of small sites; our reading is that no standard reference exists to displace them.

That is the measurement: a pool with no incumbent to displace. Our reading of why is that no standard reference for this category has been published, so the engines stitch answers together from whoever has said something specific. The run cannot confirm that, because it records what was cited and not what was available to cite.

Behind linkedin.com at 28 citations, Reddit is second at 18. Read from their titles, the cited posts share a shape: an operator stating a position in the first line. One cited post title, "GTM engineering is the new architecture function", is a fair sample. We later coded all 37 cited posts by their opening words, and 15 of the 37 open by stating a position; the count is in its own post.

What does the engine search on your buyer's behalf?

definitionFan-out query

A fan-out query is a search the engine composes and runs itself while answering, fanning one buyer question out into several specific web searches. Engines report these alongside the response, which makes them a record of how the machine translates a buyer's question into search language, and a target list for anyone deciding what to publish.

Our run captured 56 unique fan-out queries. 15 of them are explicitly versus-shaped, and most of the rest are comparison-adjacent: pricing lookups on two named tools, in-house against agency, build against buy. A sample, verbatim: gtm engineer ai go-to-market systems agency vs hire 2026 b2b saas, revenue operations outsource vs hire revops consultancy factors 2025, ai sdr tools vs human sdr cost comparison. When we checked the same phrases in conventional keyword tools, they registered no measurable volume.

A buyer asks a broad question and the engine translates it into narrow either-or comparisons before it reads a source. The comparison band is both the question type with the highest pooled citation rate in our run and the form the engines reach for on a large share of questions on the way to an answer, which is two separate reasons to write in it.

What do all three engines recommend on hire versus agency?

The question in our set closest to a purchase decision asks whether a mid-sized B2B software company should hire a GTM engineer or work with an agency to build its AI go-to-market systems. All three engines gave the same shape of answer: work with an agency to design and build the first version, then hire an internal owner to run it. The reasons they gave were speed now, institutional ownership later, and a de-risked path to the hire. That is the case we make in first meetings, and the engines assembled it from public sources with no help from us.

Each engine then closed the same way: it handed the final decision back to the buyer with a qualifier it did not resolve. One closed on depends heavily on your specific context, one on if budget allows, one on conditionals about whether the capability is core to you. Elsewhere in the run the qualifier gets named outright, as in depends on your stage, technical bandwidth, and budget on the tool-selection question. No response cites a published standard the buyer could check themselves. The closest any engine came was improvising a scoring rubric on the spot, a different one each time, which we read as demand for a standard and the absence of one in the same answer. We publish the rung-by-rung version of that decision, so we read those closings closely.

The deferral is fair. Our hypothesis is that the engines defer because the reference they would need is all but absent from the sources they draw on; a definition of maturity stages with verifiable criteria is the kind of asset a long tail does not produce. The run does not test that. The test is to re-run the same 36 questions once a published standard has had time to be indexed and see whether the closings change.

Why is the deferral the gap?

A buyer who asks the engines before your first meeting arrives having heard the case for the sequence, argued on speed and ownership grounds, with no way to act on it, because the deciding variable, their stage, was named and left undefined. Our attempt at that definition is the 17 verification criteria behind the five-stage model, each phrased as evidence a team either has or does not have.

What we are doing with the findings, in the order the data suggests. Build for the comparison band first, since its pooled citation rate, 72%, was the highest of the four bands and it is where the buying decision happens. Write LinkedIn posts on the questions from the audit and treat them as citable assets rather than a separate channel. Keep the definition pages for Perplexity and Claude, which cited on 100% and 67% of definition questions in our run. And re-run the identical 36 questions at 60 days, same models and same wording, because a changed question set produces a new baseline rather than a delta.

The engines are already in your sales cycle. In our audit they held no loyalty to any incumbent, cited whoever had said something specific, made the argument we would have made, and handed the decision back with a qualifier none of them defined. Defining that qualifier in public is the open move, in this category and probably in yours.

ShareLinkedIn

Get the research as it publishes

One email a month with what we published, sometimes one more when a piece deserves it. Next up: the 60-day re-run of the AI answer audit, due October 24, 2026.

Unsubscribe anytime. See our privacy policy.