What the AI engines tell your buyers before your first call
We ran 36 real buyer questions through ChatGPT, Claude and Perplexity. 108 answers, 703 citations, no incumbent, and all three engines defer the one variable nobody has defined.
[ key takeaways ]
- Across 108 responses, Perplexity cited sources on every answer while ChatGPT and Claude cited on 31%, so the engine matters more than the question.
- Inside ChatGPT the split is sharp: comparison questions earned citations 75% of the time and definition questions never did, while Claude inverted the pattern.
- 82% of the 489 cited domains appeared exactly once, so no incumbent owns this category in the engines' eyes.
- All three engines recommended agency first, then hire, and then handed the decision back with a qualifier none of them defines, which is the gap a published standard fills.
What do the AI engines tell your buyers before the first meeting?
By the time a buyer books a call, they have usually already asked an AI engine the question they are about to ask you. We wanted to know exactly what the engines say back, so we measured it. We took 36 questions drawn from real sales conversations, the questions buyers ask when an operator leaves or a system stalls, and ran each one through ChatGPT (gpt-5.5), Claude (claude-sonnet-5) and Perplexity (sonar-pro) with web search on. The run produced 108 usable responses carrying 703 citations across 489 unique domains, and it cost $5.04. Gemini was in the design and was dropped after three consecutive upstream rate limits, so this is a three-engine result and we read it as one.
Four things came back that we have not seen published anywhere else. Which engines cite sources at all. Which question types trigger a citation in which engine. Who is being cited in the AI go-to-market category. And what all three engines recommend when a buyer asks whether to hire or to work with an agency, which turned out to be the most useful finding in the set.
Which engine actually cites its sources?
Perplexity cited sources on all 36 of the 36 questions we asked it. ChatGPT cited on 31% of questions, and Claude on 31%. Perplexity is a search product with a language model attached, and it behaves like one. The other two are language models with a search tool attached, and they decide, question by question, whether to browse the web or answer from memory.
Which engine your buyer prefers therefore changes what your buyer hears. A Perplexity user gets an answer stitched from live sources every time. A ChatGPT user gets the model's memory most of the time, which means the category as it existed whenever the training data was cut, with no path for anything published since to reach them.
[ fig. 01 · citation rate by engine ]
The headline rates hide the useful part. The 31% we measured for each of ChatGPT and Claude sounds like two engines behaving the same way. They are making nearly opposite decisions about when to look.
Which questions make an engine go looking?
We banded the 36 questions into four types before the run. Comparison and decision questions, where a buyer is choosing between two options. Diagnostic questions, where something is underperforming and the buyer wants to know why. How-to questions. And definition questions, where the buyer wants a term explained. Pooled across all three engines, comparison questions earned a citation 72% of the time, diagnostics and how-tos 41% each, and definitions 56%. Inside a single engine the spread is much sharper.
ChatGPT cited 75% of comparison questions and 0% of definitions in our run. Claude ran the other way, citing 42% of comparisons and 67% of definitions. Perplexity cited 100% of everything.
[ fig. 02 · citation rate by engine and question type ]
A page that answers a comparison question can earn citations in all three engines. A page that answers a definition can earn citations in Perplexity and Claude and, in our run, never in ChatGPT, which cited no definition question at all. Nothing on the list is wasted, but the ceiling differs by question type, and that ordering should decide which pages get built first.
The split also explains a common misreading. A team that publishes definitional content and checks its AI visibility inside ChatGPT will conclude that AI search ignores it, and inside ChatGPT that conclusion is correct. Two engines away, the same page can be earning citations on every relevant question.
Across the four bands in our audit, the average cited answer drew on 5.9 to 7.3 sources. A citation puts you on a panel of six or seven voices rather than making you the answer. The panel is still worth joining, because of who currently sits on it.
Who do the engines cite in this category?
Almost nobody, repeatedly. The 703 citations in our audit spread across 489 unique domains, and 400 of those domains, 82% of them, were cited exactly once. The top 20 domains together account for 22% of all citations. By type, 82.6% of citations went to long-tail agencies and independent sites, 9.4% to vendor blogs, 6.8% to community threads, and 1.1% to analyst firms, with Gartner's domain appearing three times in the 703.
[ fig. 03 · who gets cited, by source type ]
The long tail in one number: 400 of the 489 cited domains appeared exactly once, and the single most-cited domain, linkedin.com at 28 citations, is individual operators' posts rather than any company's page.
There is no incumbent to displace, and the generous reading is the accurate one. Nobody has published the standard reference for this category, so the engines stitch answers together from whoever has said anything specific.
Consider what that pool means for the buyer on the other end. These are decisions that set headcount and budget for a year, hire an engineer, retain an agency, rebuild an outbound motion, and the research layer answering them is assembled from single posts by single operators, most of them cited exactly once in our data. The engines are doing their job with the material available. The material is thin.
The single most-cited domain in the run is linkedin.com, with 28 citations, and the cited items are posts by individual operators rather than company pages. Reddit is second at 18. The cited posts share a shape we saw across the run. A first-person operator makes one specific claim, with a number or a named mechanism, and takes a clear position. One cited post title, "GTM engineering is the new architecture function", is a fair sample of the genre. A LinkedIn post written as the answer to a question buyers actually ask does distribution and AI citation work at the same time, from the same asset.
What does the engine search on your buyer's behalf?
- definitionFan-out query
A fan-out query is a search the engine composes and runs itself while answering, fanning one buyer question out into several specific web searches. Engines report these alongside the response, which makes them a record of how the machine translates a buyer's question into search language, and a target list for anyone deciding what to publish.
Our run captured 56 unique fan-out queries, and nearly all of them are versus-shaped. A sample, verbatim: gtm engineer ai go-to-market systems agency vs hire 2026 b2b saas, revenue operations outsource vs hire revops consultancy factors, ai sdr tools vs human sdr cost comparison. When we checked the same phrases in conventional keyword tools, they registered no measurable volume. The vocabulary a buyer's research actually runs on is priced at zero by the tools most teams use to decide what to write, which means the target list is sitting in the open with nobody bidding on it.
The shape matters as much as the list. A buyer asks a broad question and the engine translates it into narrow either-or comparisons before reading anything. The comparison band is both the most-cited question type in our run and the form the engines impose on every other question on the way to an answer, which is two separate reasons to write in it.
What do all three engines recommend on hire versus agency?
The question in our set closest to a purchase decision asks whether a mid-sized B2B software company should hire a GTM engineer or work with an agency to build its AI go-to-market systems. All three engines gave the same shape of answer. Work with an agency to design and build the first version, then hire an internal owner to run it. The reasons the engines gave for the sequence were speed now, institutional ownership later, and a de-risked path to the hire, phrased by one engine as reducing hiring risk and by another as de-risking the timeline. That is the case we make in first meetings, and the engines assembled it from public sources without ever having heard of us.
Each engine then closed the same way. Having recommended the sequence, it handed the final decision back to the buyer with a qualifier it never resolved. One closed on depends heavily on your specific context. One closed on if budget allows. One closed on conditionals about whether the capability is core to you. Elsewhere in the run the qualifier gets named outright, as in depends on your stage, technical bandwidth, and budget on the tool-selection question. What no response does is cite a published standard the buyer could check themselves. The closest any engine came was improvising a scoring rubric on the spot, a different one each time it was needed, which proves the demand for a standard and the absence of one in the same answer. We publish the rung-by-rung version of that decision, so we read those closings with some attention.
Sit in the buyer's seat for the next step. The recommendation was clear, the qualifier sounded reasonable, and the natural follow-up, what stage are we actually at, sends the buyer straight back into the same long tail the first answer was assembled from.
The deferral is fair on the engines' part. They defer because the reference they would need does not exist in their sources. The citation pool we measured is 82.6% long tail, individual posts and small agency sites, and a definition of maturity stages with verifiable criteria is exactly the kind of asset a long tail does not produce. The engines behave the way a careful analyst behaves with no standard to cite. They recommend the pattern and defer the particulars.
Why is the deferral the gap?
Look at what the engines have already settled and what they have left open. They already recommend the sequencing model, an agency-built first version with an internal owner hired to run it. They already argue for it on key-person-risk grounds. A buyer who asks the engines before your first meeting arrives having heard that case. What the buyer has not heard is any way to act on it, because the deciding variable, their stage, was named and then left undefined.
That undefined variable is the same gap we meet inside companies. Teams report one maturity stage and their evidence verifies another, a gap we call stage drift, and it opens because nobody defined what good looks like at each rung, so every team answers by feel. It is why the five-stage model carries 17 verification criteria, each phrased as evidence a team either has or does not have, and why the assessment exists. A buyer who has heard "it depends on your stage" from three engines goes looking for a way to establish their stage, and the asset that wins that moment is the published standard the engines had nowhere to find.
What we are doing with the findings, in the order the data suggests. Build for the comparison band first, since it is the only band that earned citations in all three engines and it is where the buying decision happens. Treat LinkedIn posts as citable assets rather than as a separate channel, written on the questions from the audit. Keep the definition pages for Perplexity and Claude, which cited definitions at 100% and 67% in our run. And re-run the identical 36 questions at 60 days, same models and same wording, because a changed question set produces a new baseline rather than a delta.
The re-run matters more than it sounds, because a measurement this small and this repeatable turns a snapshot into an index. The run cost $5.04. The identical run in two months costs about the same and buys the delta, which is the number that shows whether anything published in between moved a citation.
The engines are already in your sales cycle. In our audit they held no loyalty to any incumbent, cited whoever had said something specific, made the argument we would have made, and deferred exactly one variable. Defining that variable in public is the open move, in this category and probably in yours.