We audited 42 AI GTM products and our hypothesis failed
We scored 42 AI GTM homepages against our five stage model expecting agent washing. The review refuted 8 of 9 large gaps. What survived: 26 products market at Stage 4 and 24 require a Stage 3 buyer.
[ key takeaways ]
- In our audit of 42 AI GTM product homepages, 26 market at Stage 4 while only one requires a Stage 4 buyer and 24 require Stage 3: a one-rung offset across the 42 products, which we read as ordinary positioning rather than deception.
- The agent-washing hypothesis failed: nine products scored gaps of two rungs or more on first pass and eight did not survive adversarial review.
- Six of seven apparent Stage 5 claims reduced to Stage 4 on review, because asserting autonomous execution and asserting a system that rewrites its own scoring from outcomes are different claims.
- The audit corrected our own map: two Stage 5 placements did not survive their own vendors' words and moved down.
What do you do when your own hypothesis fails?
You publish it. In August 2026 we fetched the homepage of every product on our AI GTM stack map, 46 pages in one pass on one day, of which 42 returned usable marketing copy. We scored each page twice against the five stage AI GTM Maturity Model: once for the stage the copy claims to deliver, and once for the stage a buying organization must already be operating at for that promise to land. Every gap of two rungs or more then went to an adversarial reviewer with one instruction, refute this gap using the vendor's own page, and default to refuted when uncertain.
We went in expecting agent washing: Stage 3 products dressed in Stage 5 language across the corpus.
The hypothesis failed. Nine of the 42 products scored a gap of two rungs or more on first pass. On adversarial review, eight of the nine did not survive as scored. After correction the whole corpus holds five two rung gaps and zero gaps of three rungs, and 31 of the 42 products sit one rung above what they require. What survived the review is smaller than the agent washing we expected and, for a buyer, more expensive.
- definitionRequired stage
The required stage is the AI GTM maturity stage a buying organization must already occupy for a product's promise to land: the signals the copy assumes are being captured, the destinations it assumes exist for an automated action, and the owners it assumes will work whatever the system fires. A homepage claims a stage with its verbs. It requires a stage with its assumptions, and both are readable in the same copy.
[ fig. 01 · the two scores, applied to eight phrases ]
Stage the copy claims
Stage the buyer must already be at
Where does the market claim to be, and where does it need you to be?
In our audit of 42 products, 26 market themselves at Stage 4, Signal-Driven Systems, while 24 require a Stage 3 buyer and one requires Stage 4.
[ fig. 02 · claimed stage against required stage, 42 products ]
Stage the copy claims
Stage the buyer must already be at
The Stage 4 claims arrive in three recurring grammatical moves. The first names the actor: the system rather than a person is the subject of the sentence that builds pipeline, advances deals, and works around the clock. The second names the trigger: a site visit, a reply, a threshold crossing starts the work, so execution begins without anyone deciding to begin it. The third replaces headcount, promising scale without hiring. A fourth pattern carries less information: in several of the Stage 4 pages we scored, the claim lives in a product name rather than in a sentence with a verb, and a noun says little about who is still in the loop.
Six of the 42 products claim exactly the stage they require, and two of those argue against the rung above them, explicitly bounding their agent vocabulary to workflows a human authors and can trace.
Did anyone actually claim Stage 5?
Seven products appeared to on first pass, and six of the seven were reduced to Stage 4 on review, because their copy asserts autonomous execution, a Stage 4 property, where Stage 5 requires outcomes changing the system's own scoring, routing, or creative with no human editing them. The check that did most of the work in the review was a plain word search: three of the pages we had scored at Stage 5 contain no instance of retrain, adapt, self improve, gets smarter, or over time anywhere in the captured text.
One Stage 5 claim survived, resting on a named reinforcement learning capability, and it belongs to the only product in the set that requires a Stage 4 buyer. Because only gaps of two rungs or more went to review, that row was never adversarially tested, so it should be read as the least stress tested row in the audit rather than the most secure.
Why is a one rung offset expensive?
[ fig. 03 · which claim is sold to which operating model ]
columns: stage required of the buyer
21 of 42 products share one cell, a Stage 4 claim over a Stage 3 requirement. The outlined diagonal, copy that claims exactly what it requires, holds 6 products.
The measured fact is the cell: 21 of the 42 products we scored sit at a Stage 4 promise over a Stage 3 requirement, and the condition is only visible in aggregate. Our reading of why is that each page sells the destination rather than the starting point, ordinary positioning one rung ahead of the product. That is a hypothesis, and the audit does not control for its own rubric: a rule that reads claims from verbs and requirements from assumptions could produce a one rung offset by construction. A second scorer working blind on the same pages, or the same rubric run on product documentation instead of homepages, would test it.
What the offset costs a buyer is also unmeasured here, so read this as the mechanism we expect. Stage 3 is the price of admission for most of what is on sale, and Stage 3 means shared workflows, governed context, and repeatable processes, an operating model where an automated action has somewhere to land. A team at Stage 2 buying a Stage 4 product gets software that works as advertised while the outcome does not arrive, because a signal with no destination becomes an unread channel rather than pipeline. The prerequisite is a rung the team has not reached, and the copy carries it as an assumption rather than stating it. What this does not account for is a vendor's onboarding or a services partner supplying the missing rung; following Stage 2 buyers of Stage 4 products after purchase would settle it.
How was the audit scored?
[ fig. 04 · the method, end to end ]
The limits. This is homepage copy, one snapshot, on one day, and a marketing page is not product documentation: a vendor whose homepage stops short of a claim may still ship the capability, and one that reaches may still deliver on it. Classification is judgment applied against a published ladder, and 7 of our 42 rows carry low confidence, several because the captured text was thin. The set is 42 products selected for a market map, a chosen list rather than a sample, so it supports claims about these pages and no inference about the category at large. We hold commercial relationships with two vendors in the set; both of their first-pass gaps shrank on review, which lowers the offset rather than inflating it, so any partner bias in the scoring works against the finding, not for it.
What did the audit do to our own map?
Our own stack map had placed seven products at Stage 5 before the audit. Read against the same retrain test, two of those placements did not survive their own vendors' words, so we moved them down. The row was arguably wrong in both directions, because the one product in the corpus whose copy asserts that outcomes change its own decisions is a product we had placed a rung below where its copy reaches. Our map already said Stage 5 was nearly empty; the audit says it is emptier than the map admitted.
What should a buyer do with this?
Read the verbs rather than the nouns. Agent in a product name told us little in this corpus, while drafts, routes, pauses, and retrains told us which stage was being claimed and who is still in the loop.
Before buying anything positioned at Stage 4, write down what the tool will fire into: which CRM, which sequence, which named owner, and what happens within a day of an alert. In our audit, 21 of the 42 products assume that answer already exists inside your company, and when it does, the offset is harmless, because a Stage 3 operating model is what a Stage 4 promise needs underneath it. When the answer does not exist, the better purchase is one stage lower, and the destination gets built first.
Place yourself before you place the vendors. The free assessment takes about two minutes and scores where your evidence puts you on the same five stages this audit used.
We expected to write about agent washing. The evidence handed us an offset instead, 42 products describing themselves one rung ahead of their buyers, and a correction to two rows of our own map.