We audited 42 AI GTM products and our hypothesis failed
We scored 42 AI GTM homepages against our five stage model expecting agent washing. The review refuted 8 of 9 large gaps. What survived: 26 products market at Stage 4 and 24 require a Stage 3 buyer.
[ key takeaways ]
- In our audit of 42 AI GTM product homepages, 26 market at Stage 4 while only one requires a Stage 4 buyer and 24 require Stage 3: a market-wide one-rung offset that needs no deception to exist.
- The agent-washing hypothesis failed: nine products scored gaps of two rungs or more on first pass and eight did not survive adversarial review.
- Six of seven apparent Stage 5 claims reduced to Stage 4 on review, because asserting autonomous execution and asserting a system that rewrites its own scoring from outcomes are different claims.
- The audit corrected our own map: two Stage 5 placements did not survive their own vendors' words and moved down.
What do you do when your own hypothesis fails?
You publish it. In August 2026 we fetched the homepage of every product on our AI GTM stack map, 46 pages in one pass on one day, of which 42 returned usable marketing copy. We scored each page twice against the five stage AI GTM Maturity Model: once for the stage the copy claims to deliver, and once for the stage a buying organization must already be operating at for that promise to land. Every gap of two rungs or more then went to an adversarial reviewer with one instruction, refute this gap using the vendor's own page, and default to refuted when uncertain.
We went in expecting agent washing. The market has spent two years attaching the word agent to everything that ships, and our working hypothesis was that we would find Stage 3 products dressed in Stage 5 language across the corpus.
The hypothesis failed. Nine of the 42 products scored a gap of two rungs or more on first pass. On adversarial review, eight of the nine did not survive as scored. After correction the whole corpus holds five two rung gaps and zero gaps of three rungs, and 31 of the 42 products sit one rung above what they require, which is what software marketing has looked like for as long as software has had marketing.
The failure is the reason to keep reading. An audit that only found fault would be a marketing document. What survived the review is smaller than the accusation we expected and, for a buyer, considerably more expensive.
- definitionRequired stage
The required stage is the AI GTM maturity stage a buying organization must already occupy for a product's promise to land: the signals the copy assumes are being captured, the destinations it assumes exist for an automated action, and the owners it assumes will work whatever the system fires. A homepage claims a stage with its verbs. It requires a stage with its assumptions, and both are readable in the same copy.
Where does the market claim to be, and where does it need you to be?
In our audit of 42 products, 26 market themselves at Stage 4, Signal-Driven Systems. One product requires a Stage 4 buyer. Twenty four require a Stage 3 buyer, 41 of 42 require Stage 3 or below, and nothing in the corpus requires Stage 5.
[ fig. 01 · claimed stage against required stage, 42 products ]
Stage the copy claims
Stage the buyer must already be at
The Stage 4 claims arrive in three recurring grammatical moves. The first names the actor: the system rather than a person is the subject of the sentence that builds pipeline, advances deals, and works around the clock. The second names the trigger: a site visit, a reply, a threshold crossing starts the work, so execution begins without anyone deciding to begin it. The third replaces headcount, promising scale without hiring. A fourth pattern carries less information than the other three, the word agent used as a noun in product names and navigation menus. In a substantial minority of the Stage 4 population we scored, the claim lives in a product name rather than in a sentence with a verb, and a noun says almost nothing about who is still in the loop.
The counterweight deserves equal visibility. Six of the 42 products claim exactly the stage they require, and two of those argue against the rung above them, explicitly bounding their agent vocabulary to workflows a human authors and can trace. Several more stop short of claims the category's vocabulary would have allowed them. Across all 42 pages we found less reaching than the discourse around this market suggests.
Did anyone actually claim Stage 5?
Seven products appeared to on first pass, and six of the seven were reduced to Stage 4 on review, for the same reason each time. Their copy asserts autonomous execution, which is a Stage 4 property. Stage 5 requires a different thing, outcomes changing the system's own scoring, routing, or creative with no human editing them, and the check that did most of the work in the review was a plain word search: three of the pages we had scored at Stage 5 contain no instance of retrain, adapt, self improve, gets smarter, or over time anywhere in the captured text.
One Stage 5 claim survived, resting on a named reinforcement learning capability, and it belongs to the only product in the set that requires a Stage 4 buyer. Because only gaps of two rungs or more went to review, that row was never adversarially tested, so it should be read as the least stress tested row in the audit rather than the most secure.
The top of the ladder is close to unclaimed. The market's center of gravity is autonomous action, and on the pages we captured, almost nobody asserts that outcomes retrain anything. Whatever the discourse implies, the homepages themselves mostly stop one rung short of the adaptation claim.
Why is the offset expensive if nobody is lying?
[ fig. 02 · which claim is sold to which operating model ]
columns: stage required of the buyer
21 of 42 products share one cell, a Stage 4 claim over a Stage 3 requirement. The outlined diagonal, copy that claims exactly what it requires, holds 6 products.
Half of the corpus, 21 of the 42 products we scored, sits in a single cell of that matrix, a Stage 4 promise sold to a Stage 3 operating model. Nobody has to be lying for that cell to be the most expensive fact in the category. Each vendor is doing the normal thing, selling the destination rather than the starting point, one rung of aspiration over a real product. The condition only becomes visible in aggregate, and no single homepage produces it.
What the aggregate does to a buyer is specific. Stage 3 is the price of admission for most of what is on sale, and Stage 3 means shared workflows, governed context, and repeatable processes, an operating model where an automated action has somewhere to land. A team at Stage 2 buying a Stage 4 product gets software that works exactly as advertised while the outcome never arrives, because a signal with no destination becomes an unread channel rather than pipeline. The product did what the page said. The prerequisite was a rung the team had not reached, and it appeared nowhere on the page as a prerequisite.
The aggregate also feeds the buyer side pattern we measure in audits. Stage drift, the gap between the AI GTM maturity stage a company reports and the stage its evidence verifies, opens without anyone overstating anything, and this audit documents one supply line into it: a team calibrates its self description to the vocabulary of the homepages it evaluated software on, and that vocabulary runs one rung above the operating model underneath. Read against the 17 verification criteria, the same team scores where its evidence puts it. The distance between those two readings now has a measured, market wide source.
How was the audit scored?
[ fig. 03 · the method, end to end ]
The refutation pass is what makes these numbers publishable. The reviewer worked only from the vendor's own captured page, and a gap was recorded as surviving only when refutation failed. Under that treatment both of the three rung gaps found on first pass collapsed to two, and six of the seven Stage 5 claims fell. First pass classifiers are eager. The discipline in this audit came from the second pass, which is why the corrected scores, and only the corrected scores, are the record.
The limits, stated plainly. This is homepage copy, one snapshot, on one day, and a marketing page is not product documentation: a vendor whose homepage stops short of a claim may still ship the capability, and a vendor whose homepage reaches may still deliver on it. Classification is judgment applied against a published ladder, and 7 of our 42 rows carry low confidence, several because the captured text was thin or ended inside a navigation menu. The set is 42 products selected for a market map, a chosen list rather than a sample, so it supports claims about these pages and no inference about the category at large. And we hold commercial relationships with two vendors in the set. Both of their first pass scores were reduced on review, so the relationships cut against the finding rather than for it, and we state them anyway.
What did the audit do to our own map?
This is the section we most want read. In the audit, seven products' homepage copy read as a Stage 5 claim on first pass, and six of the seven reduced to Stage 4 against the retrain test, because asserting autonomous execution and asserting a system that rewrites its own scoring from outcomes are different claims. Separately, our own stack map had placed seven products at Stage 5 before the audit, and two of those placements did not survive their own vendors' words, so we moved them down.
The row was arguably wrong in both directions, because the one product in the corpus whose copy asserts that outcomes change its own decisions is a product we had placed a rung below where its copy reaches. Our map already said Stage 5 was nearly empty. The audit says it is emptier than our map admitted, which strengthens the claim the map was making and still forced a correction to rows we had published.
What should a buyer do with this?
Read the verbs rather than the nouns. Agent in a product name told us almost nothing in this corpus, while drafts, routes, pauses, and retrains told us exactly which stage was being claimed and who is still in the loop.
Then, before buying anything positioned at Stage 4, write down what the tool will fire into. Which CRM, which sequence, which named owner, and what happens within a day of an alert. In our audit, 21 of the 42 products assume that answer already exists inside your company, and when it does, the offset is harmless, because a Stage 3 operating model is exactly what a Stage 4 promise needs underneath it. When the answer does not exist, the honest purchase is one stage lower, and the destination gets built first.
And place yourself before you place the vendors. The free assessment takes about two minutes and scores where your evidence puts you on the same five stages this audit used. A team that knows it verifies at Stage 3 can read 26 Stage 4 homepages for what they are: real products, described one rung ahead of the operating model they assume you already run. Bought in that order, the software works and the outcome arrives.
We expected to write about deception. The evidence handed us an offset instead, a market selling one rung ahead of its buyers the way software always has, and a correction to two rows of our own map. Reporting that accurately seemed worth more than the post we planned to write.