The standard
How to verify a stage
17 criteria across five stages. Every one names the evidence an auditor has to be shown, and the near-miss that looks like proof and is not. A description invites a self-report. A standard can be checked.
Why the readings disagree
Self-reported stages drift up, and nobody lied
Ask a revenue team which stage they are at and the answer comes back one to two rungs above what an audit finds. That is not dishonesty, and being precise about it matters, because the accusation is both wrong and useless.
[ fig. 00 · the mechanism ]
what teams ask themselves
- Do we have scoring?
- Do we have workflows?
- Do reps use AI?
Every one of these is a question about possession, and every honest answer is yes.
what an audit asks
- Does scoring change what happens next?
- Did a workflow finish last week with no person in the middle?
- What would stop if your best operator left?
These are questions about operating model, and they are the only ones that predict anything.
possession questions predict nothing. operating-model questions predict everything.
So every criterion below carries three parts. The claim a team would make, the evidence that settles it, and the near-miss that looks like proof and is not. The third part does the work. Without it, an honest optimist passes every check.
[ fig. 01 · stage 1 verification · 1 criterion ]
Manual AI
Execution is manually initiated in chat tools. Knowledge, tools, and outputs depend on the operator.
the claim
“We are using AI in go-to-market.”
evidence that settles it
At least one person uses an AI tool in their weekly work, and can show you what they produced with it this month.
Licences purchased. A seat count is a procurement record, not a usage record. Check consumption before you count it.
if you only ask one question: if the person who uses ai most went on leave tomorrow, what would stop?
[ fig. 02 · stage 2 verification · 4 criteria ]
Assisted Execution
AI assists inside individual tasks, but prompts, context, and quality still live with individuals.
the claim
“Our team uses AI to produce work faster.”
evidence that settles it
A named person is accountable for the quality of AI-assisted output in at least one function, and can say what they are accountable for.
Everyone is encouraged to use AI. Encouragement without an owner produces variance, and variance is the definition of Stage 1.
the claim
“We have prompts and context that people reuse.”
evidence that settles it
A shared prompt or context asset that at least two people used in the last 30 days, with the dates.
A prompt library that exists and that nobody has opened. Existence is not adoption.
the claim
“We know what good output looks like.”
evidence that settles it
A written definition of good for at least one output type, specific enough that two reviewers reach the same verdict.
We know it when we see it. If it is not written, it cannot be delegated, and undelegatable quality is a person, not a system.
the claim
“AI is making us more productive.”
evidence that settles it
A measured before and after on one named workflow.
A company-wide productivity claim with no per-workflow measurement. Individual acceleration is real and is not organizational maturity.
if you only ask one question: can two different people produce comparable output from the same starting point?
[ fig. 03 · stage 3 verification · 4 criteria ]
Orchestrated Workflows
Repeatable AI workflows with shared knowledge, tooling, and governance across the team.
the claim
“We have AI workflows.”
evidence that settles it
One workflow that runs end to end without a human step in the middle. Show the trigger, the steps, and a completion log with dates.
A workflow somebody kicks off manually each time. A person pressing go is the Stage 2 operating model with better tooling behind the button.
the claim
“Our knowledge is shared.”
evidence that settles it
A knowledge base the workflow reads from, with a named owner and a stated update cadence.
Prompts and context living in personal accounts. The capability is real and it belongs to an individual, so it leaves when they do.
the claim
“We have a single source of truth.”
evidence that settles it
For each key field, one name, one written definition, one owner.
Six different columns all called champion, none authoritative. Observed in the field. Every team was right, and none of them agreed.
the claim
“The team runs on shared systems.”
evidence that settles it
The same workflow serving more than one person, with usage across at least two names.
Three teams each building their own version of the same thing because nobody owns the shared layer. Observed in the field, and none of the three could be trusted.
if you only ask one question: name one workflow that completed last week without a person in the middle of it.
[ fig. 04 · stage 4 verification · 4 criteria ]
Signal-Driven Systems
Scoring, tagging, and routing trigger execution automatically from buyer signals, not from someone remembering to act.
the claim
“We act on buyer signals.”
evidence that settles it
Signals resolve to a person, not only to a company. Show the resolution rate.
Company-level intent with no person-level clarity. Observed in the field: reps cannot act on an account that is warm in the abstract, so they quietly stopped opening the tool while the renewal was defended on usage nobody could find.
the claim
“We score accounts and leads.”
evidence that settles it
A scoring output that changes what happens next automatically. Show the routing rule and a log of records it moved.
A score field that exists and that nothing consumes. A number nobody acts on is a decoration with a refresh schedule.
the claim
“The system knows what to prioritise.”
evidence that settles it
A written measurement framework: what the system is meant to move, how it is read, and who reads it.
No framework. Most organisations do not have one, which is why most Stage 4 claims cannot be checked in either direction.
the claim
“Execution triggers itself.”
evidence that settles it
Measured time from signal to action, without anyone remembering to act.
A weekly meeting where a human reviews the signal list. That is a calendar reminder wearing a system's clothes.
if you only ask one question: when a signal fires, what happens next, and who decided that without being asked?
[ fig. 05 · stage 5 verification · 4 criteria ]
Adaptive GTM Engine
Triggers, knowledge, and tools improve continuously from outcomes while humans govern strategy, thresholds, and exceptions.
the claim
“Our system learns from outcomes.”
evidence that settles it
One threshold, score weight or routing rule that changed without a human editing it. Show the before, the after, the outcome that caused it, and the date.
A quarterly manual model retune. That is Stage 4 with a recurring calendar invite.
the claim
“The system adapts.”
evidence that settles it
An audit trail that lets you replay the change and see why it happened.
A vendor's claim about its own product. Gartner named the vendor-side version of this agent washing. Watch it change its own behaviour, or record it as not evidenced.
the claim
“Humans focus on strategy, not execution.”
evidence that settles it
A list of what humans actually approved in the last 30 days, showing thresholds and exceptions rather than individual actions.
Humans still approving each action, one at a time, quickly. Speed of approval is not absence of approval.
the claim
“This capability is ours.”
evidence that settles it
The operator-departure test. Name the single person whose exit would hurt most, then state precisely what would stop. If the honest answer is nothing, the capability is in the system.
A confident answer with no specifics behind it. If nobody can name what would break, nobody has checked.
if you only ask one question: show me one threshold that changed because of an outcome, without a person editing it.
The distance is the number that matters
Two readings come out of this. The stage a team reports when asked, and the stage the evidence supports. The distance between them is the only figure worth planning against, because it is the one that tells you what to build next and roughly what it will cost.
A team that reports Stage 4 and verifies at Stage 2 does not have a software problem. It has two rungs of operating model to build, and it is currently paying Stage 4 licence fees for the privilege of finding that out at renewal.
A half-rung gap and a two-rung gap are different pieces of work. Which is why the honest first step is measuring the distance rather than buying anything.