← Articles
September 13, 2026 · 5 minute read

The Label That Doesn't Show Up in the Demo

Of every thousand vendors marketing themselves today as "AI agents," Gartner estimates fewer than one hundred thirty actually are. The fifteen-minute demo doesn't tell the two apart. Neither does the right question — it isn't in the demo either.

A word that became a label before it became a criterion

Of every thousand products sold today as "AI agents," Gartner estimates fewer than one hundred thirty earn the name. It's not an accusation of widespread bad faith. It's the picture of a word that became a category label before it became a technical criterion — the phenomenon the firm itself named "agent washing."

The confusion has structure, and that structure is what lets you separate signal from noise without having to trust the vendor.

The three layers the demo doesn't distinguish

Gartner splits what's sold today under the same label into three layers, and the difference between them isn't subtle. Automation handles routine: it follows a fixed script, without interpreting anything new. An assistant reacts when someone asks: it answers, suggests, compares, but doesn't act on its own. An agent decides on its own within a boundary someone drew beforehand: it notices a condition, picks an action, and executes it, without waiting to be asked.

Most of what reaches the market under the name "agent" is automation, sometimes an assistant. Rarely the third layer. In a fifteen-minute demo, with a rehearsed script, all three look exactly the same: someone asks, the system delivers something impressive, everyone nods.

The wrong kind of oversight is worse than too little

Here's the most expensive mistake Gartner itself documents: applying the same level of control to every system called an "agent," without distinguishing what each one actually decides on its own, is what most often kills AI projects inside companies. It isn't a lack of governance. It's governance in the wrong place.

A system that only suggests, and waits for approval on every action, doesn't need the same level of oversight as a system that executes on its own within a wide boundary. Treating the two as equally risky costs twice over: excess control on the first slows the team down without reducing any risk, because the risk was already low, and a lack of control on the second unleashes real autonomy under the supervision that would have only served the first.

The question that replaces "does this have AI?"

Before asking a vendor whether the product "has AI" or "is an agent," the question that separates the real thing from the repackaged one is different, and it has two parts. What does this system decide without consulting anyone? And what does it only execute because a person already decided beforehand?

The answer is rarely in the sales material. It's in asking to see, live, the system running into a situation outside the demo script, and watching who decides the next step when that happens.

What this puts on your desk

Pick a system your company already calls an "AI agent" and put this question in writing to whoever operates it: which decision does it make on its own, today, without anyone needing to approve it? If the answer is "none," you bought automation with a new name, and there's nothing wrong with that, as long as the price and the oversight are calibrated to what the system actually is.

The risk isn't buying automation thinking it's less than it is. It's supervising automation as if it were an agent, spending attention where it protects nothing, or letting a real agent run loose because the label said "assistant."

Where this piece's yardstick comes from

The question "what does it decide on its own, and what does it only execute because someone already decided" is the observable half of the same criterion that organizes the rest of this series.

The first axis is the reach of the error: automation that follows routine, an assistant that suggests, and an agent that decides within a boundary can't live under the same permission, because the cost of an error is a different order of magnitude at each layer. The second axis is the control left on the human side, which ranges from approving every action to only noticing the result once it's done.

The rule linking the two fits in one line: the control that remains has to be greater than or equal to the irreversibility of what the system can reach. Uniform governance ignores that rule by definition, because it treats different layers as if they were one.

The four reach bands, the ceiling for each one, and the questions that make the yardstick bite are published, with dates, on the method page.

The estimate that fewer than one hundred thirty of every thousand self-declared vendors are real agents, and the concept of "agent washing," come from Gartner (2025-2026). The finding that uniform governance increases the chance of agent project failure comes from the Gartner report "Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure," May 2026. The three-layer taxonomy (automation, assistant, agent) comes from the same set of publications.

This piece's yardstick is published, with dates

The four reach bands, the ceiling for each one, and the questions that make the yardstick bite. None of that is secret: what can't be copied is the practice of applying it case by case.

Read the method
Apply it to your own case