← Articles
August 4, 2026 · 6 minute read

The Ladder Has a Trapdoor

Every AI maturity ladder promises that climbing means improving. The company that invested the most in it found out, in practice, that this isn't always true.

Every AI maturity ladder out there tells the same story. Five, seven, eight rungs, and the promise is always that climbing means improving. At the bottom, the manual company. At the top, the adaptive company, where agents work within boundaries and humans handle what matters.

It's a good design. And it has a problem none of those diagrams show: the rung above isn't always better than the one below. Sometimes it's just more expensive, more fragile, and harder to undo.

If you want to see this happening, you don't need to look for a company that's behind. Just look at the one that invested the most.

What happened at the company that spends the most

Amazon set a target for more than 80% of its developers to use AI tools every week. To track it, it built an internal leaderboard ranking people by usage.

Adoption went up, as you'd expect. And something else went up with it: engineers started firing off agents on low-value tasks just to improve their own ranking. The practice got an internal nickname, tokenmaxxing, and produced an effect nobody had designed for: the leaderboard raised the company's costs.

The leaderboard got shut down. And the executive responsible had to say, out loud, a sentence that sums up the whole problem: don't use AI just for the sake of using AI.

Worth noting: the company didn't give up on measuring. It swapped the token count for a measure of work actually delivered with AI. The fix wasn't to stop measuring. It was to stop measuring the input it was paying for.

And it's not one company's anomaly. Another giant in the sector ran an equivalent ranking, with roughly eighty-five thousand people ranked by usage, also pulled down once it made the news. One market analyst put it better than any consultant could: you get the behavior you built the incentive for.

Before that, two incidents. In December 2025, an agent deleted an infrastructure environment on its own, and an entire region went down for thirteen hours. In March 2026, six hours of failures in checkout, login, and pricing, with more than twenty-one thousand user reports.

The fix adopted after the second incident is the part that matters: junior and mid-level engineers now need a senior's sign-off before pushing AI-assisted code to production.

Read that again. The answer wasn't more AI. It wasn't a better model. It was bringing human approval back into the middle of the process.

The trapdoor

On the ladders out there, human approval at every step is a sign of immaturity. It's rung two, rung three, the place you're supposed to leave behind.

The most AI-intensive company on the planet climbed, hit the ceiling, and turned back. Not out of incompetence, but because it discovered in practice what the diagram doesn't show: the right rung isn't the highest one the technology can reach. It's the highest one the reach of its error can afford.

An agent that drafts a memo and an agent that executes a half-million-dollar transaction can't live under the same permission, the same monitoring, and the same off switch. That sounds obvious written out like this, and it's still the thing people get wrong most, because the maturity ladder has no column for the size of the damage.

What the metric taught

There's a second lesson, and it's more uncomfortable than the first, because it isn't about technology.

That leaderboard didn't fail because it was poorly built. It did exactly what a leaderboard does: people optimize for what's measured. If the metric is usage, the rational behavior becomes using — including where using doesn't help.

Amazon has sixteen leadership principles, written, public, and lived for decades. One of them asks people to operate at all levels, stay connected to the details, and audit frequently. Another says to accomplish more with less, because constraints breed resourcefulness.

A token-usage leaderboard contradicts both at once. And it won, for as long as it existed.

The conclusion isn't that the culture was weak. It's that no set of values is self-enforcing. A solid culture loses to a daily metric, every day, at any company, because the metric is what shows up in the review and the culture is what's on the wall. When the leaderboard was shut down, the principles started mattering again — not because anyone rewrote them, but because they stopped competing with a number.

What this demands of whoever decides

If you lead a department that's adopting AI, three questions are worth more than any ladder:

What AI metric have you put in place, and which of your own values is it contradicting right now? If the answer is "none," it's worth checking what the metric rewards when nobody's watching.

Could whoever approves what the AI produces have produced it themselves? If they couldn't, approving isn't examining. It's signing. And the difference between the two only shows up on the day of the incident.

Was the gain you already cut from the budget measured, or is it still a promise? Cutting capacity against a projected gain is an easy decision to make and a slow one to undo, and it shifts onto people the cost of a bet the company made.

The right rung

The merit of that fix isn't that it avoided the mistake, because it didn't. It's that it detected it and undid it. Two reversals in a few months: the leaderboard shut off, human approval brought back.

That's the most underrated thing in this entire conversation. The question that matters isn't how far AI can go. It's how far back you can go if it goes too far.

A ladder without a trapdoor doesn't exist. What exists is knowing which rung you're standing on, and whether it can hold your weight.

The ideas in this piece have names

I wrote the piece above without jargon, on purpose, because the argument holds without it. But the ideas didn't appear here for the first time, and it's worth saying where they come from.

The reversibility yardstick is what's behind "the right rung is the highest one the reach of its error can afford." It classifies a decision by the reach of the damage: R1 draft, R2 internal, R3 operational, R4 external effect. And the autonomy ceiling stops being fixed and becomes conditional on that classification, decision by decision, instead of one policy applying to the whole company.

Remaining Control is the other half, and it answers how much is left on the human side: C4 examine case by case, C3 authorize in batches, C2 having designed the boundaries without seeing individual cases, C1 only reviewing the record afterward. The rule linking the two is short: the control that remains has to be greater than or equal to the effect's irreversibility.

Approval theater is the name for what happens when someone signs off without examining, and the piece above shows the organizational version of it: a company that reports operating two or three rungs above where it actually operates.

These concepts are original development, built between July and August 2026, on top of the AI adoption yardstick Mike Taylor and Every published in 2024 and 2025. They're the foundation of the method I use with executives, and this piece is the first time they've appeared in public, with a date.

If you want to apply the yardstick to your own case, it fits in one question: what's the reach of this process's error, and how much control is still left on your side?

The facts cited come from published reporting. Primary source: Rafe Rosner-Uddin, "Amazon scraps AI leaderboard to stop workers chasing usage scores", Financial Times, May 28, 2026, republished by InfoWorld, CIO, Tom's Hardware, heise, and Fortune. Leadership principles as published by the company itself.

This piece's yardstick is published, in full

The four reach bands, the ceiling for each one, and the questions that make the yardstick bite. None of that is secret: what can't be copied is the practice of applying it case by case.

Read the method
Apply it to your own case