The Override

By Benjamin Evans

What happens while you wait

The AI flagged the transaction as suspicious. Unusual amount. New recipient. Velocity pattern matched known scam behavior. Confidence score: 87%.

The rep overrode it. The customer lost $4,000.

Last Tuesday, a different rep saw a different flag. Similar score, similar pattern. She overrode it too.

It was a grandmother sending rent money to her daughter. The flag was wrong. The override was right.

Same action. Same interface. Same button.

Opposite outcomes.

A button with a philosophy

Every AI product has an override moment — the point where a human looks at what the machine decided and says: no.

In financial products, it's the fraud flag a rep can dismiss. In healthcare, the diagnostic suggestion a doctor can reject. In content moderation, the automated removal a reviewer can reverse.

In almost every product I've seen, this interaction is a binary. Accept/reject. The machine speaks. The human answers yes or no.

This is the most consequential interaction in the entire system. And it's designed like a light switch.



The machine knows the pattern. The human knows the person. The interface lets only one of them speak.

Dylan Field told Lenny Rachitsky that AI "lacks the judgment" for these moments — but most products aren't designing for the judgment they still need.

Accept/reject flattens a complex judgment into a simple one. It asks "is the machine right?" without giving the human the information to answer well.

That first rep didn't have access to the model's reasoning. She saw 87% — a number that communicated certainty without communicating understanding. Yet What does 87% mean? That thirteen percent of similar flags are wrong? That this combination of features has historically preceded fraud 87% of the time?

The agent had context the model didn't: the customer sounded calm, had an explanation, felt normal. But "felt normal" is exactly the cognitive territory where scam victims are most vulnerable — socially engineered to sound calm, to have an explanation, to make the whole thing feel normal.

Both were operating on partial information. The interface forced one of them to win.

The override isn't where the human corrects the machine. It's where two incomplete views need to be reconciled. And we've given that reconciliation the interface of a coin toss.

A 2024 study in Computers in Human Behavior found that merely knowing advice is AI-generated causes people to follow it — even when it contradicts their own assessment. The Interaction Design Foundation names the consequence, summarizing Don Norman: "when behavior is frustrating or the system appears recalcitrant, the result is negative affect."

The system asked the human to exercise judgment, then gave them every reason not to.

The entire value of the override lives in one quadrant. The interface gives no help reaching it.

The hard design problem is the four-quadrant matrix:

Model right, human accepts. The system works.

Model right, human overrides. Dangerous. The human's judgment failed.

Model wrong, human accepts. Invisible. The failure is silent.

Model wrong, human overrides. The save. Why the override exists.

The entire value lives in that fourth quadrant. But the interface gives the human no help distinguishing it from the second — no help knowing whether their instinct to override is the right call or the dangerous one.

Gary Klein's recognition-primed decision model — built from studying how firefighters make life-or-death calls in seconds — shows that experienced professionals don't weigh options. They pattern-match. But Klein's model depends on rich situational information. The firefighter sees the structure, smells the smoke, feels the floor. The override interface gives the human a confidence score and a button.1

What would actually help: What did the model weigh? What didn't it have access to? What's the base rate — how often is the model right in comparable cases? What are the consequences, presented not as a legal disclaimer but as a human reality? If you override and the model was right, a person loses money they can't afford. If you accept and the model was wrong, a grandmother can't pay rent.

That's not a button. That's a conversation.


Rules tell people what to do. Frameworks help people make the right call when the pressure is on.

At Cash App, the design work spanned thirteen domains, and every one contained some version of this problem. In the dispute resolution flow, we had a system that flagged incoming disputes as potentially fraudulent. The model was good. But it also flagged a woman who'd been scammed out of $2,600 through a fake rental listing — and the flag made it easier for the reviewing agent to treat her like the problem.

The flag was technically correct. It matched fraud signals. But it didn't carry the context to distinguish a scammer exploiting the system from a victim using the system as designed. The agent had to make that call with a binary interface and a confidence score.

So we built decision frameworks. When confidence is high and stakes are high, block and explain. When confidence is moderate and stakes are high, warn and give the human a path to proceed with additional verification. When confidence is low, step back entirely. The frameworks didn't eliminate the hard calls. They gave the human a structure for making better calls faster.

The best frameworks don't tell people what to do. They help people make the right call faster when the pressure is on.

Almost no AI product is building this.

"The model isn't accountable. The product team isn't accountable. The person who pressed the button is."

Who is accountable when an override goes wrong? The human who pressed the button. The model made a recommendation, not a decision. The product team built a tool, not a mandate. Accountability flows to the person with the least power and the least information in the chain.

The people closest to the human impact — the rep, the moderator, the fraud analyst — bear the weight of decisions that were set up to fail by the system around them. Casey Newton's reporting at The Verge exposed this: content moderators at Cognizant developing PTSD, leading to a $52 million settlement from Facebook. Over a third score in the moderate-to-severe range for psychological distress.2 The system creates the conditions. The individual absorbs the consequences.

A 2025 review of automation bias — 35 studies, nearly 20,000 participants — found that non-specialists are the most susceptible to automation bias. Product School names the same tension: "AI handles analysis and execution speed, while humans provide judgment." But the people with the least expertise defer most to the system. And they're the ones on the front line.

Designing the override well means designing the organizational context around it. What support does the human have in the moment? Is the override data being used to improve the system, or to evaluate the human?

The frameworks we built at Cash App only worked because the team had psychological safety to override the system when their judgment said the system was wrong — even knowing some overrides would be mistakes. You can't build that through interface design alone. You build it through culture, through how you treat the person who made the wrong call for the right reasons.

The better the model gets, the less anyone listens to the human. That's exactly backwards.

Models are improving. Flags are getting more accurate. The natural response is to trust the model more and the human less — to narrow the override window, to add friction to disagreement.

This is rational in the short term. The model's aggregate accuracy probably is better across all cases.

But averages obscure the cases that matter most. A 2026 MIT study found that AI chatbots provide less accurate information to vulnerable users — the exact population most likely to need the override and least likely to have the power to exercise it.

Most products fail vulnerable users in silence. The customer doesn't complain. They don't file a ticket. They find another way. The product team never sees the data because they never built the instrument to capture it. Every narrowing of the override window makes the instrument weaker.

We design AI products as if the machine and the human are in competition. But the best outcomes come from the moment where each contributes what the other can't — where pattern recognition and contextual understanding are both present, both respected, and both visible in the interface.

We're not designing that yet.

We're designing the button.