Why An Accurate AI Prediction Still Isn’t Enough?

Built By Evolution — Decision Science Field Note

An AI hands you a number. Before you let it make the call, run it through the five questions a prediction can't answer for itself.

7 min read · Source: Prediction Machines — Agrawal, Gans & Goldfarb · Companion video below


The anatomy of a decision

Input Prediction
(AI)
Judgment Action Outcome Feedback

AI is compressing one box in this loop — prediction — into something fast and cheap. The other five didn't get automated. They got quieter, which is easy to mistake for solved.


An AI says a job candidate has an 87 percent chance of succeeding. Do you hire them?

You can't actually answer that from the number alone — not because the prediction is wrong, but because a prediction isn't a decision. It can't tell you what success is worth, what failure costs, or who answers for it if the number turns out wrong.

That's the trap sitting inside almost every AI tool you touch. A hiring score, a medical flag, a maps ETA, a chatbot's confident answer — each one is a prediction wearing the costume of a decision. The five questions below catch the difference. Each one targets a part of the decision that prediction alone was never built to settle. Run them on the next number an AI hands you, whether the stakes are a hiring committee or a dinner recommendation.

Five questions, one for each part prediction can't settle

They follow the loop above in order — skipping only the box AI already handles.

Q1 · INPUT
What is this prediction actually built from?

Every prediction is a summary of the information fed into it, and it is silent about everything left out. A hiring model trained on resumes and interview scores has never met the candidate's references, never watched them handle a hard week, never heard the context behind a résumé gap. The context, the history, the exception a person in the room would normally catch — none of it is in the number, because none of it was ever in the input.

Before trusting a prediction, ask what it saw, and just as important, what it structurally couldn't have seen.

A maps app predicts the fastest route from historical traffic patterns. It has never seen this morning's accident three blocks ahead.

Q2 · ACTION
What happens because of this, and who can stop it?

A prediction only matters once it's wired to an action. Ask what actually changes because the number came back this way, and who, or what, is positioned to intervene before that action takes effect. A model that flags a transaction but leaves a human free to review it before anything moves is a fundamentally different system from one that blocks the card automatically, even at identical accuracy.

If nothing can be stopped and no one is positioned to stop it, the prediction has more authority than it earned. That authority should track the stakes: safe to automate when the action is low-stakes and reversible, worth a human review pass when it's noticeable but correctable, and something that should inform a person rather than replace one when it's severe or hard to undo.

A streaming pick that's wrong costs you ten minutes. A parole-risk score that's wrong costs someone years.

Q3 · OUTCOME
What are we actually trying to create, and is that what's being measured?

This is where a proxy quietly stands in for the real goal. "87 percent likely to succeed" only means something once you answer a question the number itself can't: succeed at what? If the model is scoring "still in the role and hitting targets at twelve months," that's a real, usable prediction, and a narrower one than "will be a good hire."

Mistaking the proxy for the goal is a judgment failure, not a prediction failure, and it's exactly the kind of gap a perfectly accurate model will never flag for you. The model was never wrong about what it was measuring. It was just measuring something smaller than what you meant.

"Clicked the video" is a proxy for "found the video worth their time." A recommendation engine can be excellent at the first and terrible for the second.

Q4 · JUDGMENT
What does each kind of mistake actually cost?

Every imperfect prediction fails in one of two directions. A false positive says something's there when it isn't. A false negative misses something that is. Those two failures are rarely priced the same: a song recommendation that misses costs you a skip; a medical flag that misses costs something you can't get back.

Accuracy alone can hide which direction the mistakes are even happening in. If a real emergency shows up in one case out of a hundred, a model that says "no emergency" every single time is already 99 percent accurate, and catches nothing. "Mostly accurate" means nothing until you've priced the false positive and the false negative separately — the stated accuracy was never built to tell you that on its own.

Two models can both be 90 percent accurate. One is picking a playlist. One is flagging an ER case. They are not equally safe to trust.

Q5 · FEEDBACK
Does what actually happens make the system smarter?

A decision system isn't finished when the action is taken. The real outcome has to make it back to whatever produced the prediction, or the next prediction is no better than this one. Ask what result comes back, who actually sees it, and whether it changes the model, the threshold, or the process next time.

A dashboard nobody checks, a rejected candidate nobody follows up on, a flagged case that gets overturned but never logged — each of those is a feedback loop that was built and then left open. A process with no path back from outcome to input isn't a trustworthy decision system, no matter how good the model looked on day one.

A spam filter gets smarter every time you mark something "not spam." A hiring model only gets smarter if someone tracks what happened to the candidates it rejected.

Back to the candidate

87 percent predicted to succeed.

Run the five questions and the gap becomes visible. Which outcome was actually predicted? What action follows, and who can catch it before it lands? How costly is rejecting someone who'd have excelled, against hiring someone who struggles? Can the decision be reviewed? Will what actually happens improve the system, or just confirm what it already assumed?

The prediction was useful. It was never the decision.

And this isn't only a hiring-committee problem. The same five questions apply every time an app names the fastest route, a streaming service names what to watch next, or a chatbot gives you a confident-sounding answer, even when the stakes are small enough that you'd never think to ask. That's the actual habit worth building: not a rule for one dashboard, but a question set for every number a machine hands you.


Watch the full breakdown

AI Doesn't Make Decisions: It Makes This One Part Cheaper

The video version of this piece — the London black-cab story, why accuracy isn't enough, and where AI should actually sit in your decisions. ~10 minutes, Built By Evolution.

▶ Watch on YouTube


Built By Evolution — Decision Science
Adapted from the "anatomy of a decision" framework in Prediction Machines: The Simple Economics of Artificial Intelligence by Ajay Agrawal, Joshua Gans, and Avi Goldfarb. Understand More. Decide Better. Live Wiser.

Next question: if AI makes prediction cheaper, does human judgment get more valuable? That's the next decision.

Comments

Popular posts from this blog

Tesla vs. Waymo: The Sensor Bet That Decides Driverless Reliability

An AI Boss Fired a Human—but It Wasn't Acting Alone

Welcome to Built By Evolution: Every Story Hides a Principle