The Decision Audit Prompt
Built By Evolution — Decision Science Field Note
The five-question Decision Audit works as a mental checklist. It works better as a prompt. Below is one prompt template that runs the audit in either Claude or ChatGPT, followed by a worked example so you can see what a real run looks like before you try your own.
If you haven't seen the framework itself yet, start with the companion piece, "Why An Accurate AI Prediction Still Isn't Enough?", or the video it's built from: Watch on YouTube.
The prompt
Copy everything in the box below, paste it into a new Claude or ChatGPT chat, and fill in your own case at the bottom. It's plain natural-language instructions — no special syntax — so it runs the same way in both.
You are running a Decision Audit on an AI-generated prediction, score, or recommendation. Use the following framework, adapted from the "anatomy of a decision" in Prediction Machines (Agrawal, Gans, Goldfarb): Input → Prediction → Judgment → Action → Outcome → Feedback I will give you the AI prediction itself, the decision it's meant to inform, and whatever context I have. Do not just restate or validate the prediction. Audit it by answering these five questions, in order, using only the specific case I describe below — no generic disclaimers: 1. INPUT — What information does this prediction appear to be based on, and what relevant context, history, or exception might be missing from it? 2. ACTION — What action follows from this prediction, and who (or what) can review, override, or stop that action before it takes effect? Given how reversible and how high-stakes the action is, should this prediction be allowed to Recommend, get Reviewed, or be Automated? 3. OUTCOME — What is the real goal here? Is the prediction measuring that goal directly, or a proxy for it? Name the gap if one exists. 4. JUDGMENT — What does a false positive cost here, versus a false negative? Are those costs roughly equal, or is one much worse? What does that imply about how the prediction should be used? 5. FEEDBACK — After this decision is acted on, what result should be tracked, who should see it, and how would it improve the next prediction? Finish with a one-paragraph verdict: how much authority should this specific prediction get, and what's the single biggest blind spot in relying on it as-is? Here's the case: [PASTE YOUR AI PREDICTION / SCORE / RECOMMENDATION AND CONTEXT HERE]
How to use it
- Swap in your own case. Replace the bracketed line with the actual number or recommendation an AI gave you, plus a sentence on what decision it's feeding into. The more specific, the sharper the audit — "our fraud model flagged this $4,200 transaction, 2 days before the customer's trip" beats "an AI flagged a transaction."
- Run it once as-is. Don't pre-answer the five questions yourself first — let the model surface what it can and can't infer from what you gave it. A vague or overconfident answer on Question 1 (Input) is itself useful information: it tells you the model doesn't have enough context either.
- Push on the verdict. If the model lands on "Automate" for something that affects a person's job, credit, or health, ask it to defend that call against the Judgment question specifically. That's usually where a too-easy verdict falls apart.
Worked example
The case, as given to the model:
Our hiring platform gives Candidate A a score: "87% predicted likelihood of success in this role," based on resume-derived skills match and performance data from past hires in similar roles. The hiring manager is deciding whether to extend an offer, run one more interview round, or pass.
The audit that came back:
1 · INPUT
The score is likely built from structured, comparable signals: parsed résumé skills, years of experience, and how similar profiles performed after being hired. It almost certainly does not include anything from an interview, a reference call, or a work sample — the things most likely to surface a mismatch that a résumé wouldn't show. Treat the 87% as "87% based on paper," not "87% based on the candidate."
2 · ACTION
The action isn't binary — it's a three-way fork: offer, interview again, or pass. That's actually a good sign; a system that lets a human insert one more information-gathering step before the irreversible action (the offer) is already closer to Review than to Automate. The audit for this case: keep it there. Don't let the score skip the interview round on its own — use it to prioritize who gets that round sooner, not to decide who skips it.
3 · OUTCOME
The real goal is "a good hire" — someone who does the job well and stays. The score almost certainly measures a proxy: something like "still employed and hitting targets at 12 months" in the training data. That's a reasonable stand-in, but it silently drops anything the past-hire data didn't capture — team fit, growth trajectory beyond 12 months, reasons a past hire left that had nothing to do with skill. The gap between "scored well on the proxy" and "was a good hire" is exactly where a false 87% would hide.
4 · JUDGMENT
A false positive here — hiring someone the score liked who doesn't work out — costs a bad quarter, a rehire, team disruption. A false negative — passing on someone who would have excelled because the score underrated them — costs a candidate you'll likely never find out you lost, since there's no feedback loop on people you didn't hire. That asymmetry matters: the model's false negatives are structurally invisible to you, which means "87% accurate" is quietly easier to believe than it should be.
5 · FEEDBACK
Track what happens to this hire at 6 and 12 months against what the score predicted — that closes the loop for future candidates scored the same way. Separately, and this is the harder discipline: track a sample of people the score ranked lowest but who got hired anyway (through a referral, say), to get any signal at all on the false-negative rate. Without that second loop, the system only ever learns from the candidates it said yes to.
Verdict. This prediction has earned a place in the Recommend / Review lane, not Automate — use the 87% to prioritize and triage, never as a standalone pass/fail gate, and keep a human-run interview between the score and the offer. The single biggest blind spot: the score is measuring "performed like past successful hires in this role," which is a proxy for "will be a good hire," not the thing itself, and that gap is invisible from the number alone.
Try it on your own case
The same five questions run the same way whether the AI gave you a hiring score, a fraud flag, a churn-risk number, a medical triage score, or a chatbot's confident answer. Paste the prompt, swap in your case, and see where the verdict lands — Recommend, Review, or Automate.
Built By Evolution — Decision Science. Understand More. Decide Better. Live Wiser.
Comments
Post a Comment