RevOps & Stack
Falsifiable forecasts: a dated, testable prediction for every deal
A 70 percent confidence score can never be wrong, which is why it never gets better. What deal forecasting looks like when every claim can be tested.
By Rishi Patel, Founder & CEO, RevSage.ai · · 7 min read
Sit in enough forecast calls and you start noticing that nobody is ever wrong in them. A rep says a deal feels like a 70. The quarter ends, the deal doesn't, and the number quietly becomes a 40 in next month's sheet, which also turns out to be defensible whatever happens. Every score was reasonable. No score was ever falsified. Nothing was learned, all year.
The problem has an old and well-worn name. A claim that no possible observation can contradict is unfalsifiable, and unfalsifiable claims don't improve with experience, because experience has no way to grade them. "This deal is 70 percent likely to close" on a single deal is exactly such a claim. Win: the 70 was right. Lose: three out of ten, what can you do.
I want to make the case for replacing deal-level percentages with something older-fashioned and much more demanding: a dated, testable prediction, one per deal, written so a specific future observation can prove it wrong.
Key takeaways
- A percentage on one deal is compatible with every outcome, so it can never be graded, so the judgment behind it never gets better. That's a structural problem, no dashboard fixes it.
- A falsifiable prediction has four parts: an observable event or non-event, a deadline, an explicit falsification test, and the evidence it rests on.
- Anchor the deadline to the deal's own last interaction, never the fiscal calendar. Quarter-end anchoring lets stale deals borrow credibility from a date that has nothing to do with them.
- Expired predictions are the agenda: grade last review's claims before making new ones. Hit rate makes calibration a discussable number.
- Prediction-writing also changes behavior before the deadline, because a specific claim tells the rep exactly what to go verify this week.
The comfort of numbers that can't lose
Percentage scores survive because they serve everyone in the room except the forecast.
For the rep, a 70 communicates optimism while committing to nothing checkable. For the manager, scores roll up into a weighted pipeline number that looks like arithmetic and reads like rigor. For the eventual post-mortem, every score turns out to have been consistent with what happened. It's a currency everyone accepts precisely because it never has to be redeemed.
The cost is invisible because it's an absence: no error signal, ever. A rep who's been systematically wrong about which deals were real for two straight years can have a defensible score history the whole way down. Nothing in the process was capable of catching it. Aggregate calibration across many deals could catch it in principle, but almost nobody runs that analysis on rep-entered scores, and the scores drift with mood and quarter pressure anyway.
I wrote before about what confidence scores from AI sales tools actually tell you, and the argument here is a cousin of that one: a number without a stated test is an opinion wearing a costume.
What a testable deal prediction looks like
Here's the shape, four parts, one sentence if you're disciplined:
An observable event or non-event. Something a third party could verify. The buyer signs the pilot agreement. The CFO joins a call. The buyer sends written confirmation of scope. Or the negative: no written response to the proposal will exist. Observables only: no "momentum," no "alignment," nothing that lives inside anyone's head.
A deadline. A real date. The claim is about the world by that date.
A falsification test. Spelled out: what evidence, existing by the deadline, proves this prediction wrong? If you can't write this clause, the prediction isn't one.
The evidence it rests on. One or two lines: what, in the record of this deal, makes you believe it. Quotes and dates beat vibes.
Assembled: "By September 12, the buyer will have confirmed pilot scope in writing. Falsified by: any written scope confirmation not existing on the 12th. Basis: she committed verbally on the August 28 call and asked for the SOW the same day."
Notice what writing this does to you, before the deadline ever arrives. The claim exposes exactly what you don't know. If your prediction rests on a verbal commitment, you suddenly care about converting it to writing this week. The prediction becomes a to-do list for its own verification. Scores never did that for anyone.
Anchor to the deal's pulse, never the quarter
One design decision matters more than it looks: where the deadline comes from.
The natural habit is fiscal. Will it close this quarter? But quarter-end anchoring has a flaw that corrupts everything downstream: it hands every deal the same clock regardless of the deal's condition. A deal with no buyer contact since March can still be "closing this quarter," and the phrase sounds like a plan. The calendar lends dead deals its own credibility.
Anchor instead to the deal's most recent real interaction. The prediction window opens at the last substantive buyer touchpoint and runs a fixed length, 14 days is a good default for mid-cycle B2B, and the claim must name something observable within that window.
This sounds like a small bookkeeping choice. Watch what it does to a stale deal. Last contact eleven weeks ago, and the prediction must name an observable event within 14 days of... eleven weeks ago. The window already expired, empty. The deal can't hide inside "this quarter" anymore, because it never got to borrow the quarter's clock. Deals age in their own time, and the forecast is forced to notice. The lagging-indicator problem I described in why healthy deals die is largely this same disease in CRM form: fields that reference the rep's hopes instead of the buyer's behavior.
A useful side effect: fixed-length windows make predictions comparable across the pipeline. Every claim is the same size. Hit rates mean something.
Grading day is the whole point
A falsifiable prediction pays out on the day it expires, and only if someone actually grades it.
So restructure the pipeline review: expired predictions first, new predictions second. For each expired claim, three possible grades. It held. It was falsified. Or, the most instructive grade, it couldn't be graded, because the claim turned out to be mushier than it looked when written. Mushy claims are how you find out who's still writing scores in sentence form.
Then the numbers you actually want start to accumulate. A rep's hit rate. A team's hit rate on "the buyer will do X" versus "the buyer won't do X" claims, the second kind turns out honest more often, which is itself a finding. Which evidence types, verbal commitments versus written ones versus meeting behavior, actually predicted. This is calibration you can discuss, coach, and improve, and it comes from the same hour the forecast call was already burning.
Two ground rules keep it honest. Falsified predictions carry zero shame, they carry information: something believed about this buyer was wrong, and now the deal strategy knows it. And nobody gets to revise a prediction mid-window. Amended claims are new claims, graded separately. The moment editing history is allowed, you've rebuilt the 70-that-becomes-a-40, with extra paperwork.
Where the evidence comes from
The weakest link in all of this is the basis clause. A prediction grounded in "seems engaged" is a score with a deadline stapled on. The basis has to come from the record: what the buyer said, did, confirmed, avoided, with dates.
Which means the discipline is only as good as your access to that record. If reconstructing what the buyer actually committed to requires re-listening to four calls, nobody will do it weekly, and the basis clauses will quietly rot back into vibes.
This is the layer we built RevSage for: it reads every call and thread in the deal and produces exactly this artifact, a dated, falsifiable prediction per deal, each with its falsification test spelled out and its evidence chain quoted and timestamped, as part of live root-cause analysis across the pipeline. The window anchors to the deal's own last interaction, never the quarter. When the evidence is thin, the prediction says so, which is itself a flag worth having.
But I want to be careful about the pitch here, because the practice stands without any tooling at all. One sentence per deal, four parts, graded on expiry. A team that does this with a spreadsheet for two quarters will know more about its own judgment than years of weighted pipeline ever taught it.
The deeper shift is cultural, and I'll admit it's the part I care about most. Percentages let a pipeline review stay comfortable indefinitely: everyone reasonable, nobody wrong, nothing learned. Predictions make the review briefly uncomfortable and permanently useful, because being wrong in a specific, dated, recorded way is the only version of wrong that teaches anything. The forecast stops being a mood report and becomes an experiment log. That trade costs some pride the first month. It returns judgment, compounding, for as long as you keep grading.
Frequently asked questions
- What is a falsifiable sales forecast?
- A forecast statement specific enough that a defined future observation can prove it wrong: a named event or non-event, a deadline, and an explicit description of what would disprove it. By this window's end date, the buyer will have confirmed the pilot scope in writing, falsified if no written confirmation exists by then. Compare that with 70 percent likely, which no outcome can ever contradict.
- What is wrong with percentage confidence scores on deals?
- A percentage on a single deal is compatible with every outcome. Win at 70 percent, the score was right. Lose at 70 percent, well, three in ten. Because no result can contradict it, the forecaster never receives a clean error signal, and judgment doesn't improve. Percentages also smuggle optimism: they feel precise while committing to nothing.
- Why should deal predictions be anchored to the last buyer interaction instead of the quarter?
- Quarter-end anchoring lets stale deals borrow credibility from the calendar: closing this quarter sounds like a plan even when nothing has happened for six weeks. Anchoring to the deal's own last touchpoint, something observable within 14 days of the last real interaction, forces every prediction to reckon with the deal's actual pulse rather than the fiscal one.
- How do you start using falsifiable predictions in pipeline reviews?
- Pick the deals past the midpoint of your pipeline. For each, have the owner write one sentence: what observable thing happens or fails to happen by a specific date, and what evidence would falsify the claim. Review the expired predictions at the start of the next session before making new ones. Hit rates make calibration discussable in a way scores never do.
About the author
Rishi Patel, Founder & CEO, RevSage.ai. Rishi has spent 11 years building and scaling B2B SaaS companies, most of it obsessing over why some reps consistently read buyers right and most don't. He founded RevSage to give every rep the buyer intuition of their best teammate.