SALES JUDGMENT CHALLENGE 009
THE SUMMARY
THAT BECAME
A SCORECARD
An AI assistant summarizes every sales call, and managers find it accurate. Then the summaries start rating deal health, those ratings feed the forecast, and leadership wants forecast results in quarterly reviews. Nothing about the tool has changed.
What would you do?
THE SITUATION
The summaries are accurate. The decision they now influence has changed.
Your company uses an AI assistant to summarize sales calls. It captures priorities, objections, next steps and decision-maker comments. Managers find it accurate and reliable.
Then managers start using the summaries to rate deal health. Those ratings feed the forecast.
Now leadership wants forecast performance to become one input in quarterly salesperson reviews. No one has found a problem with the accuracy of the summaries.
Nothing about the tool has changed. Only what it is being used to decide.
MAKE THE CALL
What should leadership do?
Choose before you continue.
WATCH THE CHALLENGE
Coming February 4.
The video for this challenge publishes February 4. Check back then — or work through the situation and questions below in the meantime.
THE JUDGMENT PROBLEM
Accurate and valid are not the same question.
This isn't "are the summaries accurate?" On everything we know, yes.
The harder question is this: an output can be accurate and still not be a valid measure of the thing you're now using it to judge.
Follow the output. The summary answers what was said on the call. Deal health asks how likely the deal is to close. The forecast asks what the team will deliver. The review asks how well this rep performed. Same output, four different questions.
What has actually been established is that the summaries capture the calls reliably. Whether they are a valid basis for judging deal health, and eventually rep performance, is a different question.
THE MEASUREMENT PROBLEM
The last step adds what the rep doesn't control.
A strong salesperson with a difficult set of accounts can produce weaker numbers than an average salesperson working stronger opportunities.
Territory quality, account mix, buyer timing and market conditions all shape forecast results. Unless the measure accounts for those differences, the number can confuse circumstance with performance.
THE JUDGMENT DIFFERENCE
Accurate is not the same as valid for this decision.
What was said on the call.
What does the score actually represent?
Nothing in the chain has to be wrong. The measurement can still be.
BETTER QUESTIONS
Before you put it in a review, ask:
What decision does this output now influence?
What does it actually measure: the rep, or the accounts the rep was given?
If the decision has changed, has the tool been evaluated for that decision?
This isn't about whether the summaries are accurate. It's about what they've been asked to measure.
SO, WHAT WOULD I DO?
Each option solves a different problem — and none solves all of them.
A limits the consequence: it caps how much the number can affect someone before it's properly understood, but it doesn't test what the number means. B checks the input: managers get a chance to challenge the summaries and the context, but an accurate summary doesn't validate everything inferred from it. C tests the measure: it asks directly whether the score says anything about rep performance. D can generate evidence, if the pilot is designed to compare the score against independent evidence of rep performance. A quarter of parallel data by itself tells you very little.
It depends doesn't mean anything goes. It means knowing what each safeguard does, and what it doesn't.
THE SALES JUDGMENT TAKEAWAY
A tool earns trust for a task. It doesn't earn it for the next one.
Ask what decision the output now influences.
Ask what it actually measures.
Don't ask only whether the output is accurate.
Ask what it actually measures.
BRING THE CHALLENGE TO YOUR TEAM
Make judgment visible.
Judgment Challenges can be explored with your team through a private 90-minute Judgment Lab, turning the scenario into a practical conversation about how a trusted AI tool ends up informing decisions it was never evaluated for, and what its output actually measures.