Every leader I know has, at some point, gotten a performance review that felt wrong. The facts were accurate. They just told an incomplete story. A manager inherits a strong team and looks brilliant. Another inherits a team in disarray, works twice as hard, and still comes out looking average. The scoreboard doesn’t know the difference, and neither, most of the time, does the person writing the review.

This is the oldest problem in people management. We measure leaders by outcomes they only partly control, then filter those outcomes through the opinions of whoever happens to be in the room. Every company says it wants an objective performance evaluation. Fewer have found a credible way to build it. Gallup found that just 2% of Fortune 500 CHROs strongly agree their performance management system actually inspires employees to improve. The fix always seems to require better opinions, and opinions are exactly the thing that’s unreliable in the first place.

I found this problem being solved somewhere I didn’t expect: professional football. I recently talked with James Hasty, a former NFL cornerback who spent 13 seasons in the league and then nine years building Coaching AI, a system that measures coaching performance objectively. His premise is simple to state and hard to build. A coach’s grade should reflect the difficulty of what he was actually given, not a league-wide average. A coach with a thin roster and a coach with a stacked one are not doing the same job, even when their win totals look similar, so grading them on the same scale has never made sense. What Hasty built instead compares each coach against the specific opposing talent his players faced, which means a good coach with a bad roster finally looks different from a bad coach with a good one.

The parallel to corporate leadership isn’t subtle, and Hasty draws it directly. Here’s what he told me.

1. Fix the yardstick before you fix the review

The most common evaluation mistake, Hasty said, is measuring everyone against the same flat average. “When wins are the only input, coaches who inherit great rosters typically are overvalued, due to their lack of experience in high-pressure situations.”

His system corrects for that by comparing each coach against the specific difficulty of what he was actually given, not a league-wide baseline. He draws the corporate parallel himself: “The equivalent of a coach with a bad roster is an executive with a broken system and poor culture.” That executive rarely gets full credit for stabilizing things, because the scoreboard only shows where the numbers ended up, not where they started.

2. Watch the decision, not just the scoreboard

Fixing the yardstick only helps if you’re also looking at the right thing. Hasty said the leaders doing the hardest, smartest work often get missed entirely, because their skill shows up in decision quality, not the final score. He pointed to coaches with varied position experience, across special teams, linebacker, defensive back and coordinator roles, who bring a kind of judgment a raw record can’t capture. Hasty’s Coaching AI system tries to observe that judgment directly, by analyzing not only play-calling, but also speech, evaluating how a coach communicates under pressure.

He described the goal as observing leadership decisiveness, assessing teaching clarity, and interpreting vocal tone in real time. His recommendation for any CEO trying to fix manager evaluation is to stop scoring managers on the outcome their team inherited, and start scoring the quality of the decisions they made along the way, especially under pressure and in chaos.

3. Know exactly where measurement should stop

The most useful thing Hasty said was about what not to measure. “Objective measurement should stop right before it starts to quantify the parts of coaching that are fundamentally human, relational, and emotional.”

Empathy, moral courage, and culture-building, he said, are visible in how someone behaves when no one is watching, and they have to be interpreted, not scored. That’s a discipline most performance-management systems lack entirely — the confidence to build a rigorous metric and the restraint to know where it shouldn’t reach.

The deeper lesson isn’t really about football or about AI. It’s that the fix for subjective evaluation was never a search for better opinions. It’s a comparison that accounts for what the leader was actually handed, paired with the judgment to know which parts of leadership shouldn’t be reduced to a number in the first place. Sports figured that out on a scoreboard. The rest of us are still working it out in a conference room.