A hockey stat sent me down a rabbit hole this weekend 🤓
It started with goalie trivia and ended with manager evaluation.
The trivia: you can’t judge a goalie on save percentage alone - a goalie behind a leaky defense faces harder shots and can look worse than he is. So analysts built models that score every shot’s observable difficulty - distance, angle, shot type, rebound/previous-event context, manpower, and in richer datasets, pre-shot movement - and predict the probability it becomes a goal. Sum the differences between expected and actual outcomes, and you’ve partially separated goalie performance from shot volume and shot quality (Goals Saved Above Expected).
Then I learned this same basic logic has been used in education for ~30 years. It’s called Value-Added Modeling. Predict each student’s expected test score from prior achievement, demographics, peers, and other available context - then estimate the teacher effect from the remaining classroom-level difference.
The People Analytics application is obvious: strip out tenure, role, comp, market, etc. - and the residual may contain a manager effect.
Tempting, but also dangerous. What VAM taught education the hard way:️
- Single-year estimates are noisy: year-to-year stability for individual teachers is often around 0.2-0.5, sometimes higher depending on grade, subject, and model.️
- Three years of data materially improves reliability - patience beats better models - but it doesn’t eliminate bias.️
- When test-based metrics became high-stakes, they created predictable gaming: more test prep, narrowed instruction, and in the broader test-based-accountability era, outright cheating scandals like Atlanta`.️
- The American Statistical Association issued a formal caution in 2014 about overinterpreting VAMs for high-stakes individual decisions.️
- Houston teachers later won a favorable due-process ruling and settlement because the proprietary model/data behind their scores could not be meaningfully challenged.
Every one of those lessons applies to manager scorecards. The model can be useful, but the governance around it is where things usually break. I can see this working for surfacing patterns, validating training programs, and allocating coaching - but going wrong at the predictable next step: using it to gate promotions or trigger PIPs on thin data, which is where many orgs will be tempted to go first, for understandable reasons.
Question for the PA folks here: Are you using “above expected” style models in PA today? For what - retention, hiring, sales, something else? If you tried and walked away, what was the deal-breaker? Sample size? Politics? Legal? The metric got gamed?
Related notes
- The Triple-Filter Test: How to prioritize HR interventions with panel data
- Unexpected protective effect of having a good manager?
- Want to maximize your impact as a leader?
- Is contagious turnover overrated? Probably only if you ignore the managers.
- Novel way to measure leadership skills via causal inference (and AI)?
📄 Read the original post with full outputs on my blog.