Session
What Actually Predicts a Good Hire: Evidence From 486 Million Job Records
Does public career history predict on-the-job performance? We tested it across 486 million job records. This talk gives the measured effect sizes, the three hypotheses we falsified, and how to build a scoring system that stays honest when the signal is small.
Quality of hire is the metric every talent organization says it wants and almost none of them can deliver. Previous research has uncovered the connection at the corporate level: analysis of 4.35 million hires into Fortune 500 organizations in 2020–2025 demonstrated that 1 point of improvement in the hiring quality index of a company correlated with an increase in revenue growth year over year by 0.14%. And the next logical step is to see if it works for the single hire.
We have spent a year exploring this question at individual-level granularity, using a database comprising around 96 million public career profiles and 486 million job records. In creating the feature layer, we received the first clear and sustainable evidence: at this scale, rule-based normalization with a model backup performs better in terms of cost and coverage compared to per-row LLM classification; and we will give the numbers to support our choice. The talk covers our findings, including those that failed.
The straightforward conclusion from our study: career trajectory data provides significant signals about job performance in the future, and the size of it is relatively low.
The correlation with employer performance ratings was r = 0.087 (AUC 0.540, n = 3,878). The discrimination of hired from non-hired candidates in a natural experiment was AUC 0.560 (n = 6,301). Departmental context in a supervised model brought up the performance to AUC 0.665, while within-department normalization resulted in perfect monotonic discrimination across quintiles, with a 13.4 point spread between first and fifth quintile (n = 3,499, p < 0.001).
We will also take a look at the negative results that we have learned from: an industry-specific model that ended up being a dead-end, a forward-looking variant of the model that never made it past our threshold, and the training population coverage problem that had quietly broken our model for senior candidates all along.
And then the engineering problem that arises from all of it: What exactly do you build when your model is real but flawed?
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top