Tobias Canavesi
Data Scientist at Crosschq
Actions
Tobias Canavesi is a theoretical physicist turned Senior Data Scientist focused on explainable, production-grade AI. At Crosschq, he builds the ML systems behind Quality-of-Hire measurement: LLM- and embedding-based candidate ranking, multi-agent pipelines in LangGraph, fraud and scam detection in hiring, and forecasting models that connect pre-hire signals to post-hire outcomes. His analysis of 432,000 hires anchors a public writing series on what résumés can and can't predict. He recently joined Crosschq's Science team to apply ML to personality and cognitive assessment.
Previously, he built a forecasting pipeline for 7,000+ SKUs and GenAI e-commerce chatbots at Heyday/Essor, and led data-science projects at Accenture across mining, airlines, and beverages, including a Chilean mining initiative that increased water recovery by ~10% (~25,000 tons/day).
He holds a PhD in Theoretical Physics (UNLP), with publications in JHEP and Classical and Quantum Gravity, teaches in the AI Specialization at the University of Buenos Aires, and spoke at AI DevWorld 2026.
Area of Expertise
What Actually Predicts a Good Hire: Evidence From 486 Million Job Records
Does public career history predict on-the-job performance? We tested it across 486 million job records. This talk gives the measured effect sizes, the three hypotheses we falsified, and how to build a scoring system that stays honest when the signal is small.
Quality of hire is the metric every talent organization says it wants and almost none of them can deliver. Previous research has uncovered the connection at the corporate level: analysis of 4.35 million hires into Fortune 500 organizations in 2020–2025 demonstrated that 1 point of improvement in the hiring quality index of a company correlated with an increase in revenue growth year over year by 0.14%. And the next logical step is to see if it works for the single hire.
We have spent a year exploring this question at individual-level granularity, using a database comprising around 96 million public career profiles and 486 million job records. In creating the feature layer, we received the first clear and sustainable evidence: at this scale, rule-based normalization with a model backup performs better in terms of cost and coverage compared to per-row LLM classification; and we will give the numbers to support our choice. The talk covers our findings, including those that failed.
The straightforward conclusion from our study: career trajectory data provides significant signals about job performance in the future, and the size of it is relatively low.
The correlation with employer performance ratings was r = 0.087 (AUC 0.540, n = 3,878). The discrimination of hired from non-hired candidates in a natural experiment was AUC 0.560 (n = 6,301). Departmental context in a supervised model brought up the performance to AUC 0.665, while within-department normalization resulted in perfect monotonic discrimination across quintiles, with a 13.4 point spread between first and fifth quintile (n = 3,499, p < 0.001).
We will also take a look at the negative results that we have learned from: an industry-specific model that ended up being a dead-end, a forward-looking variant of the model that never made it past our threshold, and the training population coverage problem that had quietly broken our model for senior candidates all along.
And then the engineering problem that arises from all of it: What exactly do you build when your model is real but flawed?
OPEN Session: An Explainable AI System for Applicant Ranking with LangGraph + Hybrid Embeddings
Hiring teams drown in resumes while job requisitions vary wildly in structure. This talk unveils a production-grade Applicant Ranking AI Suite that parses job descriptions and resumes into ontology-aligned structure and produces auditable, explainable rankings. The system combines LangGraph multi-node agents (job + resume), hybrid LLM/embedding similarity across work experience, education, and skills, and a stability penalty (tenure, gaps) to calibrate a final 1–5-star score. We also show how to design for compliance with NYC Local Law 144—including measurable impact ratios for bias audits, public transparency artifacts, and built-in notice/reporting hooks—so teams can operationalize fair, reviewable AI in hiring. Under the hood: Pydantic-guardrailed prompts, per-feature similarity matrices, MLflow-tracked latency/cost, and config-only model swaps (OpenAI/Gemini/Bedrock).
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top