Back

Would Two Recruiters Screen the Same Candidate the Same Way?

Before asking whether AI screening is reliable, it's worth asking how consistent unstructured human phone screens are. What inter-rater reliability research suggests, and why the first-round screen deserves more scrutiny.

AI Implementation5 min read
Would Two Recruiters Screen the Same Candidate the Same Way?

Here's an experiment almost no recruiting team has run. Take one candidate and have two of your recruiters phone-screen them separately, without comparing notes. Would they reach the same verdict?

Most people would guess yes, give or take. I suspect the real answer is worse than that, and it matters for a debate that has been going on for two years about whether AI screening is reliable. That debate keeps asking whether AI can assess candidates as well as a human. It rarely asks how consistent the human process was to begin with.

Consistent screening just means that two recruiters evaluating the same candidate against the same criteria reach roughly the same conclusion. It sounds like a low bar. But think about how the typical phone screen works. There's no fixed set of questions and no scoring rubric, and the format changes from one recruiter to the next. It would be surprising if that produced consistent results.

The research suggests it doesn't. One often-cited estimate puts the inter-rater reliability of unstructured interviews, meaning the correlation between two independent evaluators judging the same candidate, at around 0.37. A score of 1.0 would be perfect agreement. At 0.37, the two assessments share less than 14 percent of their variance. In any useful sense, the two recruiters aren't measuring the same thing. Each is measuring the candidate as seen through their own questions and biases, and through however that particular Tuesday afternoon is going.

Structured interviews help. When every candidate gets the same questions in the same order and is rated against a defined rubric, reliability rises to somewhere around 0.56 to 0.67. That's a real improvement. The catch is that structure takes discipline to keep up, and my guess is that many teams that adopt it drift back toward conversational screening within months. The rubric turns into a formality, and the interviews stop being structured in any way that counts.

Why nobody notices

If the problem is this large, why don't teams see it? Because nothing in the process would ever show it to them. A recruiter screens a candidate and records a pass or a fail. The candidate moves forward or doesn't. Nobody has the same person screened a second time to compare, so there's no feedback loop at all.

What does show up is indirect. Hiring managers get "qualified" candidates from different recruiters and can't work out why the bar seems to move. One recruiter's strong yes is another's maybe. Teams put this down to judgment being subjective. I think it's better described as an instrument with poor reliability: use the same tool twice and you get different readings, even though the candidate hasn't changed.

The cost builds up quietly. Some candidates who should have gone through didn't, because the recruiter who screened them was having a hard afternoon, didn't click with how they talked, or used their usual questions instead of the ones this role needed. Others who shouldn't have gone through did, for the opposite reasons. Quality of hire suffers even though the hiring team's judgment is fine, because the first filter was noisy. And since nobody measured the noise, nobody fixed it.

This is why I think the case for AI screening calls is usually argued backwards. The claim worth making isn't that AI has better judgment than an experienced recruiter. It's that AI asks the same questions the same way every time, for every candidate, whether the call happens late at night or on a Friday afternoon. That goes after the consistency problem where it starts.

Asendia is built around that idea. What reaches the recruiter is a written qualification summary for each candidate, and every candidate for a role was asked the same structured set of questions, which the recruiting team defines at the start of the campaign. The candidate's answers are attached word for word, so two summaries can be read side by side and actually compared, and the recruiter applies their own judgment on top. Behind each summary is a live spoken screening conversation that took place soon after the person applied.

The recruiter still makes the call. They just make it from a consistent record instead of reconstructing one from notes scribbled during a 30-minute conversation. Screening quality stops depending on who happened to pick up the phone that morning and starts depending on the criteria you chose.

For agencies this matters more, because inconsistency between recruiters becomes something clients see. Imagine a team of four screening for six clients at once. A bad screen on an important campaign reflects on the agency as a whole. If every first conversation runs against the same criteria, the recruiters inherit a pool that was screened one way rather than four. And once first-pass consistency is achievable, it changes which numbers are worth tracking, which I wrote about in the post on recruitment KPIs in a post-AI world.

Hiring teams put a lot of work into interview training, rubrics and debrief formats, all at stages that already run with some consistency. The first-round screen, which decides who gets to those stages, rarely gets the same scrutiny. Fixing it doesn't require giving up human judgment later in the process, only admitting that the unstructured phone screen gives readings that depend more on the recruiter than on the candidate. You could try training recruiters harder, but I think handing them a consistent starting point does more.

To find out how big the problem is on your own team, run the two-recruiter experiment on a handful of candidates. It won't take long, and I suspect the results will settle the question.

Ready to transform your hiring strategy? Schedule a Demo with our founders today!

Badis Zormati

Co-Founder, Asendia AI

Ready to transform your hiring strategy?

Schedule a Demo

Keep reading