Why AI Answer Tools Have Broken Async Video Interviews
Candidates can now use AI tools that write interview answers in real time, which changes what a one-way video recording measures. Why the weakness is the async format itself, and why live, adaptive screening holds up better.

What is a recruiter actually watching when she reviews a one-way video interview? The idea was that she's watching the candidate: how they think and how they talk when asked something. For a while that was roughly true. I don't think it is anymore.
Async video sold recruiters an efficiency trade. Instead of scheduling phone screens, you give every candidate the same recorded prompts, let them answer whenever it suits them, and review the recordings in batches. A recruiter could get through several times as many candidates as a phone process reached in the same hours, and nobody had to find a slot that worked for both sides. On paper you got the same information as a live screen with far more throughput. In 2022 that was a good deal.
It rested on an assumption so obvious that nobody bothered to state it: the answers in the recording came from the person in the recording.
Tools that write interview answers in real time are now easy to get. Some are browser extensions that put a suggested answer on screen while the question plays. Some listen through the microphone and write a response in a sidebar the camera can't see. Others polish whatever the candidate drafts before the recording uploads.
So some share of async video answers, probably a growing one, don't show how the candidate thinks or communicates. They show how the candidate's AI tool handles interview questions. And since most of these tools are built on the same few underlying models, the answers are starting to converge. A recruiter reviewing a batch sees one polished, coherent response after another and struggles to tell them apart, which makes sense, because underneath they are close to being the same answer.
We've watched this happen once already with written applications, as covered in the post on AI-generated applications flooding your ATS. The resume screen went first, and the video screen is going the same way.
What the two formats have in common is that neither reacts to what the candidate says. The question in an async video is known ahead of time, or can be worked out within seconds of it starting to play. Most platforms allow retakes. Nothing follows up on the candidate's actual answer, asks for a concrete example when they give an abstract one, or throws in a question that breaks a prepared script.
Under those conditions you aren't really screening anyone. The candidate, or their tool, has as long as the platform allows and knows exactly what's being asked. What you end up measuring is how well someone optimizes for the format, which tells you little about how they think when they have to answer on the spot.
A live conversation is much harder to game. If someone has to paste each question into a tool and wait for the answer, the pause is noticeable when the other side is waiting for them to speak. And when the next question depends on what you just said, there's no script to prepare. Cheating doesn't become impossible, but there's far less room for it.
That reasoning is behind how we built Asendia, and it's the main reason Asendia is voice AI and not another recorded format. It phones each candidate after they apply, whether that's the same evening, late at night, or over the weekend, and runs a structured screening conversation set up for the specific role. Talking in real time leaves no quiet gap for pasting a question into a tool, and because the conversation responds to what the candidate says, there's no fixed script to prepare against. A vague answer gets a follow-up asking for specifics. A claimed skill gets a question about where and how they used it, instead of being accepted and passed over. What goes into the ATS is a qualification summary with excerpts from the conversation, taken from a live exchange and not from a recording the candidate could redo as many times as they liked.
This matters most when volume is high. Suppose an agency runs a campaign and 400 applications have arrived before the recruiter opens her queue on Monday. Asendia can have a real conversation with each of those applicants within hours of their applying, without the agency adding people, and the recruiter starts the week with a shortlist built from conversations that were hard to prepare for. The pipeline in the ATS looks the same as before. The difference is that everyone in it was spoken to in real time.
I don't blame candidates for using these tools. Employers already use AI to screen and rank them, so candidates are doing the same thing from the other side. The trouble is that async video was never designed to hold up when they do. A screen doesn't have to be run by a human to be reliable. It has to happen live and respond to what the person says, so that there's nothing useful to prepare in advance, and async video as it works today does neither. If your screening process hasn't changed since 2022, try watching a batch of recent recordings with this in mind and count how many you could tell apart.
Ready to transform your hiring strategy? Schedule a Demo with our founders today!
Badis Zormati
Co-Founder, Asendia AI

