The shift to video interviewing happened fast. Pandemic necessity accelerated adoption, and efficiency gains cemented it. Async video interviews in particular have become a genuinely useful tool for recruiters. They eliminate scheduling friction, give candidates flexibility, and let hiring teams review responses on their own time.
The problem isn’t video interviewing. The problem is what happened next.
As video interviewing became standard, AI-powered scoring followed. The pitch is compelling: instead of watching hours of recordings and comparing subjective impressions, let the AI score every response and surface your best candidates automatically. In theory, that’s exactly what structured, data-driven hiring looks like. In practice, many implementations fall short of that promise. And that gap has real consequences for candidates, for recruiters, and for the organizations relying on these tools.
What video interviewing is actually measuring
The first question any HR leader should ask about an AI scoring system isn’t “how accurate is it?” It’s “what is it actually measuring?”
For a lot of video interview scoring tools, the honest answer is uncomfortable. Many platforms evaluate tone of voice, speech pace, eye contact, facial expressions, and filler words alongside, or even instead of, the actual content of what candidates say. The assumption is that these mannerisms correlate with competence or cultural fit. The evidence for that assumption is thin. The evidence that measuring these mannerisms introduces bias is much stronger.
Several major players in the AI hiring space quietly removed facial analysis features after regulatory pressure and public scrutiny. That’s an acknowledgment that these criteria were never as reliable or fair as advertised. But the problem didn’t disappear when facial analysis did. It just became less visible. Any scoring system that evaluates how a candidate presents rather than what they communicate is still measuring the wrong thing.
This matters for two reasons. Candidates deserve a fair evaluation. And organizations that believe they’re making data-driven decisions deserve to actually be making data-driven decisions.
The accountability gap
Here’s a scenario worth thinking through: a hiring decision gets challenged. It could be a candidate questioning why they weren’t selected, an internal audit, or a regulatory inquiry. The question is how the decision was made.
“The AI scored them lower” is not a satisfying answer to any of those. It can’t be traced back to specific job-relevant criteria. It can’t be explained to the candidate. It can’t withstand scrutiny from a compliance standpoint. When you can’t say, “this candidate scored lower because their response to the structured thinking question demonstrated X rather than Y,” you don’t have a data-driven hiring process. You have a process that feels data-driven because a number came out at the end.
The irony is that organizations adopting black-box tools often do so in an attempt to reduce bias and improve consistency. Those are the right goals. A tool that can’t be interrogated, explained, or audited doesn’t achieve them. It just makes the problem harder to see.
What structured scoring actually looks like
There’s a meaningful difference between AI that scores video interviews and AI that scores them the right way. The distinction comes down to structure, transparency, and what the system is actually evaluating.
A structured approach starts with job requirements. What competencies does this role demand? What does good performance actually look like? From those answers, you build rubrics: explicit criteria that describe what a strong, adequate, or weak response looks like for each competency being assessed. Those rubrics become the scoring standard before any candidate records a response.
When responses come in, the system evaluates what candidates actually said against those pre-defined criteria. Not how they sounded or looked. What they communicated. Criterion-level scores roll up to an overall assessment, giving hiring teams a ranked view they can trust because they can see exactly what produced it.
Candidate response summaries serve a similar function. Rather than requiring panel members to watch every recording, a well-generated summary captures the substance of what a candidate communicated, giving reviewers a factual record they can evaluate consistently against the same criteria. This is where the real efficiency gains live, and they don’t come at the expense of fairness or accountability.
None of this works unless recruiters can review, edit, and approve the criteria before scoring begins. The AI should generate a starting point, identifying relevant competencies and drafting rubric criteria, but humans need to own the standard. If a recruiter can’t look at a scoring rubric and say “yes, this is what we’re evaluating and why,” it shouldn’t be deployed.
The human in the loop
I want to be clear about something that gets lost in a lot of conversations about AI in hiring: the goal is not to automate hiring decisions. It’s to free up human judgment for when it actually matters.
When recruiters spend most of their time on scheduling, coordination, and watching recordings, they have little left for the work that genuinely requires a human: building relationships with candidates, understanding team dynamics, making contextual calls that no algorithm can. Structured AI scoring, done right, shifts that balance. It handles the repetitive evaluation work so recruiters can focus on candidates who have already been thoughtfully assessed.
But this only works if the AI is transparent enough that humans can trust it, interrogate it, and override it when needed. AI should give HR professionals more confidence in their decisions, not less ability to explain them.
What to ask before you adopt AI video interviewing
For HR leaders evaluating these tools, four questions cut through a lot of noise quickly.
What specifically is being scored? Ask for an explicit list of the criteria the system evaluates. If the answer includes anything other than the content of candidate responses, ask for the validation data behind those criteria.
Is the scoring criteria tied to job requirements? Generic rubrics applied across all roles aren’t structured. They’re just standardized. Legitimate structured scoring starts with the specific competencies required for the specific job.
Can I see, edit, and approve the criteria before scoring begins? If the rubrics are fixed and opaque, you’re not in control of your evaluation standard. You should be able to review and modify what the system will use before a single candidate is evaluated.
Can I explain any score to a candidate or a regulator? This is the accountability test. If the answer requires “the AI said so” rather than pointing to specific, documented criteria and how a candidate performed against them, the process isn’t defensible.
These aren’t difficult questions. Good tools will answer them clearly. The ones that don’t are telling you something important.
What we’re seeing in the market
HR leaders are asking harder questions about AI now. They want transparency, auditability, and clear connections between what the AI measures and what actually predicts success. That shift is good for candidates, good for organizations, and good for AI in HR overall.
At Cangrade, we’ve seen that transparency isn’t just an ethical requirement. It’s what makes the efficiency gains real and sustainable. When recruiters can trust a score because they can see exactly how it was produced, they use it. When they can’t explain it, they don’t. Tools that claim to save time but can’t be trusted don’t actually save time. The tools that will earn a long-term place in this space are the ones that can stand behind their decisions.
Explore Hrtech Articles for the latest Tech Trends in Human Resources Technology

