6 mins read
/
What AI Voice Screening Actually Gets Wrong

Abhimanyu Roat
Co-founder & CEO

The accuracy problem nobody quotes upfront
Speech-to-text is the foundation of any AI voice screener. Everything downstream — the scoring, the transcript, the shortlist — depends on how accurately the system heard the candidate.
In India, that foundation is shakier than vendors admit.
Indian candidates do not speak in clean, broadcast-quality English. They speak in heavy regional accents. They code-switch between English and Hindi mid-sentence. They take calls from noisy environments — construction sites, shared rooms, bus stands. Standard speech-to-text models were not built for this.
Word error rates for domain-specific Indian speech range from 7.69% for well-trained models to 24.5% for standard systems that were not tuned on regional accents or Hinglish. At the high end, that is one word in four transcribed incorrectly.
When the transcript is wrong, the score is wrong. A candidate from Bihar who pronounces "forklift" in a way the model does not recognise gets flagged as unqualified. The system produces what looks like a clean rejection — but it is a false negative built on bad audio processing. The agency misses a good candidate and never knows why.
The fix is not to abandon AI screening. It is to require human review on any candidate who scores close to the threshold. The borderline cases are exactly where transcription errors cluster, and they are also exactly where a recruiter's judgment matters most.
Rigid scripts and the candidates who just hang up
A human recruiter can improvise. If a candidate asks a question the recruiter was not expecting, they answer it and move on. If a candidate sounds frustrated, they adjust their tone.
A voice bot follows a script. When the conversation goes off-branch — when a candidate asks something the bot was not trained to handle, or gives an answer that does not fit the expected options — most systems loop awkwardly or respond with something generic.
For blue collar candidates doing a basic availability check, this rarely matters. The questions are simple and the branching is predictable.
For mid-market white collar candidates, it is a different problem. A sales manager with four years of experience who gets looped twice by a bot asking if they have "any relevant experience" will just end the call. They have other options. The agency loses the candidate without knowing the conversation fell apart.
The solution is scope management. AI voice screening works cleanly for binary qualification: location, notice period, salary range, availability. It does not work well for anything that requires improvisation. Keep the bot inside those boundaries and the drop-off rate stays low.
The compliance gap most agencies have not noticed
The Digital Personal Data Protection Act, 2023 — the DPDP Act — changed the legal environment around candidate data in India. Most agencies know it exists. Most have not worked through what it means for AI voice screening specifically.
The short version: an AI voice screener is a data collection system. It records voices. It generates transcripts. It stores information about individuals. Under the DPDP Act, doing any of this without explicit prior consent from the candidate is a violation.
Candidates are not employees. That matters because the DPDP Act has a carve-out for processing employee data in the context of an employment relationship — but candidates are not yet employees, so the carve-out does not apply. Explicit, granular, purpose-specific consent is required before the call begins.
There is also a data retention requirement that agencies routinely ignore. If a candidate is screened and not placed, the agency cannot keep their voice recording and transcript indefinitely. The law requires deletion within a defined window — typically 12 to 24 months — unless the candidate has given fresh consent to stay in the database.
Penalties under the DPDP Act are not symbolic. They reach up to ₹250 crore for failure to implement reasonable security safeguards. For a recruitment agency that has automated thousands of screening calls without a compliant consent flow or a deletion schedule, that exposure is real.
Building compliance into the screening workflow is not complicated. Consent can be captured at the start of the call. Automated deletion schedules are standard in most data infrastructure. The agencies that are exposed are not the ones who investigated and decided it was too hard — they are the ones who have not looked at it yet.
What this means in practice
AI voice screening works. The agencies using it well are producing shortlists faster, handling more mandates with the same headcount, and running a screening stage that costs a fraction of what manual calling costs.
But it is not a plug-and-play tool. It needs calibration for the Indian audio environment. It needs scoping to the questions where automation actually performs reliably. And it needs a compliance layer before the first call goes out.
The agencies that will struggle are the ones that deploy it as a black box, trust every score at face value, and assume the legal question is someone else's problem. The ones that will win are the ones that treat it as a fast, flawed assistant — and build the human checkpoints to catch what it misses.
Frequently Asked Questions
How accurate is speech-to-text for Indian accents and Hinglish?
It depends heavily on the model and how it was trained. Word error rates for domain-specific Indian speech range from 7.69% at the optimistic end to 24.5% for standard models. That means one in four words transcribed incorrectly at the high end — enough to affect how a candidate is scored.
Does the DPDP Act apply to AI voice screening of candidates?
Yes. Candidates are not employees, so they do not fall under the legitimate employment use exemption. Agencies must capture explicit consent before recording or transcribing a candidate's voice. Transcripts of rejected candidates must be deleted within 12 to 24 months. Penalties reach up to ₹250 crore.
Should agencies stop using AI voice screening because of these risks?
No. The answer is guardrails, not avoidance. Human review of borderline transcripts catches false negatives from accent errors. DPDP-compliant consent flows are straightforward to build. Treat AI screening as a fast data collection tool, not a replacement for final recruiter judgment.
Share this post


