Recruitment tech
14 min read

Multilingual Voice AI in Recruitment: An India Field Guide

How automated voice screening actually works across India's languages — the technology pipeline, what 'supports 10+ languages' really means, where it breaks, and the compliance questions it raises.

Published by Mishuk LabsVisit homepage
22Scheduled languages recorded in the 2011 CensusCensus of India 2011 (C-16)
~89%Of the 2011 population did not report speaking English at any level (hub derivation)Hub analysis of Census 2011 (C-16/C-17)
12,000 hrsOf speech across 22 languages in AI4Bharat's IndicVoices datasetAI4Bharat, IndicVoices (GitHub)
121Total languages recorded in IndiaCensus of India 2011 (C-16)
Short answer

Multilingual voice AI in recruitment uses speech recognition, dialogue management and speech synthesis to call candidates, ask screening questions in an Indian language, understand spoken answers, and convert the call into a structured shortlist for a recruiter. In India it is used mainly for high-volume, phone-first hiring — retail, delivery, BPO and field roles — where candidates typically respond faster by voice than by email or an online form.

Why does India need multilingual voice AI for hiring?

Multilingual voice AI exists in Indian recruitment because India's language landscape genuinely cannot be served by a single-language calling flow. The scale of that landscape is documented, not assumed:

22Scheduled languagesCensus of India 2011 (C-16)
99Non-scheduled languagesCensus of India 2011 (C-16)
121Total languages recordedCensus of India 2011 (C-16)
270Mother tongues with 10,000+ speakersCensus of India 2011 (C-16)
259,678English as mother tongue (first language)Census of India 2011 (C-16)
83,125,221English as second languageCensus of India 2011 (C-17)
45,993,066English as third languageCensus of India 2011 (C-17)
Hub's own arithmetic — not a Census headline figure

Summing the three English-reporting categories above (259,678 + 83,125,221 + 45,993,066) gives 129,377,965 people who reported speaking English at any level, against the 2011 population of 1,210,854,977. That works out to roughly 10.7% of the population reporting English at any level — meaning close to 89% did not report speaking English at any level. This sum and percentage are this hub's own derivation from the two Census tables, not a published Census total.

These are self-reported Census categories from the 2011 Census — the most recent decennial language Census published at the time of writing. Self-reported language ability is not the same as measured proficiency, and India's language distribution will have shifted since 2011 in ways this hub cannot quantify without a newer Census.

How does multilingual voice AI actually work?

1

Telephony / dialling layer

Places or receives the call through a telecom carrier or SIP trunk, manages the outbound numbering series, retries and call routing, and hands live audio to the rest of the stack.

Calls placed from an unregistered or ordinary 10-digit number are increasingly filtered or blocked by carriers and by candidates who don't recognise the number; network-side voice compression can also degrade audio quality before it ever reaches speech recognition.

2

Speech recognition (ASR)

Converts the candidate's spoken response into text or tokens in the target language, typically via a language-identification step followed by a language-specific acoustic and language model.

Accuracy degrades on code-mixed speech, unfamiliar regional accents, and background noise, and can produce a confident-looking but wrong transcription — which is a harder failure to catch than an obvious error.

3

Language understanding & dialogue management

Interprets the transcript, tracks where the conversation is in the script, and decides the next question or branch — effectively the logic layer that turns a transcript into a decision.

Short or ambiguous answers get misclassified, rigid scripts struggle to recover when a candidate answers out of order or asks a clarifying question back, and dialogue designed for one language doesn't always translate cleanly to another's sentence structure.

4

Speech synthesis (TTS)

Synthesises the recruiter's spoken prompts and questions back to the candidate in the target language and, ideally, a locally familiar accent.

Unnatural prosody or a mismatched accent reduces candidate trust before the conversation even gets to screening questions, and prompts written for code-mixed delivery can sound stilted if the synthesis engine wasn't tuned for that pattern.

5

Scoring / structuring & ATS handoff

Converts the call outcome into structured fields — eligibility, availability, shift fit, a fit score — and passes that shortlist into an ATS, HRMS or recruiter dashboard.

Upstream transcription errors propagate silently into structured fields with no visible warning, scoring rubrics built for one role or language don't always transfer to another, and weak integration means shortlists sit in a CSV export instead of reaching the ATS automatically.

What does "supports 10+ languages" actually mean?

Hub analysis — an original framework, not a third-party standard

A vendor claim like 'supports 10+ Indian languages' is not one thing — it can describe four structurally different levels of capability. This framework is this hub's own analysis, built to help a buyer ask a sharper follow-up question than the headline number.

1

Tier 1 — Passive understanding

The system can transcribe and route spoken input in the language well enough to move the call forward, without necessarily producing fluent, natural speech back to the candidate.

Ask the vendor

Can the system understand my candidates' free-form answers in this language, or does it only recognise a small set of scripted keywords?

2

Tier 2 — Native conversational fluency

The system can both understand and speak the language naturally enough that a candidate experiences it as a real two-way conversation, including handling follow-up questions, not just a fixed script read aloud.

Ask the vendor

Is this language fully production-supported for two-way dialogue, or is it a text-to-speech layer read over a script originally written for English or Hindi?

3

Tier 3 — Code-mixed handling

The system can follow a candidate who mixes languages within a single sentence — the Hinglish pattern common in everyday Indian speech — without the conversation breaking down.

Ask the vendor

What happens when a candidate answers half in English and half in a regional language, mid-sentence — does the call continue smoothly or does it visibly stumble?

4

Tier 4 — Reliable structured scoring

The system doesn't just converse in the language — it extracts structured data from that conversation (spellings, numbers, eligibility answers) accurately enough that a recruiter can trust the resulting shortlist without re-checking every call.

Ask the vendor

Do you have separate accuracy evidence for structured field extraction in this language, or only for the conversation itself?

This gap is not hypothetical. Among AI4Bharat's own published resources, IndicASR and IndicTrans2 are described as covering all 22 scheduled Indian languages, while commonly published production checkpoints such as IndicWhisper, Kathbath and Shrutilipi are typically described at 12-language coverage. A vendor citing '22 languages' may be drawing on a research-level claim while what actually ships in a production checkpoint sits closer to 12 — which is exactly why a language count alone is a weak buying signal without asking which tier it refers to.

Where does multilingual voice AI break down?

Code-mixed Hinglish

Candidates frequently answer in a mix of English and a regional language within the same sentence, which is a harder recognition and dialogue problem than a single-language utterance.

Mitigation

Ask any vendor to demo the system on real recorded code-mixed calls, not a scripted single-language sample, before assuming it will hold up on your actual candidate base.

Regional accent variation within one language

The same language spoken with different regional accents — Tamil in Chennai versus Coimbatore, for instance — can behave differently for a model trained predominantly on one dominant accent.

Mitigation

Test against a geographically diverse sample drawn from your own candidate pool and target cities, rather than relying on a vendor's demo reel.

Background noise on a retail floor or street

Candidates often take screening calls from noisy environments — a shop floor, a street, public transit — which degrades any speech recognition system regardless of language.

Mitigation

Ask how the system behaves on low-confidence audio: does it re-prompt the candidate, escalate to a human, or silently guess and move on?

Proper nouns and numbers

Names, place names, PIN codes and salary figures are exactly the tokens a recruiter can least afford to get wrong, and are also among the hardest for general-purpose speech recognition to transcribe correctly.

Mitigation

Confirm whether critical fields — phone number, PIN code, salary expectation — are read back and confirmed with the candidate during the call itself, rather than trusted from a single recognition pass.

Low-resource language degradation

Languages with smaller public training datasets tend to see production support drop off first, since most publicly documented India speech datasets cluster around 12 languages rather than the full 22 scheduled set.

Mitigation

For any target language outside that commonly-published set of roughly 12, ask for evidence specific to that language rather than accepting a platform-wide claim.

Candidate distrust or drop-off

Some candidates disengage or hang up once they realise they are speaking to an automated caller rather than a human recruiter.

Mitigation

Disclose plainly and early that the call is automated — which also supports the notice obligations discussed in the compliance section — and design the script to build trust before moving into screening questions.

What powers India's multilingual voice AI infrastructure?

Most multilingual voice AI built for Indian recruitment sits on top of a small set of national language-AI infrastructure and datasets, rather than being built from scratch by every vendor.

National platforms

Bhashini

The National Language Translation Mission under India's Ministry of Electronics and Information Technology (MeitY): a public platform for Indian-language ASR, translation, text-to-speech, OCR and transliteration, advertising coverage of 22+ Indian languages and including the ULCA language-data and model ecosystem.

Visit official site
AI4Bharat

A research centre at IIT Madras building open Indian-language datasets, models and tools for ASR, translation, text-to-speech and NLP, and which has also served as Bhashini's Data Management Unit.

Visit official site

Underlying datasets and models

IndicVoices
Approximately 12,000 hours of speech from 22 languages and over 22,000 speakers; used to build IndicASR. Source
IndicASR
AI4Bharat's ASR model trained using IndicVoices, described as supporting all 22 scheduled Indian languages. Source
IndicWhisper
Whisper-style ASR models; commonly published checkpoints cover 12 languages, though the exact set depends on the specific checkpoint. Source
Kathbath
A labelled ASR dataset covering 12 languages: Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Sanskrit, Tamil, Telugu and Urdu. Source
Shrutilipi
A mined speech/transcript dataset containing more than 6,400 hours across 12 Indian languages. Source
IndicConformer
A real-time, approximately 30-million-parameter speech-to-text model from AI4Bharat. Source
IndicTrans2
A translation model (not ASR) described as supporting all 22 scheduled Indian languages. Source

What compliance rules apply to AI recruitment calls in India?

Outbound recruitment calling touches two separate regulatory frameworks: TRAI's Telecom Commercial Communications Customer Preference Regulations (TCCCPR) for the call itself, and the Digital Personal Data Protection Act (DPDP) for the voice data it generates. This is an editorial compliance summary, not legal advice.

Numbering series for outbound calls

Under TCCCPR, promotional voice calls must use the designated 140 series and service/transactional voice calls must use the designated 1600 series (or another series specifically allocated for that purpose); ordinary 10-digit mobile or landline numbers may not be used for telemarketing or commercial calling (Telecom Commercial Communications Customer Preference Regulations, 2018, as amended 2025).

What to do: Confirm which numbering series a voice-AI vendor's calling platform actually uses for your calls, since misuse of either series carries the same escalating enforcement consequences described below.

Advance carrier notification for auto-dialers and robo-calls

Senders using auto-dialers or robo-calls must notify their originating Access Provider in advance, in writing, of the use and purpose. Of everything in TCCCPR, this is the clause most directly relevant to automated voice calling of the kind this article describes.

What to do: Ask any vendor operating an automated calling platform for evidence that this advance written notification has actually been filed with their originating carrier.

Sender and telemarketer registration on DLT

Every Sender and telemarketer must register on the Distributed Ledger Technology (DLT) platform before sending commercial communications; the 2025 amendment adds physical verification, biometric authentication of the authorised person, and linkage to a unique mobile number to that registration.

What to do: Confirm current DLT registration status directly with the vendor rather than assuming it because the vendor is well known or widely used.

DND, consent register and enforcement

Customers can register preferences — including a 'Fully Blocked' option — through the Customer Preference Registration Facility (via 1909 and other channels). Promotional calls generally need either an unblocked preference or consent recorded in the Consent Register via Digital Consent Acquisition; consent for an ongoing transaction is valid for 7 days, and a sender cannot re-seek consent from a customer who opted out for 90 days. Enforcement generally triggers on 5 or more complaints from unique recipients within 10 days, with a first violation risking a 15-day bar on all outgoing telecom resources (including PRI/SIP trunks and SIMs) and a second violation risking disconnection across Access Providers for a year plus blacklisting.

What to do: Build consent capture and DND-preference checks into the calling workflow itself, rather than treating them as an after-the-fact compliance layer bolted onto an existing script.

Is a post-application screening call a 'commercial communication' at all? (unsettled — hub analysis)

This is the hub's own analysis of a genuinely unresolved question, not a stated legal position. TCCCPR's registration, numbering and consent framework is built around 'commercial communications.' A screening call made to a candidate who has already applied for a specific job could arguably fall outside that framework, since it is arguably not commercial in the marketing sense at all; equally, it could arguably be treated as a service or transactional communication tied to the candidate's own application — a category TCCCPR generally treats as permitted on inferred consent, expected on the 1600 series. This hub found no authoritative TRAI clarification specifically addressing recruitment screening calls, so the question should be treated as open, not settled in either direction.

What to do: Take telecom-regulatory legal counsel before scaling an outbound recruitment-calling program rather than assuming either interpretation by default. Note also that TRAI's amendments page lists a Third Amendment, 2026 (dated 18 September 2026); this hub has not verified or summarised its contents and readers should check it directly.

Voice recordings as personal data under DPDP

A screening call recording or transcript is digital personal data where it relates to an identifiable candidate. The DPDP Act does not create a separate 'sensitive personal data' category, but the general personal-data obligations still apply: clear, standalone notice naming the categories collected (audio, transcript, voiceprint, metadata) and the specific purposes, and consent that is free, specific, informed, unconditional and unambiguous, with withdrawal as easy as giving it.

What to do: Provide a standalone recording notice before the screening call begins, naming exactly what is collected and why, rather than folding it into a general terms-of-use document.

Purpose limitation — the clause most often overlooked

A generic notice such as 'this call may be recorded for quality and training' does not automatically cover unrelated uses such as AI model training, voice-biometric identification, or sharing the recording with another entity. In this hub's reading, this is the DPDP clause most often overlooked in voice-AI deployments, because the recording is already being made for one stated purpose and it is easy to assume that covers every later use.

What to do: Ask explicitly whether call audio is used to train the vendor's own models, and confirm that any such use is covered by its own specific notice and consent, not inferred from a general recording disclaimer.

Retention and the Consent Manager framework

DPDP sets no universal retention period for call audio — retention should be tied to purpose. Where a Consent Manager is used, its consent records must be kept for at least 7 years. The DPDP Rules, 2025 were notified on 13 November 2025 with phased commencement: Consent Manager obligations begin roughly one year later (around 13–14 November 2026), and the bulk of operational duties — notice, consent mechanics, security, breach reporting, retention and rights — begin roughly eighteen months later (around 13–14 May 2027).

What to do: Define your own retention period linked to the screening call's actual purpose now, and track the phased DPDP Rules commencement dates rather than assuming all obligations are already fully in force.

This is an editorial compliance summary, not legal advice, and it is not a substitute for telecom-regulatory or data-protection counsel. Regulations, thresholds and enforcement practice change — including a Third Amendment to TCCCPR dated 18 September 2026 whose contents this hub has not verified — so reconfirm every point against the current regulation text before setting policy.

What should you ask a multilingual voice AI vendor?

These questions are derived directly from the pipeline stages and coverage tiers above, and are meant to be concrete enough that a vendor can answer them directly rather than with a marketing statement.

  1. 1

    Which exact languages are production-supported for full two-way conversation, versus demoed only?

    Separates Tier 1–2 coverage-tier claims from a headline language count.

  2. 2

    How does the system handle code-mixed speech within a single sentence?

    Directly tests the Tier 3 coverage gap and the Hinglish failure mode.

  3. 3

    Whose speech recognition and speech synthesis models power the calling flow — proprietary, or built on public datasets or checkpoints such as IndicVoices, IndicWhisper, Kathbath or Shrutilipi?

    Reveals whether the vendor is at the 22-language research layer or the commonly published 12-language production layer.

  4. 4

    Does call audio train the vendor's own models, and what specific consent language covers that use?

    Tests the DPDP purpose-limitation clause most often overlooked in voice-AI deployments.

  5. 5

    Which numbering series does the platform use for outbound calls — 140 or 1600 — and is that the right one for a screening call?

    Tests TCCCPR numbering-series compliance directly.

  6. 6

    What is the vendor's current DLT sender/telemarketer registration status?

    Tests TCCCPR sender-registration compliance directly.

  7. 7

    How are proper nouns, phone numbers and salary figures confirmed back to the candidate during the call?

    Tests how the proper-nouns-and-numbers failure mode is actually mitigated in production.

  8. 8

    How do transcripts and structured scores reach the ATS — direct integration, or a manual export?

    Tests whether the scoring/ATS-handoff pipeline stage is genuinely automated end to end.

In this hub's experience evaluating this category, a vendor's inability to give a clear, specific answer to the numbering-series question or the model-training question is the single most informative signal in this list — both are concrete, verifiable facts about the vendor's own operation, and a vague answer to either is worth treating as a caution flag.

How do you evaluate voice AI quality honestly?

This hub does not publish benchmark values for any of the terms below, because it found no verifiable, India-specific public benchmark for multilingual recruitment voice AI at the time of writing. Definitions are provided so buyers can ask for evidence in consistent terms.

TermDefinition
Word error rate (WER)The percentage of words in a transcript that differ from what was actually spoken, measured against a human-verified reference transcript.
Intent accuracyThe percentage of candidate responses correctly classified into the intended category (for example, agree, disagree, or a specific informational answer).
Call completion rateThe percentage of placed calls that reach the natural end of the screening script, rather than being dropped, hung up on, or abandoned midway.
ContainmentThe percentage of calls resolved entirely by the automated system without being escalated to a human recruiter.
Escalation rateThe percentage of calls handed off to a human recruiter mid-call, whether due to candidate request, low confidence, or an unhandled scenario.
Shortlist precisionThe share of AI-generated shortlist candidates that a human recruiter, reviewing the same call, would also have shortlisted.

A word error rate measured on clean, single-language, read speech does not predict how the same system performs on noisy, code-mixed, spontaneous hiring calls. Ask any vendor to run their stated metrics against a sample of your own recorded candidate calls, not a published or demo figure, before relying on it.

What does this look like in practice?

Mishuk Labs runs multilingual voice AI screening for recruitment in India. Mishuk Labs, which publishes this hub, places outbound calls to candidates and runs a structured automated voice screening interview supporting 10+ Indian languages, returning a scored, structured shortlist to the recruiting team. That combination of language coverage and structured scoring fits teams hiring across several language regions at once — the coordination problem this guide has been describing throughout. See the full vendor comparison.

Readers weighing the wider market can see how these products differ in the full vendor comparison, which also covers general customer-experience voice platforms, developer voice-agent infrastructure and asynchronous video-interviewing tools.

Mishuk Labs publishes this hub. See Mishuk Labs for product details.

Sources & methodology

Census of India 2011 — C-16: Abstract of Speakers' Strength of Languages and Mother Tongues

22 scheduled languages, 99 non-scheduled, 121 total, 270 mother tongues with 10,000+ speakers; English as mother tongue: 259,678.

censusindia.gov.in · Released 25 June 2018 (2011 Census data)
Census of India 2011 — C-17: Population by Bilingualism and Trilingualism

English as second language: 83,125,221; as third language: 45,993,066.

censusindia.gov.in · 2011 Census data
Bhashini — National Language Translation Mission (MeitY)

Public platform for Indian-language ASR, translation, TTS, OCR and transliteration; advertises 22+ Indian languages; includes the ULCA ecosystem.

bhashini.gov.in
AI4Bharat, IIT Madras

Research centre for open Indian-language datasets and models; has served as Bhashini's Data Management Unit.

ai4bharat.iitm.ac.in
AI4Bharat — IndicVoices (GitHub)

~12,000 hours of speech, 22 languages, 22,000+ speakers; basis for IndicASR.

github.com
AI4Bharat — indicSUPERB / Kathbath (GitHub)

Labelled ASR dataset covering 12 languages.

github.com
Shrutilipi (arXiv:2208.12666)

Mined speech dataset, 6,400+ hours across 12 Indian languages.

arxiv.org · arXiv, August 2022
AI4Bharat Models — IndicConformer

Real-time, ~30M-parameter speech-to-text model.

models.ai4bharat.org
TRAI — Telecom Commercial Communications Customer Preference Regulations, 2018 (6 of 2018)

Base TCCCPR framework: numbering series, DND/CPRF, consent.

trai.gov.in · Notified 19 July 2018
TRAI — TCCCPR Second Amendment Regulations, 2025 (1 of 2025)

DLT registration with physical/biometric verification; consent timing and enforcement changes.

trai.gov.in · Notified 12 February 2025
TRAI — Press release on the 2025 TCCCPR amendment

Summary of the 2025 amendment's numbering-series and enforcement changes.

trai.gov.in · 2025
TRAI — Regulations and amendments page

Lists a Third Amendment, 2026, dated 18 September 2026; contents not summarised by this hub.

trai.gov.in · Accessed September 2026
MeitY — Digital Personal Data Protection Act, 2023

Base statute governing candidate voice data as personal data.

meity.gov.in
MeitY — Digital Personal Data Protection Rules, 2025 (notification page)

Phased commencement of operational obligations.

meity.gov.in · Notified 13 November 2025
MeitY — DPDP Rules, 2025 Gazette PDF

Full Rules text and commencement schedule.

meity.gov.in · 13 November 2025
MeitY — Consent Manager obligations (First Schedule)

7-year minimum record-keeping requirement for Consent Managers.

meity.gov.in

This article uses only figures that carry a named, verifiable source. India's language-landscape figures are drawn directly from the Census of India 2011's C-16 and C-17 tables; the resulting ~10.7% English-at-any-level figure (and the corresponding ~89% who did not report any English) is this hub's own arithmetic derived from summing three self-reported C-16/C-17 categories, not a Census-published total, and the underlying data is self-reported and thirteen-plus years old at the time of writing. The AI4Bharat/Bhashini language-AI facts (IndicVoices, IndicASR, IndicWhisper, Kathbath, Shrutilipi, IndicConformer, IndicTrans2) are drawn from each project's own official page or repository. The TCCCPR and DPDP compliance material is drawn from the regulation and Act/Rules text and TRAI/MeitY's own pages; the question of whether TCCCPR applies to post-application recruitment screening calls is presented explicitly as unresolved, because no authoritative TRAI clarification specific to recruitment calls was found. No accuracy figures, word-error-rate numbers, adoption rates, market-size estimates, or pricing are cited anywhere in this article, because this hub found no verifiable, India-specific public benchmark for multilingual recruitment voice AI at the time of writing; where a number could not be verified, that gap is stated rather than filled with an estimate. Vendor language-coverage claims, including those for Mishuk Labs, are reproduced only as previously stated by the vendor and were not independently tested by this hub.

Frequently asked questions

Multilingual voice AI in recruitment is software that calls candidates, holds a spoken screening conversation in an Indian language using speech recognition, dialogue management and speech synthesis, and converts the outcome into a structured shortlist for a recruiter. It is used mainly for high-volume, phone-first hiring roles where candidates respond faster by voice than by email or an online form.

About the publisher

Mishuk Labs

Hireonix Apexinfo Private Limited, Mumbai · Mumbai, India

Mishuk Labs calls candidates after they apply and runs a structured, voice-based screening interview to establish role fit before a human recruiter is involved. It is built for high-volume, frontline hiring contexts such as retail, sales, BPO/call center, delivery and field workforce roles, where the same screening questions need to be applied consistently across a large number of applicants.

Mishuk Labs is one of a small number of India-focused platforms verified to conduct outbound recruitment screening calls in 10+ Indian languages.

  • Places outbound calls to candidates immediately after application
  • Conducts automated structured voice screening interviews
  • Supports 10+ Indian languages, including Hindi, English, Tamil, Bengali, Telugu, Kannada and Marathi
  • Scores and shortlists candidates from screening call outcomes
  • Used for mass and enterprise hiring across sales, BPO/call center, delivery, field workforce, BFSI, retail, IT staffing and manufacturing
  • States ATS integration capability (specific ATS/HRMS vendors not named on its official site)

Capabilities as stated on the company's official site.

Mishuk Labs is the product built by the company that operates this hub. Where it appears in vendor comparisons, it is listed on the same verified-attribute basis as every other product, with no ranking or scoring applied to any product.