Data Engineer(Speech/Language Data)
Check Your Resume Match Score
Scan your resume against ATS criteria for this Data Engineer(Speech/Language Data) role at Gnani Innovations.
Apply for this position
Job Description
Remote
3 - 5 years experience
Responsibilities
Data Engineer / Analyst — Speech & Language Data
Location: Bengaluru Type: Full-time Function: AI Team
About the role
Every speech model we ship is downstream of a decision someone made about data. Which hours got recorded. Which transcripts were trusted. Whether the Tamil test set accidentally shares speakers with the Tamil training set. Whether a number was written as २५ or 25. Those decisions set the ceiling on what our ASR and TTS models can ever achieve — and no amount of GPU time recovers a corpus that was built carelessly.
Training accurate ASR across 22+ Indian languages takes tens of thousands of hours of audio. Training TTS voices that sound human takes carefully recorded, tightly aligned, phonetically balanced speech. Right now that data lives across many places and many heads. We want one person to own it.
This role is the data owner for Gnani’s AI team. You source it, clean it, label it, version it, catalogue it, and defend its provenance — and you are the single point of contact when any researcher needs hours, a benchmark, or an answer about where something came from. It is a hands-on engineering and analysis role with real judgment attached: you will listen to audio, read transcripts, and be the person who notices that something is wrong before a model trains on it.
Core mandate
Sourcing & acquisition
Find, evaluate, and bring in speech and text data across 22+ Indian languages — public, licensed, customer, and commissioned
Pipelines & data quality
Build the cleaning, segmentation, alignment, and validation pipelines that turn raw audio into trainable hours
Annotation & benchmarks
Own transcription and labelling quality, and maintain the frozen evaluation sets the whole team is judged on
Data ownership & governance
Be the single point of contact for data — catalogued, versioned, consented, and compliant by default
What you’ll drive
Sourcing & acquisition
Be permanently on the hunt for data — track new public Indic speech and text corpora as they release, evaluate them for quality and licence terms, and bring the good ones in fast
Know the landscape cold: the AI4Bharat ecosystem (IndicVoices, Kathbath, Shrutilipi, GramVaani, IndicTTS), LIMMITS, SYSPIN, Rasa, Common Voice, FLEURS, and whatever lands next — and maintain an honest view of what each is actually good for
Assess and onboard customer-contributed data — what consent covers it, what the DPA permits, what has to be de-identified before it can be used, and what cannot be used at all
Spec and manage commissioned data collection and studio recording — script design for phonetic and dialect coverage, speaker recruitment criteria, session QC, and vendor throughput
Keep a running gap analysis: hours per language, per dialect, per channel, per domain — so the team argues about priorities with numbers instead of impressions
Pipelines, cleaning & curation
Build and own the pipelines that turn raw audio into trainable hours — format conversion and resampling, loudness normalisation, VAD-based segmentation, diarisation, and forced alignment
Automate quality screening at scale: SNR estimation, clipping and DC offset, truncated utterances, silence-heavy segments, language-ID mismatches, and audio–transcript drift
De-duplicate seriously — audio fingerprinting and near-duplicate text detection, so the same recording is not both trained on and evaluated on
Own the text side, which for Indic data is most of the work: orthographic convention, numeral form, transliteration and romanised code-switch, punctuation, casing, inverse text normalisation, and lexicon maintenance — inconsistent conventions show up as WER that has nothing to do with the model
Run large-scale pseudo-labelling and weak supervision — transcribe unlabelled audio with existing models, filter on confidence and cross-model agreement, and route the uncertain cases to human review
Version and catalogue everything: reproducible dataset snapshots, training manifests, dataset cards recording source, licence, consent basis, and known limitations
Annotation, listening & quality
Own transcription and labelling quality end to end — write the annotation guidelines, train annotators against them, and revise them when reality disagrees
Listen to audio and annotate yourself, as needed. Sampling real data is how you find the problems dashboards hide, and it is part of this job rather than beneath it
Measure annotation quality properly — gold sets, inter-annotator agreement, WER between independent transcriptions — and manage vendors on quality and cost per hour, not volume alone
Build the TTS-specific quality gates: alignment tightness, transcript fidelity, speaker and style consistency, prosody and emotion tagging, and codec round-trip checks
Coordinate listening panels for subjective evaluation (MOS / CMOS / preference tests) — recruitment, screening, sample randomisation, and result analysis
Benchmarks & evaluation sets
Own Gnani’s internal benchmark suite for ASR and TTS — curated, frozen, versioned, speaker-disjoint from training data, and representative across language, dialect, accent, channel, and domain
Guarantee split hygiene. No test speaker in train, no leaked utterance, no silently mutated eval set — and be willing to block a result that violates it
Track external and public benchmarks, reproduce them faithfully, and flag when a published comparison is not apples to apples
Analyse results, not just produce them — slice error by language, speaker, channel, and domain, and tell the team where the corpus is failing them
Tooling & internal UIs
Write the Python that holds all of this together — ingestion and processing scripts, validation checks, catalogue tooling, and reporting
Build small internal web tools with modern coding assistants: annotation and QC interfaces, dataset explorers, audio A/B listening and diffing tools, and demo UIs for internal reviews and customer conversations
Turn recurring manual work into tooling. If you have done it by hand three times, it should be a script or a screen by the fourth
Governance, security & customers
Treat voice data as sensitive PII by default — voice recordings, transcripts, and speaker embeddings, the last of which are biometric data and carry the highest sensitivity
Build de-identification and redaction into the pipeline: account and card numbers, Aadhaar and PAN, names, addresses, and health or financial detail out of transcripts before anyone works with them
Maintain provenance and consent records well enough to survive an audit — what data came from where, under what licence or DPA, retained how long, and deleted on what trigger
Apply data standards and recognised governance practice to how corpora are documented, accessed, retained, and disposed of — aligned with DPDP and, where relevant, RBI, IRDAI, and HIPAA expectations
Work with the security team on access control, encryption at rest and in transit, least-privilege data rooms, and audit logging — and support customer security reviews and data-handling questions directly
Being the single point of contact
Be the person the team comes to for data. Take requests from ASR, TTS, LLM, and product, understand what is actually needed, and deliver the right subset with the right metadata
Communicate clearly and often across a team of researchers and engineers with different needs and different vocabularies, and keep everyone working from one shared picture of what data exists
Say no when a request would compromise split hygiene, consent, or compliance — and explain why in terms the requester accepts
Who you are
EXPERIENCE WE’RE LOOKING FOR
3–6 years in data engineering, data analysis, or ML data work — with real ownership of a dataset that other people depended on
Strong, practical Python: pandas or polars, scripting, automation, and comfort working with files and pipelines at scale rather than only in notebooks
Genuine ownership instinct. This role has no one above it to catch a data problem — you are that person, and you should want to be
Excellent written and verbal communication, and the patience to work across several stakeholders with competing data needs
Careful, detail-obsessed temperament — the kind of person who spots that two files have subtly different transcript conventions before anyone trains on them
WHAT MAKES A STANDOUT CANDIDATE
Hands-on experience with audio or speech data — segmentation, alignment, transcription workflows, or annotation operations
Fluency in one or more Indian languages beyond English, and an ear for dialect and code-switching
Experience running annotation vendors or an in-house labelling team, including quality measurement and cost management
Comfort building small web UIs and demos — Streamlit, Gradio, or a lightweight FastAPI plus React app, coding assistants very much welcome
Exposure to data governance and privacy in a regulated setting — DPDP, consent management, PII redaction, or supporting customer security reviews
Comfortable operating with high ownership in a fast-moving, post–Series B environment
TECHNICAL FLUENCY — A MUST-HAVE
Python & data: pandas / polars, PyArrow and Parquet, JSONL manifests, SQL, and clean reusable scripting
Audio tooling: ffmpeg, sox, librosa or torchaudio — resampling, segmentation, loudness, and basic signal sanity checks
Pipelines & storage: object storage, sharded formats for training throughput, workflow orchestration (Airflow, Prefect, or Dagster), Git, Docker
Dataset discipline: versioning and snapshots, train / dev / test split design, speaker-disjoint splits, deduplication, dataset documentation
Analysis & reporting: WER and CER computation and error slicing, coverage and gap reporting, agreement metrics, and dashboards people actually read
Nice to have: NVIDIA NeMo manifest conventions, forced alignment tools, Spark or Ray for large jobs, and basic familiarity with how ASR and TTS models consume data
Skills Required
Primary Skills
Annotation Formatting
NLP/STT
Azure TTS
Data Preparation
ASR/STT APIs
Data Cleaning
Frequently Asked Questions
How to apply for Data Engineer(Speech/Language Data) at Gnani Innovations?
Click the "Apply via CareerScan" button on this page.
What is the salary for this role?
Salary details will be discussed during the interview.
What experience is required?
4 years of experience is required.
Is this position still open?
Yes, currently active and accepting applications.
Similar Openings
Explore related active roles in data annotation
Data Engineer(Speech/Language Data)
Gnani Innovations · Remote