Senior Staff Engineer, Text-to-Speech (TTS)
www.gnani.ai
Location: Bengaluru Type: Full-time(work from Office) Function: AI Team
About the role
While you're reading this, Gnani is talking to customers across India — in Hindi, Tamil, Telugu, and dozens more languages, switching between them mid-sentence without missing a beat. That voice is our TTS stack. It's the difference between an AI that sounds robotic and one that sounds human, expressive, and unmistakably Indian.
We're building the most natural voices in Indian languages, at production latency and enterprise scale — and we want you to lead the research behind them. As Senior Staff Engineer for TTS, you own the stack end-to-end: advancing quality and expressiveness, engineering it for real-time performance, and building commercial-grade voice cloning. This is a senior individual-contributor role: you set technical direction, design experiments, and mentor the researchers around you — without a people-management load.
Core mandate
Research
Advance TTS quality, expressiveness, and multilingual coverage — experiment across architectures
Production & performance
Take models to production with low latency, high throughput, and reliability at scale
Voice cloning & customization
Build commercial-grade cloning and voice design for media and enterprise use
Technical leadership
Guide juniors, design rigorous experiments, and own the standard for how TTS gets built here
What you'll drive
Research & modeling
- Own the TTS modeling roadmap — architecture selection, quality, expressiveness, and multilingual / code-switched coverage across Indian languages
- Experiment across TTS architectures (autoregressive / LM-based, non-autoregressive, diffusion / flow-matching, neural-codec–based) and translate findings into shippable systems
- Push on the hard problems: prosody and emotion control, naturalness, code-switching (e.g. Hinglish), G2P for Indic scripts, voice design, audio tags, and speaker / style consistency
- Define and track quality evaluation — MOS / CMOS, intelligibility (WER via ASR), speaker similarity — and make architecture calls grounded in data
- Participate in benchmarks and publish the work in international conferences
Production & performance engineering
- Design TTS systems for production: streaming synthesis, low first-audio latency, favourable RTF, and high concurrency on GPU
- Own the performance envelope — batching, quantization, kernel / inference optimization, and cost-per-hour of audio
- Optimize production inference using ONNX, TensorRT, NVIDIA Triton, quantization, batching, and GPU-efficient serving architectures
- Partner with platform and infra teams on serving, scaling, and reliability — latency SLAs, uptime, error budgets
- Build evaluation and regression harnesses so quality and latency don't silently regress release to release
Voice cloning & customer quality
- Build and improve commercial-grade voice cloning for media and enterprise applications — few-shot / zero-shot cloning, voice design from a brief, and controllable style
- Work directly with customers to resolve production issues, fine-tune voice quality, and address cloning fidelity, consent, and safety concerns
- Turn recurring customer quality issues into systematic model and pipeline improvements
Technical leadership & research presence
- Technically guide junior researchers and engineers — experiment design, tracking, code and research review, and career-shaping mentorship
- Set the bar for rigor: reproducible experiments, clean baselines, honest evaluation, and clear write-ups
- Stay current with TTS research and bring the relevant frontier into Gnani's roadmap
- Publish and represent Gnani's TTS work externally — international conferences, talks, and community engagement
Who you are
EXPERIENCE WE'RE LOOKING FOR
- 8–10 years in speech / ML research or engineering, with substantial contribution to production-grade TTS systems
- A track record of taking TTS from research to production with a focus on latency, quality, and performance at scale
- Hands-on experience across multiple TTS architectures — you've built, compared, and shipped, not just read papers
- Experience building commercial-grade voice cloning, ideally for media applications
- Good publications at international conferences (e.g. Interspeech, ICASSP, NeurIPS, ICML, ACL) in speech / audio / ML
- Strong system-design instincts for the engineering challenges of productionizing speech models
WHAT MAKES A STANDOUT CANDIDATE
- Deep familiarity with the Indian voice landscape — Indic languages, code-switching, dialect and accent diversity
- Demonstrated ability to guide juniors and lead experiment design and tracking across a team
- Up-to-date with TTS research, and able to separate durable advances from hype
- Experience with neural codecs / tokenizers (RVQ and similar) and LLM-style TTS
- Comfortable operating with high ownership in a fast-moving, post–Series B environment
TECHNICAL FLUENCY — A MUST-HAVE
- Deep understanding of modern TTS: acoustic modeling, neural vocoders and codecs, autoregressive vs non-autoregressive synthesis, diffusion / flow approaches, LLM / codec-token TTS, and prosody / duration / pitch control
- Fluency with TTS evaluation — MOS / CMOS, WER (intelligibility via ASR), RTF and latency, speaker similarity for cloning — and what drives improvements or regressions on each
- Working knowledge of production ML systems: streaming inference, GPU serving, batching, quantization, and the latency / quality / cost trade-offs that decide what ships
- Judgment on evaluation and benchmark design: good test sets, fair comparison across languages / accents / domains, and not overfitting to a benchmark
Why Gnani, why now
Gnani.ai is one of the few companies in India doing original speech research at depth — our models are trained on real Indian language data at scale, not adapted from English-first architectures. As Senior Staff Engineer for TTS, you'll own one of the most technically demanding surfaces in Indian voice AI, with the freedom to set direction and the impact of seeing your model's power real conversational products across banking, healthcare, logistics, and government services.
High visibility, high impact, and a clear path to the most senior technical roles as the company scales.