Search by job, company or skills

Senior Staff Engineer(TTS)

  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

Senior Staff Engineer, Text-to-Speech (TTS)

www.gnani.ai

Location: Bengaluru Type: Full-time(work from Office) Function: AI Team

About the role

While you're reading this, Gnani is talking to customers across India — in Hindi, Tamil, Telugu, and dozens more languages, switching between them mid-sentence without missing a beat. That voice is our TTS stack. It's the difference between an AI that sounds robotic and one that sounds human, expressive, and unmistakably Indian.

We're building the most natural voices in Indian languages, at production latency and enterprise scale — and we want you to lead the research behind them. As Senior Staff Engineer for TTS, you own the stack end-to-end: advancing quality and expressiveness, engineering it for real-time performance, and building commercial-grade voice cloning. This is a senior individual-contributor role: you set technical direction, design experiments, and mentor the researchers around you — without a people-management load.

Core mandate

Research

Advance TTS quality, expressiveness, and multilingual coverage — experiment across architectures

Production & performance

Take models to production with low latency, high throughput, and reliability at scale

Voice cloning & customization

Build commercial-grade cloning and voice design for media and enterprise use

Technical leadership

Guide juniors, design rigorous experiments, and own the standard for how TTS gets built here

What you'll drive

Research & modeling

  • Own the TTS modeling roadmap — architecture selection, quality, expressiveness, and multilingual / code-switched coverage across Indian languages
  • Experiment across TTS architectures (autoregressive / LM-based, non-autoregressive, diffusion / flow-matching, neural-codec–based) and translate findings into shippable systems
  • Push on the hard problems: prosody and emotion control, naturalness, code-switching (e.g. Hinglish), G2P for Indic scripts, voice design, audio tags, and speaker / style consistency
  • Define and track quality evaluation — MOS / CMOS, intelligibility (WER via ASR), speaker similarity — and make architecture calls grounded in data
  • Participate in benchmarks and publish the work in international conferences

Production & performance engineering

  • Design TTS systems for production: streaming synthesis, low first-audio latency, favourable RTF, and high concurrency on GPU
  • Own the performance envelope — batching, quantization, kernel / inference optimization, and cost-per-hour of audio
  • Optimize production inference using ONNX, TensorRT, NVIDIA Triton, quantization, batching, and GPU-efficient serving architectures
  • Partner with platform and infra teams on serving, scaling, and reliability — latency SLAs, uptime, error budgets
  • Build evaluation and regression harnesses so quality and latency don't silently regress release to release

Voice cloning & customer quality

  • Build and improve commercial-grade voice cloning for media and enterprise applications — few-shot / zero-shot cloning, voice design from a brief, and controllable style
  • Work directly with customers to resolve production issues, fine-tune voice quality, and address cloning fidelity, consent, and safety concerns
  • Turn recurring customer quality issues into systematic model and pipeline improvements

Technical leadership & research presence

  • Technically guide junior researchers and engineers — experiment design, tracking, code and research review, and career-shaping mentorship
  • Set the bar for rigor: reproducible experiments, clean baselines, honest evaluation, and clear write-ups
  • Stay current with TTS research and bring the relevant frontier into Gnani's roadmap
  • Publish and represent Gnani's TTS work externally — international conferences, talks, and community engagement

Who you are

EXPERIENCE WE'RE LOOKING FOR

  • 8–10 years in speech / ML research or engineering, with substantial contribution to production-grade TTS systems
  • A track record of taking TTS from research to production with a focus on latency, quality, and performance at scale
  • Hands-on experience across multiple TTS architectures — you've built, compared, and shipped, not just read papers
  • Experience building commercial-grade voice cloning, ideally for media applications
  • Good publications at international conferences (e.g. Interspeech, ICASSP, NeurIPS, ICML, ACL) in speech / audio / ML
  • Strong system-design instincts for the engineering challenges of productionizing speech models

WHAT MAKES A STANDOUT CANDIDATE

  • Deep familiarity with the Indian voice landscape — Indic languages, code-switching, dialect and accent diversity
  • Demonstrated ability to guide juniors and lead experiment design and tracking across a team
  • Up-to-date with TTS research, and able to separate durable advances from hype
  • Experience with neural codecs / tokenizers (RVQ and similar) and LLM-style TTS
  • Comfortable operating with high ownership in a fast-moving, post–Series B environment

TECHNICAL FLUENCY — A MUST-HAVE

  • Deep understanding of modern TTS: acoustic modeling, neural vocoders and codecs, autoregressive vs non-autoregressive synthesis, diffusion / flow approaches, LLM / codec-token TTS, and prosody / duration / pitch control
  • Fluency with TTS evaluation — MOS / CMOS, WER (intelligibility via ASR), RTF and latency, speaker similarity for cloning — and what drives improvements or regressions on each
  • Working knowledge of production ML systems: streaming inference, GPU serving, batching, quantization, and the latency / quality / cost trade-offs that decide what ships
  • Judgment on evaluation and benchmark design: good test sets, fair comparison across languages / accents / domains, and not overfitting to a benchmark

Why Gnani, why now

Gnani.ai is one of the few companies in India doing original speech research at depth — our models are trained on real Indian language data at scale, not adapted from English-first architectures. As Senior Staff Engineer for TTS, you'll own one of the most technically demanding surfaces in Indian voice AI, with the freedom to set direction and the impact of seeing your model's power real conversational products across banking, healthcare, logistics, and government services.

High visibility, high impact, and a clear path to the most senior technical roles as the company scales.

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 151733447