Search by job, company or skills

Applied AI Scientist - TTS

Early Applicant
  • Posted 7 days ago
  • Be among the first 10 applicants

Job Description

About the Role:

We are building state-of-the-art Voice AI technologies for Indic languages, enabling natural, intelligent, and scalable voice interactions through Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Small/Large Language Models (SLMs/LLMs). As an Applied AI Scientist specializing in TTS, you will own the production success of multilingual speech synthesis models. You will continuously improve speech quality, optimize latency, resolve production issues, and deliver natural, expressive voice experiences.

If you enjoy owning AI systems end-to-end, solving real-world problems, and delivering measurable customer impact, we'd love to hear from you.

Key Responsibilities

Own Production AI Systems

  • Own the technical performance of assigned AI models throughout their lifecycle.
  • Continuously improve model quality, robustness, latency, scalability, and production reliability.
  • Monitor model performance and proactively identify opportunities for improvement.

Deliver AI Model Releases

  • Plan, execute, validate, and support regular model releases across multiple languages and product lines.
  • Ensure models meet quality benchmarks and production readiness standards.
  • Collaborate with Engineering and DevOps teams to enable seamless deployments.

Solve Customer-Centric AI Problems.

  • Investigate production issues, customer feedback, and model failures.
  • Perform root cause analysis and implement effective solutions.
  • Design and execute experiments to improve model performance and user experience.

Cross-functional Collaboration

  • Work closely with Product Managers to translate customer needs into AI improvements.
  • Partner with Engineering teams to integrate AI models into production systems.
  • Collaborate with Research Scientists to transition research innovations into production-ready solutions.

Experimentation & Optimization

  • Design, execute, and analyze experiments to improve AI models.
  • Build evaluation pipelines and benchmark model performance.
  • Optimize models for inference speed, resource efficiency, and scalability.

Technical Ownership

  • Take end-to-end ownership of assigned initiatives.
  • Document technical designs, experiments, and implementation decisions.
  • Share knowledge and contribute to engineering best practices across the AI team.

Domain Responsibilities: TTS

  • Own production quality of the LLM-backbone TTS pipeline end-to-end (tokens → flow-matching decoder → vocoder or vocoder-free path); isolate root cause across stages when issues arise.
  • Optimize/own both streaming (low time-to-first-audio) and non-streaming paths; maintain distilled low-NFE decoders and watch for quality drift.
  • Evaluate adopting vocoder-free approaches (e.g., Supertonic-style) where they beat the current pipeline on latency/quality/simplicity.
  • Own and maintain vocabulary control (pronunciation/lexicon overrides, OOV handling) and emotion control (expressive/style conditioning) as customer-facing features, fixing mispronunciations and expressiveness issues in production.
  • Own the inference/serving stack for the TTS pipeline; deploying and tuning model servers (e.g., Triton/PyTriton, vLLM-style serving for the LM component) for throughput, latency, and GPU utilization.
  • Scale voice cloning across languages/speakers; build automatic eval pipelines (speaker similarity, ASR-based, UTMOS-style) and run MOS/CMOS/MUSHRA panels for releases.

Required Skills & Qualifications

  • Bachelor's or Master's degree or PhD in Computer Science, Artificial Intelligence, Machine Learning, or a related field.
  • 2–4 years of hands-on experience developing machine learning or deep learning systems.
  • Experience deploying or supporting AI models in production environments.
  • Strong proficiency in Python.
  • Strong experience with PyTorch or equivalent deep learning frameworks.
  • Solid understanding of machine learning and deep learning fundamentals.
  • Experience training, fine-tuning, evaluating, and debugging AI models.
  • Familiarity with Linux, Git, and GPU-based development environments.
  • Experience working with large datasets and data processing pipelines.
  • Understanding of model evaluation methodologies and performance metrics.
  • Exposure to cloud platforms and modern AI development workflows is desirable.

Preferred Qualifications

  • Experience in Speech AI, Generative AI, NLP, or Large Language Models.
  • Experience optimizing AI models for production environments.
  • Familiarity with distributed training or inference optimization.
  • Contributions to open-source projects, technical blogs, or research publications are a plus.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151619115

Beware of Scammers

We don’t charge money for job offers