Search by job, company or skills

Lead Voice AI Engineer

Early Applicant
  • Posted a day ago
  • Be among the first 10 applicants

Job Description

About the role

We are looking for a Lead Voice AI Engineer to build production-grade Voice Agents for frontline customer and employee support. You will lead the design of low-latency, real-time voice systems combining ASR, TTS, LLMs, conversational AI, enterprise workflows, knowledge retrieval, compliance, and human handoff.

This is a hands-on technical leadership role for someone who can take Voice AI from architecture to production.

Responsibilities

  • Design and build the real-time voice runtime for live conversations.
  • Build and optimize streaming ASR, TTS, VAD, endpointing, turn-taking, and barge-in.
  • Develop Voice Agents that support multi-turn, multi-intent conversations, context switching, clarification, and recovery.
  • Integrate Voice Agents with workflows, APIs, CRM, ITSM, knowledge bases, and enterprise systems.
  • Build secure identity verification, consent, privacy, audit, and compliance controls.
  • Implement warm transfer, callback, queue routing, and seamless human handoff with full conversation context.
  • Optimize multilingual voice quality across accents, noisy environments, latency, and naturalness.
  • Build evaluation frameworks for WER, intent accuracy, response latency, containment, resolution, escalation, and CSAT.
  • Establish production observability across the full call path: ASR → LLM → tools → TTS.
  • Evaluate and integrate leading speech, telephony, and AI technologies.
  • Define architecture, engineering standards, and production readiness for the Voice AI platform.
  • Mentor engineers and lead critical technical design reviews.

Minimum qualifications

  • 7+ years of software engineering experience.
  • Strong experience building production distributed or real-time systems.
  • Hands-on experience with Conversational AI, Voice AI, Speech AI, or LLM-based agents.
  • Strong programming skills in Python, Java, Go, or equivalent.
  • Experience with APIs, streaming systems, asynchronous architectures, and cloud-native platforms.
  • Strong understanding of system design, scalability, reliability, and observability.

Preferred qualifications

  • Experience with Deepgram for real-time ASR and streaming speech recognition.
  • Experience with LiveKit for WebRTC, real-time audio, voice-agent runtime, and session orchestration.
  • Experience with ElevenLabs for low-latency, natural TTS and conversational voice experiences.
  • Experience with OpenAI, Azure Speech, Google Speech, or similar ASR/TTS technologies.
  • Experience with WebRTC, SIP, RTP, WebSockets, Twilio, or contact-center platforms.
  • Experience with LLM agents, tool calling, RAG, LangGraph, or similar orchestration frameworks.
  • Experience integrating enterprise systems such as Salesforce, ServiceNow, Jira, Zendesk, or Workday.
  • Experience with multilingual speech, accent handling, noisy environments, PII redaction, and call-recording controls.
  • Experience building high-scale, multi-tenant SaaS platforms.

What success looks like

  • Voice Agents feel natural and responsive in real-time conversations.
  • Users can interrupt naturally and change context without breaking the conversation.
  • The system works reliably across languages, accents, and noisy environments.
  • Voice Agents securely execute enterprise workflows and grounded knowledge retrieval.
  • Complex cases escalate to humans with full context.
  • The platform meets measurable targets for latency, accuracy, reliability, containment, resolution, and customer satisfaction.

Engineering principle

Voice is not chat with audio. Voice is a real-time interaction model with different requirements for latency, interruption, identity, compliance, failure handling, and human handoff.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153452147

Beware of Scammers

We don’t charge money for job offers