Search Jobs

Search by job, company or skills

AI Engineer Agentic AI, RAG & Low-Latency Voice Systems | Okhla | 5+ yrs

AI Engineer Agentic AI, RAG & Low-Latency Voice Systems | Okhla | 5+ yrs

Mobiloitte
Early Applicant
  • Posted a day ago
  • Be among the first 10 applicants

Job Description

AI Engineer — Agentic AI, RAG & Low-Latency Voice Systems

Mobiloitte Technologies (I) Pvt. Ltd. | New Delhi (on-site / hybrid) | Full-time | 5+ years

Compensation: No ceiling for the right candidate.

Two things decide this role: how well you handle latency, and how far past basic RAG you can build.

Mobiloitte builds production AI for enterprise clients across India, the US, UK, UAE, Singapore and South Africa. We're hiring a senior AI Engineer to own our agentic AI, retrieval and voice stack — systems that go live, carry real call volume, and stay up.

1. LATENCY IS THE JOB, NOT A DETAIL

Our voice bots handle inbound and outbound calls where a 400ms difference decides whether the conversation feels human. We need someone who treats latency as an engineering discipline:

- Owns end-to-end voice latency from user speech to first token of bot audio — and can quote the numbers they achieved

- Profiles and attacks each hop: STT, retrieval, inference, TTS, network

- Streaming responses, speculative/partial generation, response chunking to shorten time-to-first-audio

- Barge-in and interruption handling that actually works mid-sentence

- Model caching, warm pools, batching, concurrency control under real call load

- Knows when the fix is architectural, not a bigger GPU

If you can't tell us what your p95 latency was and what you did to bring it down, this isn't the role.

2. AGENTIC AI — BEYOND BASIC RAG

Single-shot retrieve-and-answer is table stakes. We're building systems that plan, act and self-correct:

- Agent orchestration — LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, or your own framework (we care that you've built it, not which one)

- Tool and function calling, API/DB actions, structured outputs

- Multi-step reasoning, task decomposition, planning loops with retry and fallback

- Multi-agent patterns — routing, supervisor/worker, handoffs

- Conversation and long-term memory management across sessions

- Advanced retrieval: agentic and query-rewriting RAG, hybrid search, reranking, graph/multi-hop retrieval, chunking strategy chosen deliberately

- Evaluation and guardrails — hallucination reduction, groundedness scoring, tracing and observability (LangSmith, Langfuse or similar)

- Cost control: token budgeting, context-window discipline, model routing

3. VOICE — INBOUND AND OUTBOUND

- Full pipeline: STT → LLM/agent/RAG → TTS

- Whisper / Faster-Whisper / Deepgram; Piper / Coqui / XTTS or equivalent

- SIP and telephony integration, WebSockets, real-time streaming

- Built the stack rather than wrapping Vapi, Retell or Bland end to end

4. PRIVATE / SELF-HOSTED DEPLOYMENT

Many of our clients cannot send data to public APIs. You should be able to:

- Deploy Llama, Qwen or Mistral inside a client VPC or on-prem

- Run vLLM, Ollama or TensorRT-LLM in production

- Size GPUs and VRAM, apply INT8/INT4 quantization, tune inference

- Handle concurrency, autoscaling, secure API exposure, auth and data isolation

- Explain how you'd replace an OpenAI dependency with a self-hosted architecture — this earns strong preference

ON COMPENSATION

We've deliberately not published a band. For an engineer who has genuinely built agentic systems, driven latency down in production, and run models on their own infrastructure, compensation is not the constraint — we'll match what the work is worth.

Be honest with yourself before applying: if your AI work is mainly calling third-party APIs and writing prompts, we'll find that out in the first twenty minutes. If you've built agents, fought latency, and put a model on your own GPU, we want to talk.

HOW TO APPLY

Apply here and include one thing: a link, repo, demo video or short architecture note for a system you personally built — and tell us what you did on it.

Thanks

Team-HR

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Agent orchestration

Faster-Whisper

LangSmith

Whisper

Tool and function calling API

LangGraph

INT4 quantization

OpenAI Agents SDK

WebSockets

Langfuse

Qwen

Mistral

Coqui XTTS

Multi-step reasoning

TensorRT-LLM

CrewAI

SIP and telephony integration

Advanced retrieval

vLLM

AutoGen

Ollama

Llama

Deepgram

Piper

About Company

Similar Jobs

2-5 yrs
Delhi, India
Skills:
PythonLangChainLLM Application DevelopmentAutoGen or equivalentAgentic Multi-step AI WorkflowsREST API IntegrationRAG Retrieval Augmented GenerationAI Telemetry MonitoringPrompt Engineering
8-10 yrs
Gurugram, Gurugram, India
Skills:
Power AutomatePower PlatformCopilot Agent BuilderGenAI-powered assistantsLLMsagentic AI conceptsPower AppsMicrosoft GraphMicrosoft 365 ecosystemprompt engineeringSharepointTeamsagent orchestrationCopilot Studioconversational systems
6-8 yrs
Gurugram, Gurugram, India
Skills:
PostgreSQLKafkaRedisELTDockerFastAPIKubernetesPythonEtlLangChainHuggingFaceGPTWhisperRasa
5-9 yrs
Noida, India
Skills:
UipathNode.jsAI MLAngularReactRESTGcpDockerFlaskFastAPIJavascriptPythonKubernetesChatbotAgentGRPCLlmWebSocketsAgentic AIRAGCI/CDStreamlitLangraph
4-6 yrs
Delhi, India
Skills:
ServicenowAmazon Web ServicesGoogle Cloud PlatformTerraformAzureLangChainAI orchestration frameworksRegula test casesRAG vector databasesRego policiesAnthropicPagerDutyAgentic modelsLangGraphLLM APIsOpen Policy AgentGenerative AI workflowsOpenAICloud-native security servicesAtlassian Suite