Search by job, company or skills

AI Engineer Agentic AI, RAG & Low-Latency Voice Systems | Okhla | 5+ yrs

  • Posted 33 minutes ago
  • Be among the first 10 applicants

Job Description

AI Engineer — Agentic AI, RAG & Low-Latency Voice Systems

Mobiloitte Technologies (I) Pvt. Ltd. | New Delhi (on-site / hybrid) | Full-time | 5+ years

Compensation: No ceiling for the right candidate.

Two things decide this role: how well you handle latency, and how far past basic RAG you can build.

Mobiloitte builds production AI for enterprise clients across India, the US, UK, UAE, Singapore and South Africa. We're hiring a senior AI Engineer to own our agentic AI, retrieval and voice stack — systems that go live, carry real call volume, and stay up.

1. LATENCY IS THE JOB, NOT A DETAIL

Our voice bots handle inbound and outbound calls where a 400ms difference decides whether the conversation feels human. We need someone who treats latency as an engineering discipline:

- Owns end-to-end voice latency from user speech to first token of bot audio — and can quote the numbers they achieved

- Profiles and attacks each hop: STT, retrieval, inference, TTS, network

- Streaming responses, speculative/partial generation, response chunking to shorten time-to-first-audio

- Barge-in and interruption handling that actually works mid-sentence

- Model caching, warm pools, batching, concurrency control under real call load

- Knows when the fix is architectural, not a bigger GPU

If you can't tell us what your p95 latency was and what you did to bring it down, this isn't the role.

2. AGENTIC AI — BEYOND BASIC RAG

Single-shot retrieve-and-answer is table stakes. We're building systems that plan, act and self-correct:

- Agent orchestration — LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, or your own framework (we care that you've built it, not which one)

- Tool and function calling, API/DB actions, structured outputs

- Multi-step reasoning, task decomposition, planning loops with retry and fallback

- Multi-agent patterns — routing, supervisor/worker, handoffs

- Conversation and long-term memory management across sessions

- Advanced retrieval: agentic and query-rewriting RAG, hybrid search, reranking, graph/multi-hop retrieval, chunking strategy chosen deliberately

- Evaluation and guardrails — hallucination reduction, groundedness scoring, tracing and observability (LangSmith, Langfuse or similar)

- Cost control: token budgeting, context-window discipline, model routing

3. VOICE — INBOUND AND OUTBOUND

- Full pipeline: STT → LLM/agent/RAG → TTS

- Whisper / Faster-Whisper / Deepgram; Piper / Coqui / XTTS or equivalent

- SIP and telephony integration, WebSockets, real-time streaming

- Built the stack rather than wrapping Vapi, Retell or Bland end to end

4. PRIVATE / SELF-HOSTED DEPLOYMENT

Many of our clients cannot send data to public APIs. You should be able to:

- Deploy Llama, Qwen or Mistral inside a client VPC or on-prem

- Run vLLM, Ollama or TensorRT-LLM in production

- Size GPUs and VRAM, apply INT8/INT4 quantization, tune inference

- Handle concurrency, autoscaling, secure API exposure, auth and data isolation

- Explain how you'd replace an OpenAI dependency with a self-hosted architecture — this earns strong preference

ON COMPENSATION

We've deliberately not published a band. For an engineer who has genuinely built agentic systems, driven latency down in production, and run models on their own infrastructure, compensation is not the constraint — we'll match what the work is worth.

Be honest with yourself before applying: if your AI work is mainly calling third-party APIs and writing prompts, we'll find that out in the first twenty minutes. If you've built agents, fought latency, and put a model on your own GPU, we want to talk.

HOW TO APPLY

Apply here and include one thing: a link, repo, demo video or short architecture note for a system you personally built — and tell us what you did on it.

Thanks

Team-HR

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153918583

Similar Jobs

Delhi, India

Skills:

PythonLangChainLLM Application DevelopmentAutoGen or equivalentAgentic Multi-step AI WorkflowsREST API IntegrationRAG Retrieval Augmented GenerationAI Telemetry MonitoringPrompt Engineering

Noida, India

Skills:

AWS FargateNumpyAWS SNSPythonAws LambdaPandasAzure DevopsAWS SQSAzure PipelinesGrogAdvanced GenAI Agentic Framework ConceptsCI CDAzure AI FoundryAWS EventBridgeWorkflow Agentic FrameworksCloud Application Integration DeploymentVector DatabasesAzure CLILangChainFine tuning Model CustomizationAzure OpenAI ServiceAI Agents Tool CallingAI Search IndexAgentic AI SystemsAWS KinesisAmazon BedrockPrompt Engineering

Beware of Scammers

We don’t charge money for job offers