AI Engineering Lead Agentic AI Platform (Telecom BSS/OSS)
the right hire limited- Posted 3 hours ago
- Be among the first 10 applicants
Job Description
Hiring Location: Pune, Mumbai, Bengaluru
On-site: UAE
We are hiring a Technical Engineering Lead for an AI-native digital transformation platform at [a Tier-1 telecom operator in the UAE]. In line with the UAE National AI Strategy, the operator is building agentic AI and sovereign compute into its core connectivity and BSS/OSS platforms. You will be the platform's senior technical authority: you own the design and the delivery, and you lead the squads who build it.
This is not a hands-off management role. You will design multi-agent workflows and review the code that ships them.
WHAT YOU'LL DO
Design the platform
• Design cloud-native, secure, scalable architecture that brings LLMs, multi-agent systems and autonomous workflows into Network Operations (NOC), Service Assurance and customer lifecycle management.
• Design agent networks (LangGraph, LangChain, AutoGen) that run real-time, intent-based network and BSS workflows: automated recovery playbooks, predictive fault management, Zero-Touch Provisioning and autonomous network slicing.
• Align designs with TM Forum ODA and CAMARA APIs, and integrate telco-specific LLMs for alarm correlation, complex event processing and network analytics.
• Deploy across regional sovereign cloud and hybrid environments (AWS, Azure OpenAI, OCI).
Lead engineering delivery
• Own the technical roadmap end to end, and give architecture- and code-level guidance in Python, Go and Node.js.
• Make AI-assisted development (Cursor, GitHub Copilot) the team's standard way of working, with clear standards for prompt and context engineering, prompt versioning, semantic chunking and RAG pipelines.
• Partner with product managers, data scientists, network domain experts and senior stakeholders to turn telecom business problems into technical blueprints.
Make AI production-grade
• Build LLM evaluation into CI/CD: automated test beds that score models on hallucination, toxicity, compliance, latency and context recall before anything is released.
• Redefine QA for non-deterministic systems, with probabilistic testing of agents that handle billing, SIM activation and VVIP service assurance.
• Run LLMOps at scale (Databricks or Kubeflow): feedback loops, intent classification models, fine-tuning and guardrail validation.
• Instrument multi-step reasoning chains with tracing (OpenInference, Arize Phoenix, LangSmith), and control token cost through prompt caching, dynamic model routing and throughput scaling while meeting local latency targets.
Govern and protect
• Make sure every data layer complies with TDRA regulations and sovereign data-privacy requirements.
• Put real-time guardrails in place (e.g. NeMo Guardrails) for data masking, jailbreak prevention and responsible-AI behaviour.
Build the team
• Lead, mentor and upskill software and AI engineers on engineering quality and responsible AI, and help develop early-career talent through initiatives such as the AI Graduate Programme.
WHAT YOU'LL BRING
Essential
• [12+] years of hands-on software engineering, including [4+] years as a Tech Lead or technical design authority leading squads on complex production systems, ideally at a Tier-1 telecom operator or a large enterprise.
• A track record of delivering RAG and agentic systems in production: LangChain/LangGraph or LlamaIndex, vector databases (Milvus, Pinecone, Qdrant or similar), and LLMOps/MLOps pipelines.
• Strong Python, plus Node.js and/or Go.
• Deep experience with AWS, Azure or GCP, plus Kubernetes, Docker, service meshes and Terraform across hybrid environments.
• Working knowledge of telecom architecture in at least one area: BSS/OSS stacks, 5G and network slicing, NFV/SDN, service assurance, or event-driven telemetry and alarm correlation.
• Daily use of AI coding assistants, and evidence that you have raised a team's productivity with them.
• A record of coaching engineers and building strong teams.
Highly valued
• Hands-on TM Forum ODA / Open APIs or CAMARA experience.
• LLM evaluation, observability and guardrail tooling running in production.
• Fine-tuning or intent-classification work on Databricks or Kubeflow.
• Delivery in sovereign or regulated cloud environments, ideally in the GCC.
WHY THIS ROLE
• Production, not proofs of concept: agents running live network and customer operations on a national network.
• A mandate to set the engineering and AI standards for the platform from day one.
• Visible, strategic work tied directly to a national AI agenda.
HOW TO APPLY
Apply via LinkedIn or send your CV to [Confidential Information].
In a few lines, tell us about one agentic or RAG system you took to production: what it did, the scale it ran at, and one thing that went wrong and how you fixed it.
#Hiring #AgenticAI #GenAI #LLMOps #Telecom #BSS #OSS #EngineeringLeadership #UAEJobs
