Search by job, company or skills

Senior Technical Lead

8-12 Years
Early Applicant
  • Posted 6 days ago
  • Be among the first 10 applicants

Job Description

Job Description

We are looking for an AI Platform Senior Lead to join our engineering team and take ownership of the infrastructure, pipelines, and platform capabilities that power our AI and LLM-based solutions. This is a hands-on engineering role for someone who thrives at the intersection of cloud infrastructure, data engineering, and AI systems — not someone who builds models, but someone who builds the platforms that make AI systems reliable, observable, and scalable in production.

You will work closely with application teams, data scientists, and product stakeholders to ensure AI workloads run efficiently, cost-effectively, and at scale.

What You Will Do:

AI Platform & Infrastructure

  • Design, build, and maintain scalable platforms that support LLM inference, embedding pipelines, RAG systems, guardrails, and multi-tenant AI services
  • Architect and manage APIs and microservices that abstract AI capabilities for consumption by application teams
  • Own platform reliability — uptime, latency SLAs, cost per request, and performance benchmarks
  • Build and maintain multi-tenant infrastructure with strong isolation, quota management, and usage tracking across teams and clients

Data Engineering & Processing

  • Design and implement data ingestion, transformation, and processing pipelines that feed AI systems
  • Build and manage vector databases, document stores, and retrieval infrastructure for RAG and semantic search use cases
  • Ensure data quality, lineage, and governance across AI data pipelines
  • Optimize data pipelines for throughput, latency, and cost at scale

Cloud & DevOps

  • Own cloud infrastructure on AWS — including ECS/EKS, Lambda, API Gateway, S3, RDS, SQS, and AI/ML services like Bedrock and SageMaker
  • Implement Infrastructure as Code using Terraform or CDK
  • Build CI/CD pipelines for AI workloads including model serving, evaluation, and deployment automation
  • Manage Kubernetes clusters and containerised workloads for AI services

Observability & Operations

  • Implement comprehensive observability for AI systems — logging, tracing, metrics, cost dashboards, and alerting
  • Build evaluation pipelines to monitor LLM output quality, hallucination rates, latency, and token consumption in production
  • Proactively identify and resolve performance bottlenecks, reliability gaps, and cost inefficiencies

Collaboration & Technical Leadership

  • Partner with product, architecture, and application teams to translate requirements into scalable platform capabilities
  • Define and enforce platform engineering standards, patterns, and best practices
  • Mentor junior engineers and contribute to technical design reviews

Must Have

  1. 8–12 years of hands-on software and platform engineering experience
  2. Strong proficiency in Python — async patterns, API development, microservices
  3. Deep experience with cloud platforms — AWS preferred (ECS/EKS, Lambda, API Gateway, S3, SQS, RDS)
  4. Solid data engineering background — pipeline design, ETL/ELT, vector databases (Pinecone, Weaviate, pgvector), document processing
  5. Experience building and operating production AI or LLM-based systems — API gateways, embedding pipelines, RAG architectures, inference serving
  6. Strong understanding of multi-tenancy, rate limiting, quota management, and tenant isolation patterns
  7. Experience with containerisation and orchestration — Docker, Kubernetes, Helm
  8. Hands-on with observability tooling — Prometheus, Grafana, OpenTelemetry, Datadog or similar
  9. Infrastructure as Code experience — Terraform or AWS CDK
  10. Hands-on experience with LLM orchestration frameworks — LangChain, LangGraph, or similar — to understand agentic patterns, tool use, and multi-agent workflows at a platform level
  11. Knowledge of AI guardrails, content filtering, and responsible AI patterns in production — including toxicity filtering, PII detection, and output validation

Good To Have

  1. Experience with LLM evaluation frameworks — RAGAS, LangSmith, or custom eval pipelines
  2. Experience with voice AI pipelines or multimodal systems
  3. AWS certifications — Solutions Architect, ML Engineer, or equivalent

Mandatory Skills: Python , Prometheus , Docker

Years Of Experience: 8 to 12 Years

More Info

Job Type:
Industry:
Employment Type:

Job ID: 151489525

Similar Jobs

Gurugram, Gurugram, India

Skills:

.Net CoreAws LambdaAWS GlueSQL ServerMicroservicesAws RdsAWS IAMAWS DMSAWS EMREntity Framework EF CoreAgentic AIAWS CloudFrontAWS CMIAWS DynamoDB

Noida, India

Skills:

AngularCosmos DBKubernetesDockerGitAzure servicesAzure SQL ServerAzure Data ExplorerAzure IoT HubAzure Event Hub