Search by job, company or skills

Forward Deployment Engineer (SRE)

  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

Senior Associate – Forward Deployment Engineer (SRE)

Site Reliability Engineering | Forward Deployed Engineering

Location: Bangalore / Hyderabad

Experience Required

6–9 years.

Job Summary

A senior, customer-facing SRE who owns the reliability of a client's production systems end to end. Embedded as the main technical point of contact, you will design reliable infrastructure, lead incidents, apply AI-assisted operations, and mentor the team.

Key Responsibilities

  • Own the reliability of production systems for one or more enterprise customers on AWS.
  • Define and manage SLIs, SLOs, and error budgets, and ensure monitoring provides meaningful insight.
  • Lead the response to serious (P0/P1) incidents and run blameless post-mortems that lead to fixes.
  • Lead within the on-call rotation.
  • Operate and optimise Kubernetes clusters and AWS services as workloads grow.
  • Introduce AI SRE tooling and AIOps where it speeds up triage and resolution, with appropriate guardrails.
  • Advise customers on improving their reliability practices, and mentor associate engineers.
  • Share field learnings with product and engineering teams.

Required Qualifications

  • Substantial SRE experience with real ownership of production reliability on AWS.
  • Experience running Kubernetes and AWS services at scale.
  • A track record of defining and managing SLIs, SLOs, and error budgets.
  • Strong observability practice.
  • Proven incident leadership and post-mortem facilitation.
  • Strong automation skills (Python, Go, or Bash).
  • Hands-on understanding of AI-assisted operations, including introducing AI SRE tooling with sensible guardrails.
  • AWS certification at Associate level as a minimum.

Preferred Qualifications

  • Owning CI/CD pipelines and release automation.
  • Writing and maintaining Terraform modules.
  • GitOps workflows and Helm.
  • Designing telemetry collection across services.
  • Incident-management tooling such as PagerDuty.
  • Hands-on experience integrating AIOps or AI SRE tooling into production operations.
  • Familiarity with observability or evaluation of AI systems (e.g., Langfuse).
  • Experience delivering AI-based work or AI/ML-driven initiatives in production.
  • AWS Certified Solutions Architect – Professional and/or AWS Certified DevOps Engineer – Professional.

Technical Skills & Tools

  • Cloud (AWS): EC2, EKS, ECS, Lambda, S3, RDS, VPC, IAM, CloudWatch, ELB/ALB, Route 53
  • Containers & orchestration: Docker, Kubernetes (EKS), Helm
  • Observability & monitoring: Prometheus, Grafana, Datadog, OpenTelemetry, Loki, Tempo/Jaeger (tracing), ELK/Elastic
  • Reliability practices: SLIs/SLOs, error budgets, incident command, capacity planning, failure-mode analysis
  • Incident & on-call: PagerDuty, Opsgenie, blameless post-mortems
  • Automation & scripting: Python, Go, Bash, Git
  • AI-assisted operations: AI SRE tooling, AIOps, anomaly detection, automated RCA
  • DevOps (good to have): CI/CD (GitHub Actions, GitLab CI), Terraform modules, GitOps (ArgoCD), Helm

Soft Skills & Competencies

  • Confident, customer-facing communication.
  • Calm, decisive incident leadership.
  • Mentors and raises standards across the team.
  • Sound technical judgement.

More Info

Job Type:
Industry:
Employment Type:

Job ID: 152916595

Beware of Scammers

We don’t charge money for job offers