Search by job, company or skills

Site Reliability Engineer

This job is no longer accepting applications

Job Description

Job Title: Lead Site Reliability Engineer (SRE)

Location: Bangalore, Hyderabad & Chennai (Hybrid)

Experience: 8–14 Years

Job Summary

We are looking for a Lead Site Reliability Engineer (SRE) to lead a team of 5–6 engineers in a 24x7 production support environment. The ideal candidate should have strong experience in Incident Management, Linux, Kubernetes, Cloud (AWS/Azure/GCP), DevOps, Automation, and Production Operations with the ability to drive reliability, automation, and continuous improvement.

Key Responsibilities

  • Lead a team of 5–6 SREs and coordinate daily operations.
  • Act as Incident Commander during critical production incidents.
  • Drive RCA, incident triage, and rapid issue resolution.
  • Build automation to reduce manual operational tasks.
  • Monitor system health and improve platform reliability.
  • Troubleshoot Linux, Kubernetes, Cloud, Networking, and distributed systems.
  • Prepare weekly/monthly KPI reports and collaborate with global engineering teams.

Mandatory Skills

  • 8+ years in Site Reliability Engineering (SRE) / Production Support
  • Team Lead experience
  • Incident Management, RCA & Incident Command
  • Linux Administration
  • Kubernetes & Docker
  • AWS / Azure / GCP
  • DevOps & CI/CD (GitHub/GitLab/Harness/Jenkins)
  • Automation using Python/Shell/Bash
  • Networking (TCP/IP, DNS, HTTP/HTTPS, Load Balancers)
  • Experience with Java/Node.js applications is a plus

Preferred

  • Experience in Fortune 500 or large enterprise environments
  • Akamai/CDN knowledge
  • Hybrid/Multi-cloud infrastructure experience
  • AI-assisted development tools (Cursor, Claude)

Looking for professionals who can thrive in high-pressure production environments, lead critical incidents, and drive automation to improve system reliability.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151081043

Similar Jobs

Hyderabad, India

Skills:

JavaGolangGoogle Cloud PlatformKafkaSparkAzureKubernetesPythonAWSAirflowFlinkdbtMLFlowLarge Language Models

Hyderabad, India

Skills:

DevopsSite Reliability Engineering

Hyderabad

Skills:

MlJavaSpring BootJenkinsDockerTerraformECSGitlabKubernetesPythonAWSOpen TelemetryAiSite Reliability Engineering

Hyderabad, India

Skills:

MlJavaSpring BootJenkinsDockerTerraformECSGitlabKubernetesPythonAWSOpen TelemetryAiSite Reliability Engineering

Hyderabad, India

Skills:

TcpUDPDnsRtpLoad TestingGcpLoad BalancingTlsPythonincident communicationautomatic rollbackNATS-class busescapacity modelingGocanary analysisWebSocket fleetsAI-aware reliabilitymulti-cloud literacySIPstreaming pipelineschaos engineering