Search by job, company or skills

Site Reliability Engineer

Early Applicant
  • Posted 3 days ago
  • Be among the first 10 applicants

Job Description

Qualifications

• 7+ years of experience in SRE, Infrastructure Engineering, Network Engineering, or Platform Operations.

• Strong hands-on experience with Kubernetes/OpenShift.

• Experience supporting Cisco SDN, IP networking, and BGP.

• Background leading major incident management and operational escalations.

• Experience with GitOps, automation, and Infrastructure-as-Code.

• Proven technical leadership and team mentoring experience.

• Ability to support on-call and shift-based environments.

Nice to Have Skills & Experience

Preferred

• Telecommunications or large-scale enterprise platform experience.

• Experience working with globally distributed teams.

• Familiarity with observability, AIOps, and reliability engineering best practices.

Job Description

Seeking a Lead Senior Reliability Engineer to provide technical leadership for a 24x5 Service Reliability Center supporting a cloud-native operational platform. This role serves as the onshore lead for an offshore Tier 1.5 Reliability Engineering team, driving incident management, automation, platform reliability, and GitOps-based operations. Responsibilities • Lead P1/P2 incident response and serve as Incident Commander during major outages. • Provide technical leadership across Kubernetes/OpenShift, Cisco SDN, IP networking, BGP, Linux, and distributed applications. • Guide an offshore Tier 1.5 engineering team responsible for alarm monitoring, troubleshooting, and incident resolution. • Partner closely with the Insight Global Program Manager to support service delivery, operational governance, reporting, and program success. • Act as the primary day-to-day technical liaison with the client's engineering, operations, and DevOps stakeholders. • Review and approve production changes, MOPs, and operational readiness activities. • Partner with Engineering and DevOps teams to resolve complex platform issues and escalations. • Build automation, runbooks, and self-healing capabilities to improve platform reliability. • Drive adoption of GitOps and Infrastructure-as-Code practices. • Mentor engineers and lead cross-functional operational initiatives.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152079231

Similar Jobs

Hyderabad, India

Skills:

ApisTerraformDockerAnsibleMicrosoft AzureArmKubernetesSqlScriptingBicep

Hyderabad, India

Skills:

Oracle APEXOracle Database ManagementArtifactoryJIRAAngularJenkinsGitLinuxOracle Pl SqlDockerAnsibleSonarqubeFlaskMongoDBPythonAI ML workflowMS SQL-Server

Hyderabad, India

Skills:

VMwareSanNasFTPKvmPrometheusElk StackDnsGrafanaWindowsEfsDHCPTerraformDockerNfsPythonAWSCloudformationSmtpGcpShellSftpAnsibleRhelCentosPuppetAzureKubernetesHyper-VNTPOpenTelemetryRocky Linux

Hyderabad, India

Skills:

DockerDistributed SystemsIncident ManagementKubernetesinfinibandAI infrastructureCloud platformsGPUsReliability engineering

Hyderabad

Skills:

PythonAwsAnsibleKubernetesGitSRE

Beware of Scammers

We don’t charge money for job offers