Search by job, company or skills

Senior Site Reliability Engineer

Early Applicant
  • Posted 5 days ago
  • Be among the first 20 applicants

Job Description

We are seeking a Senior Site Reliability Engineer with 5+ years of experience to lead the reliability engineering strategy for our core banking infrastructure and distributed transaction processing applications. You will design self-healing architectures, establish SLOs/SLAs, lead major incident resolution, and drive zero-downtime architecture for mission-critical financial platforms.

Key Responsibilities

  • Reliability Architecture & Design: Partner with software architects to design fault-tolerant, multi-region distributed banking systems capable of processing millions of daily transactions.
  • SLO & Error Budget Management: Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets alongside product managers to balance feature velocity with system stability.
  • Advanced Automation & IaC: Implement Infrastructure as Code (IaC) using Terraform, Ansible, and Kubernetes operators. Build automated self-healing mechanisms to resolve known failure modes without human intervention.
  • Incident Leadership & Postmortems: Lead response for high-priority (P1/P2) banking outages. Facilitate blameless post-mortems and enforce long-term root cause remediations.
  • Capacity & Chaos Engineering: Forecast system growth, conduct load testing under peak banking hours, and run Chaos Engineering experiments (Gremlin, Chaos Mesh) to uncover hidden vulnerabilities.
  • Compliance & Security: Ensure infrastructure complies with banking regulatory frameworks (PCI-DSS, SOC2, Central Bank guidelines) and lead automated compliance auditing tools.

Required Qualifications & Skills

  • Education: Btech or Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience.
  • Deep Experience: 5+ years in SRE, DevOps, or Software Engineering, with at least 2 years in banking, fintech, or high-volume transactional environments.
  • Orchestration & Cloud: Expert-level knowledge of Kubernetes (CKA certified preferred) and public/hybrid cloud enterprise architectures (AWS/Azure/GCP).
  • Software Development: Strong programming skills in Go, Python, or Java for building internal SRE tooling, CLI utilities, and automated controllers.
  • Data Stores: Understanding of high-availability relational (PostgreSQL, Oracle) and distributed non-relational databases (Cassandra, Redis, Kafka).
  • Observability: Mastery of ELK/EFK stack, OpenTelemetry, Prometheus, Cortex/Thanos, and distributed tracing (Jaeger/Zipkin).

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153692907

Similar Jobs

Hyderabad, India

Skills:

ApisGrafanaSqlDockerTerraformMicrosoft AzureItil ProcessesArmScriptingKubernetesLog Analyticsobservability stacksBicepApplication InsightsAzure Monitor

Hyderabad, India

Skills:

.NETJavaPrometheusNodejsAzure Log AnalyticsGrafanaJIRADatadogGcpDockerSplunkAzureKubernetesPythonAWSPagerDuty

Hyderabad

Skills:

PythonDevopsAutomation

Hyderabad

Skills:

AnsiblePythonKubernetessite reliabilitySlurmGPU computing

Hyderabad, India

Skills:

PowerShellBashJenkinsGcpDockerTerraformAzureKubernetesPythonAWSGoArgo CD

Beware of Scammers

We don’t charge money for job offers