Search Jobs

Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

Recro
Early Applicant
  • Posted 15 hours ago
  • Be among the first 10 applicants

Job Description

Site Reliability Engineer

Experience: 3+ years

Location:Bangalore

Shift:Rotationalshifts

Role Overview

We are looking for an SRE to ensure the reliability, scalability, availability, and observability of high-traffic production systems. The role involves infrastructure automation, Kubernetes operations, monitoring, incident management, and continuous improvement of production reliability.

Key Responsibilities

  • Manage and troubleshoot AWS/GCP infrastructure and Kubernetes environments in production.
  • Work with Kubernetes, Redis, Kafka, Solr/Elasticsearch and related infrastructure components.
  • Build and maintain CI/CD pipelines, Terraform/Helm-based infrastructure, and automation.
  • Implement monitoring and alerting using Prometheus, Grafana, ELK/Loki and distributed tracing tools.
  • Participate in rotational on-call and shifts, handling P1/P2 incidents, troubleshooting, RCA, and post-incident reviews.
  • Develop SLIs/SLOs, alerts, dashboards, runbooks, and reliability improvements.
  • Automate repetitive operational tasks using Python, Shell/Bash, or Go to reduce manual toil.
  • Work on capacity planning, autoscaling, performance optimization, security patching, and cost optimization.
  • Troubleshoot Linux, networking, DNS, TCP/IP, load balancing, TLS/HTTPS, and application/infrastructure issues.
  • Drive preventive actions through RCA, automation, self-healing, and improved deployment/recovery processes.

Must-Have Skills

  • 3+ years of experience in SRE / DevOps / Infrastructure Engineering
  • Experience with high-traffic or large-scale production environments
  • Hands-on Kubernetes in production
  • Strong experience with AWS or GCP
  • Terraform and Infrastructure as Code (IaC)
  • Docker and CI/CD
  • Prometheus & Grafana
  • Linux administration and troubleshooting
  • Python / Bash / Shell scripting
  • Production incident management, RCA and on-call experience
  • Good understanding of networking fundamentals

Good to Have

  • Helm, ArgoCD/GitOps
  • ELK/Loki
  • Redis, Kafka, Solr/Elasticsearch
  • SLI/SLO, SLA and error-budget concepts
  • Ansible
  • Distributed tracing / OpenTelemetry

Note: This is a rotational-shift/on-call role, so candidates should be comfortable supporting production systems across different shifts.

More Info

Job Type:
Industry:
Employment Type:

About Company

Similar Jobs

5-7 yrs
Bengaluru, India
Skills:
Linux Shell Scripting, Problem Management, Linux Administration, Devops, Gcp, metrics, Terraform, Ansible, Agile, Azure, Python, Kubernetes, Logging, AWS, Infrastructure as Code, Alerting, Production Operations, Monitoring
2-4 yrs
Bengaluru, India
Skills:
agile environment , Continuous Delivery, Mq, Performance Monitoring, Monitoring Tools, SQL Server, Kafka, Grafana, Programming Languages, Geneos, Continuous Integration, Dynatrace, Splunk, Kubernetes, AWS, Capacity Management, Copilot AI Tools, Oracle DB2, software engineering concepts
5-7 yrs
Bengaluru, India
Skills:
Automation Tools, Apis, Terraform, Kubernetes, AWS, data network and information security best practices, DevOps tooling
3-5 yrs
Bengaluru, India
Skills:
Elk, Prometheus, Bash, Grafana, Git, Terraform, Linux, Kubernetes, Python, AWS, Go, OpenSearch
5-7 yrs
Bengaluru, India
Skills:
Linux Shell Scripting, Gcp, Terraform, Linux, Ansible, Azure, Kubernetes, Python, AWS