Senior Site Reliability Engineer
Senior Site Reliability Engineer
Falabella IndiaEarly Applicant
- Posted 2 months ago
- Be among the first 30 applicants
Job Description
We are looking for an experienced Senior Site Reliability Engineer (SRE) to join our platform engineering team. The ideal candidate will be responsible for designing, building, and operating highly available, scalable, secure, and cost-efficient cloud infrastructure platforms. This role requires strong expertise in Kubernetes, cloud platforms, observability, automation, incident management, and reliability engineering practices. The Senior SRE will collaborate closely with software engineering, security, infrastructure, and operations teams to improve system reliability, performance, scalability, and developer productivity.
Reliability Engineering
The core responsibilities for the job include the following:
Reliability Engineering
The core responsibilities for the job include the following:
- Design, implement, and maintain highly available and resilient production systems.
- Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Drive reliability improvements through automation and engineering best practices.
- Perform capacity planning and performance optimisation for critical systems.
- Conduct failure analysis and implement preventive measures.
- Design, deploy, and manage large-scale Kubernetes environments.
- Manage cloud infrastructure across AWS, GCP, or Azure environments.
- Implement Infrastructure as Code (IaC) using Terraform.
- Manage service mesh platforms such as Istio.
- Optimise infrastructure utilisation and cloud costs.
- Build and maintain monitoring, logging, and alerting platforms.
- Implement observability solutions using tools such as Prometheus, Grafana, Datadog, VictoriaMetrics, and Loki.
- Develop dashboards, alerts, and automated operational workflows.
- Monitor system performance, availability, latency, and business metrics.
- Lead production incident response and troubleshooting activities.
- Participate in on-call rotations and major incident management.
- Conduct root cause analysis (RCA) and post-incident reviews.
- Drive continuous improvements to reduce operational toil.
- Establish operational runbooks and automation frameworks.
- Develop automation scripts and tooling using Python, Go, or Bash.
- Implement CI/CD pipelines and GitOps workflows.
- Build self-service infrastructure capabilities for engineering teams.
- Automate operational tasks and infrastructure provisioning.
- Collaborate with security teams to implement DevSecOps practices.
- Ensure infrastructure compliance with organisational standards.
- Implement secure access controls, secrets management, and audit processes.
- Support vulnerability management and remediation efforts.
- Monitor cloud spending and identify optimisation opportunities.
- Implement rightsizing, reserved instance, and committed-use strategies.
- Analyse infrastructure costs and recommend cost-saving initiatives.
- Partner with engineering teams to improve resource efficiency.
- Bachelor's degree in computer science, engineering, or a related field.
- 4+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
- Strong experience managing production Kubernetes environments.
- Hands-on expertise with public cloud platforms (AWS, GCP, or Azure).
- Strong knowledge of Linux systems administration and networking.
- Experience with Infrastructure as Code tools such as Terraform.
- Expertise in observability and monitoring platforms.
- Experience implementing CI/CD and GitOps practices.
- Strong scripting and programming skills (Python, Go, Bash).
- Experience with incident management and root cause analysis.
- Experience operating large-scale distributed systems.
- Experience with service mesh technologies (Istio, Linkerd).
- Knowledge of database reliability engineering (PostgreSQL, MySQL, Redis).
- Experience with security and compliance frameworks.
- FinOps and cloud cost optimisation experience.
- Experience managing multi-region and multi-cloud environments.
- Kubernetes certifications (CKA, CKAD, CKS) are preferred.
- Cloud Platforms: AWS, Google Cloud Platform (GCP), Microsoft Azure.
- Container and Platform: Kubernetes, Docker, Helm, Istio.
- Infrastructure as Code: Terraform, Ansible.
- Observability: Prometheus, Grafana, Datadog, Elasticsearch, VictoriaMetrics, Loki.
- CI/CD and GitOps: GitHub Actions, GitLab CI, Jenkins, ArgoCD.
- Databases: PostgreSQL, MySQL, Redis.
- Programming: Python, Go, Bash.
- Strong troubleshooting and analytical skills.
- Excellent communication and stakeholder management abilities.
- Ability to lead complex technical initiatives.
- Strong ownership mindset and operational excellence.
- Ability to mentor engineers and drive engineering best practice.





