T
SRE(Linux,Terfform,Kubernet,observability,Any cloud)
T
SRE(Linux,Terfform,Kubernet,observability,Any cloud)
Tata Consultancy Services6-10 Years
Early Applicant
- Posted 11 hours ago
- Be among the first 10 applicants
Job Description
Hiring: Site Reliability Engineer (SRE) | Bangalore
We are looking for an experienced Site Reliability Engineer (SRE) with strong hands-on expertise in Linux, Terraform, Kubernetes, Observability, and Cloud platforms.
If you are passionate about building highly available, scalable, automated, and reliable cloud infrastructure, we'd love to connect with you!
Job Details
- Role: Site Reliability Engineer (SRE)
- Experience: 6–10 Years
- Location: Bangalore
- Notice Period: Immediate Joiners / Up to 30 Days
- Key Skills: Linux, Terraform, Kubernetes, Observability, Cloud
Key Responsibilities
- Build, manage, and maintain highly available and scalable production environments.
- Manage Linux-based infrastructure and troubleshoot system, application, network, and performance issues.
- Develop and maintain Infrastructure as Code using Terraform.
- Deploy, manage, monitor, and troubleshoot containerized workloads on Kubernetes.
- Implement effective monitoring, logging, alerting, and observability solutions.
- Define and monitor SLIs, SLOs, availability, reliability, and system performance.
- Automate repetitive operational activities using scripting and DevOps practices.
- Support incident management, root cause analysis, and production troubleshooting.
- Collaborate with Development, DevOps, Cloud, Security, and Infrastructure teams.
- Identify reliability risks and continuously improve platform resilience and operational efficiency.
Required Skills
- 6–10 years of relevant experience in SRE, DevOps, Cloud, or Infrastructure Engineering.
- Strong hands-on experience with Linux administration and troubleshooting.
- Strong expertise in Terraform and Infrastructure as Code (IaC).
- Hands-on experience with Kubernetes and container technologies.
- Strong understanding of Observability, Monitoring, Logging, and Alerting.
- Experience with at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Understanding of CI/CD pipelines, automation, networking, and cloud infrastructure.
- Experience with monitoring tools such as Prometheus, Grafana, ELK/Elastic Stack, Splunk, Datadog, or similar is desirable.
- Scripting knowledge in Python, Bash, or Shell is preferred.
- Strong troubleshooting, analytical, and problem-solving skills.
More Info
Key Skills
Cloud platforms
ELK Elastic Stack
Google Cloud Platform (GCP)
CI/CD pipelines
Observability
Shell



