Azure SRE (Site Reliability Engineer)
Tata Consultancy Services- Posted 20 hours ago
- Be among the first 10 applicants
Job Description
Azure SRE (Site Reliability Engineer)
Experience: 6-12 Years
Location: Chennai / Kochi
Job Description
We are seeking an experienced Azure Site Reliability Engineer (SRE) to manage, automate, and optimize cloud infrastructure and application reliability on Microsoft Azure. The ideal candidate will have strong expertise in Azure platform services, automation, observability, incident management, and DevOps practices to ensure highly available, scalable, and secure enterprise applications.
Key Responsibilities
- Design, build, and maintain reliable and scalable Azure cloud infrastructure.
- Implement SRE best practices focused on availability, performance, resiliency, and operational excellence.
- Automate infrastructure provisioning and deployment using Infrastructure as Code (IaC).
- Manage CI/CD pipelines and deployment automation.
- Establish monitoring, logging, alerting, and observability frameworks.
- Drive incident response, root cause analysis, and problem management activities.
- Collaborate with development teams to improve application reliability and performance.
- Define and track SLIs, SLOs, and SLAs for critical services.
- Implement security, governance, and compliance best practices across Azure environments.
Mandatory Skills
- Microsoft Azure
- Azure Kubernetes Service (AKS)
- Azure DevOps
- Terraform / Infrastructure as Code (IaC)
- CI/CD Pipelines
- Docker & Kubernetes
- Monitoring & Observability
- Azure Monitor
- Log Analytics
- Application Insights
- Incident & Problem Management
- Scripting (PowerShell/Python/Bash)
- Linux Administration
- Networking & Security Fundamentals
Good to Have
- GitOps
- Prometheus & Grafana
- ELK Stack
- Azure Landing Zone
- Azure Site Recovery
- Azure Automation
- Azure Policies & Governance
- FinOps & Cost Optimization
Required Experience
- 6+ years of experience in Cloud Engineering, DevOps, or Site Reliability Engineering.
- Hands-on experience managing production workloads on Azure.
- Strong expertise in infrastructure automation and cloud operations.
- Experience supporting containerized applications and Kubernetes platforms.
- Proven experience in monitoring, incident management, and reliability engineering.
- Strong communication and stakeholder management skills.
More Info
Key Skills
CI CD Pipelines
Monitoring Observability
Scripting PowerShell Python Bash
Log Analytics
Azure Kubernetes Service AKS
Terraform Infrastructure as Code IaC
Networking Security Fundamentals
Azure Monitor
Application Insights
Incident Problem Management
