Company Name: WillwareTechnologies
Role: Site Reliability Engineer (SRE)
Experience:4+ Years
Location: Chennai, Bangalore,Hyderabad,Gurugram,Noida,Pune,
WorkMode: Onsite
Job Summary
We are looking for a skilled Site Reliability Engineer (SRE) with 4-6 years of experience to ensure the reliability, availability, scalability, and performance of enterprise applications and infrastructure. The ideal candidate should have hands-on experience in cloud platforms, automation, monitoring, incident management, and DevOps practices.
Key Responsibilities
- Ensure high availability, reliability, and performance of production systems.
- Monitor application and infrastructure health using observability and monitoring tools.
- Automate operational tasks using scripting and Infrastructure as Code (IaC).
- Troubleshoot production incidents, perform root cause analysis (RCA), and implement preventive measures.
- Manage incident response, problem management, and on-call support.
- Collaborate with Development, DevOps, Cloud, and Infrastructure teams to improve system reliability.
- Implement monitoring, alerting, logging, and performance tuning solutions.
- Support CI/CD pipelines and deployment automation.
- Improve system scalability, capacity planning, and disaster recovery processes.
- Develop and maintain operational runbooks, dashboards, and documentation.
Required Skills
- 4-6 years of experience in Site Reliability Engineering (SRE) or DevOps.
- Strong knowledge of Linux/Unix administration.
- Experience with cloud platforms (AWS, Azure, or GCP).
- Hands-on experience with Kubernetes and Docker.
- Knowledge of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
- Experience with Infrastructure as Code (Terraform, Ansible, or CloudFormation).
- Experience with monitoring and logging tools such as Prometheus, Grafana, ELK, Splunk, Datadog, or New Relic.
- Proficiency in scripting using Python, Bash, or Shell scripting.
- Experience with Git version control.
- Strong understanding of networking concepts (DNS, HTTP, TCP/IP, Load Balancers).
- Knowledge of incident management, RCA, and production support.
- Experience working in Agile/Scrum environments.
Preferred Skills
- Site Reliability Engineering (SRE) principles and best practices.
- Kubernetes administration and container orchestration.
- Microservices architecture.
- Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Observability and distributed tracing.
- ITIL processes.
- Cloud security and compliance knowledge.
Education
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
Preferred Certifications
- AWS Certified Solutions Architect / SysOps Administrator
- Microsoft Azure Administrator
- Google Professional Cloud Engineer
- Certified Kubernetes Administrator (CKA)
- Terraform Associate
- Red Hat Certified System Administrator (RHCSA)