Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

People Group
2-7 Years
Not Disclosed
Quick Apply
  • Posted a month ago
  • Over 200 applicants have applied

Job Description

Key Responsibilities:

  • System Reliability: Monitor, maintain, and enhance system uptime and availability, minimizing downtime.
  • Infrastructure as Code (IaC): Design, implement, and manage infrastructure using tools such as CloudFormation, Terraform, Ansible, or Puppet.
  • Automation: Develop and maintain CI/CD pipelines and deployment scripts to streamline software releases.
  • Containerization: Manage and orchestrate application containers using Docker Swarm and AWS ECS.
  • Monitoring and Alerting: Set up and maintain monitoring tools like CloudWatch, Datadog, Zenduty, and New Relic for proactive issue resolution.
  • Scalability and Performance: Optimize application and infrastructure performance collaboratively with development teams.
  • Security: Implement and maintain security best practices across development and operations pipelines.
  • Incident Management: Participate in incident response, root cause analysis, and preventive measures.
  • Documentation: Maintain clear documentation of system architecture, deployment processes, and best practices.
  • Collaboration: Facilitate communication and knowledge sharing between development, operations, and other teams.

Qualifications:

  • Bachelor's degree in Computer Science, Information Technology, or related field.
  • Minimum 2 years of experience in DevOps, System Operations, or SRE roles.
  • Strong Linux knowledge and shell scripting skills.
  • Proficiency in AWS cloud platform and AWS services.
  • Understanding of networking, security, and secure infrastructure best practices.
  • Knowledge of ELK stack and Kafka.
  • Experience with Docker Swarm or AWS ECS for containerization.
  • Hands-on experience with CloudFormation, Terraform, Ansible.
  • Familiarity with CI/CD pipelines and version control (GitLab CI, Jenkins, Git).
  • Working knowledge of databases (MySQL, Postgres).
  • Willingness to participate in L1 incident response rotation.
  • AWS certifications (e.g., AWS Certified DevOps Engineer, AWS Certified Solutions Architect) are a plus.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

About Company

People Interactive, a world-class consumer internet company in India, is built on an Idea to help Indians find a life partner. Our vision is to build the worlds first togetherness company! We want to build a leading platform to find a life partner, discover love and share joy. It was conceived not as a business but as an idea to make people find their match for life, leveraging the digital space most competently. We are doing it with our two brands

Similar Jobs

8-10 yrs
Mumbai, India
Skills:
Cloudformation, Prometheus, Grafana, Datadog, Jenkins, Cloudwatch, Terraform, Docker, ECS, Gitlab, Dynatrace, Splunk, Kubernetes, Networking Technologies, AWS cloud services, container and container orchestration, enterprise-authorized AI capabilities, reliability engineering workflows, telemetry collection, continuous integration and continuous delivery tools, infrastructure automation, Python programming language, white and black box monitoring, observability, SLO alerting
3-5 yrs
Mumbai, India
Skills:
.NET, Sql, React, Git, Rest Apis, Postman, Terraform, Arm, Azure DevOps, Log Analytics, Application Insights, Azure Monitor, Infrastructure as Code, Bicep
6-8 yrs
Mumbai, India
Skills:
Amazon Web Services, Grafana, Autosys, Sql, Linux, Shell scripting, Dynatrace, Databricks, Splunk, Kubernetes, Python, Airflow, OpenTelemetry, Amazon Managed Workflows for Apache Airflow
6-8 yrs
Mumbai
Skills:
Autosys, Databricks, Sql, Splunk, Shell scripting, Amazon Web Services, Grafana, Dynatrace, Linux, Python, Kubernetes, Airflow, OpenTelemetry, Amazon Managed Workflows for Apache Airflow
3-5 yrs
Mumbai, India
Skills:
New Relic, Confluence, Dynatrace, Bash, Grafana, Jira, Datadog, Sql, Python, Sumo Logic