Site Reliability Engineer

Highradius

Hyderabad, India

5-8 Years

Save

Posted a day ago
Be among the first 10 applicants

Early Applicant

Job Description

Job Summary:

We are looking for a highly skilled and adaptable Site Reliability Engineer (5 - 8 Years) to become a key member
of our Cloud Engineering team. In this crucial role, you will be instrumental in designing and refining our cloud
infrastructure with a strong focus on reliability, security, and scalability. As an SRE, you'll apply software
engineering principles to solve operational challenges, ensuring the overall operational resilience and continuous
stability of our systems. This position requires a blend of managing live production environments and contributing
to engineering efforts such as automation and system improvements.

Key Responsibilities:

. Cloud Infrastructure Architecture and Management: Design, build, and maintain resilient cloud
infrastructure solutions to support the development and deployment of scalable and reliable applications.
This includes managing and optimizing cloud platforms for high availability, performance, and cost
efficiency.
. Enhancing Service Reliability: Lead reliability best practices by establishing and managing monitoring and
alerting systems to proactively detect and respond to anomalies and performance issues. Utilize SLI, SLO,
and SLA concepts to measure and improve reliability. Identify and resolve potential bottlenecks and
areas for enhancement.
. Driving Automation and Efficiency: Contribute to the automation, provisioning, and standardization of
infrastructure resources and system configurations. Identify and implement automation for repetitive tasks
to significantly reduce operational overhead. Develop Standard Operating Procedures (SOPs) and
automate workflows using tools like Rundeck or Jenkins.
. Incident Response and Resolution: Participate in and help resolve major incidents, conduct thorough root
cause analyses, and implement permanent solutions. Effectively manage incidents within the production
environment using a systematic problem-solving approach.
. Collaboration and Innovation: Work closely with diverse stakeholders and cross-functional teams,
including software engineers, to integrate cloud solutions, gather requirements, and execute Proof of
Concepts (POCs). Foster strong collaboration and communication. Guide designs and processes with a
focus on resilience and minimizing manual effort. Promote the adoption of common tooling and
components, and implement software and tools to enhance resilience and automate operations. Be
open to adopting new tools and approaches as needed.

Required Skills and Experience:

. Cloud Platforms: Demonstrated expertise in at least one major cloud platform (AWS, Azure, or GCP).
. Containerization Tools: Extensive experience with containerization (Docker) and orchestration (Kubernetes) technologies.
. Automation & IaC: Proficiency in scripting languages (shell and Python). Experience with configuration
management tools (Ansible or Puppet). Must have exposure to Infrastructure as Code (IaC) tools
(Terraform or CloudFormation).
. Monitoring & Observability: Experience setting up and configuring monitoring tools (Prometheus, Grafana,
or the ELK stack). Hands-on experience implementing OpenTelemetry for observability. Familiarity with
monitoring and logging tools for cloud-based applications.
. Service Reliability Concepts: A strong understanding of SLI, SLO, SLA, and error budgeting.
. Infrastructure Management: Proven proficiency in on-premises hosting and virtualization platforms
(VMware, Hyper-V, or KVM). Solid understanding of storage internals (NAS, SAN, EFS, NFS) and protocols
(FTP, SFTP, SMTP, NTP, DNS, DHCP). Experience with networking and firewall technologies. Strong hands-on
experience with Linux internals and operating systems (RHEL, CentOS, Rocky Linux). Experience with
Windows operating systems to support varied environments.
. Soft Skills & Mindset: Excellent communication and interpersonal skills for effective teamwork. We value
proactive individuals who are eager to learn and adapt in a dynamic environment. Must possess a
pragmatic and adaptable mindset, with a willingness to step outside comfort zones and acquire new
skills. Ability to consider the broader system impact of your work. Must be a change advocate for reliability.