Search by job, company or skills

6-8 Years
Early Applicant
Quick Apply
  • Posted 8 days ago
  • Be among the first 40 applicants

Job Description

Key Responsibilities:

  • Monitoring & Alerting:
  • Develop, maintain, and enhance monitoring and alerting systems using Datadog to proactively identify and address potential issues, ensuring optimal system performance.
  • CI/CD Pipelines:
  • Participate in the design and implementation of CI/CD pipelines using Azure DevOps, enabling automated and reliable software delivery.
  • Incident Response:
  • Lead efforts in incident response and troubleshooting to quickly diagnose and resolve production incidents, minimizing downtime and impact on users.
  • Reliability Initiatives:
  • Take ownership of reliability initiatives by identifying areas for improvement, conducting root cause analysis, and implementing solutions to prevent recurrence of incidents.
  • Collaboration:
  • Collaborate with cross-functional teams to ensure security, compliance, and performance standards are met throughout the development lifecycle.
  • On-call Support:
  • Participate in on-call rotations and provide 24/7 support for critical incidents, ensuring rapid response and resolution.
  • SLOs & SLIs:
  • Work with development teams to define and establish Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to measure and maintain system reliability.
  • Documentation:
  • Contribute to the documentation of processes, procedures, and best practices to enhance knowledge sharing within the team.

Qualifications:

  • Education:
  • Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent work experience.
  • Experience:
  • Minimum of 4 years of experience in a Site Reliability Engineer or similar role, managing cloud-based infrastructure on AWS with EKS.
  • AWS Expertise:
  • Strong expertise in AWS services, especially EKS, including cluster provisioning, scaling, and management.
  • Monitoring & Observability:
  • Proficiency in using monitoring and observability tools, with hands-on experience in Datadog or similar tools for tracking system performance and generating meaningful alerts.
  • CI/CD Experience:
  • Experience in implementing CI/CD pipelines using Azure DevOps or similar tools to automate software deployment and testing.
  • Containerization & Orchestration:
  • Solid understanding of containerization and orchestration technologies (e.g., Docker, Kubernetes) and their role in modern application architectures.
  • Troubleshooting:
  • Excellent troubleshooting skills and the ability to analyze complex issues, determine root causes, and implement effective solutions.
  • Scripting & Automation:
  • Strong scripting and automation skills (e.g., Python, Bash).
  • IaC (Infrastructure as Code):
  • Familiarity with Infrastructure as Code (IaC) tools such as Terraform or CloudFormation.
  • Incident Management:
  • Experience with incident management, post-incident analysis, and implementing improvements based on lessons learned.
  • Security & Compliance:
  • Good understanding of security best practices and compliance standards in cloud environments.
  • Communication:
  • Exceptional communication skills and the ability to collaborate effectively with cross-functional teams.
  • On-call Rotations:
  • Willingness to participate in on-call rotations and provide off-hours support when necessary.

Preferred Qualifications:

  • Relevant certifications such as:
  • AWS Certified DevOps Engineer
  • AWS Certified SRE
  • Kubernetes certifications
  • Experience with other cloud platforms (e.g., Azure, Google Cloud Platform).
  • Familiarity with microservices architecture and service mesh technologies.
  • Prior experience with application performance tuning and optimization.

About Company

We are a global niche consultancy specialising in services that span all facets of Digital transformation, DevOps, and System Integration. Our mission is to develop innovative, enterprise-wide Digital Transformation services, in response to business and technical challenges faced by the biggest and most complex organisations. We have a reputation for tackling the most complex issues and can optimise IT delivery across the enterprise. Our core services include: -Digital Transformation -DevOps -Mainframe DevOps -DataOps -System Integration -Deep Banking & Telco Domain expertise We are a global business with offices in The UK, and India. We have an excellent reputation for our work in the Financial and Telco. sectors and more generally for our work in mainframe environments. We partner with some of the strongest organisations in our industry whilst our clients are internationally recognised brand names. We promote equality, diversity and integrity and are committed to the pursuit of excellence in everything we do. We are hiring the best talent whilst moving to an agile, self-managed organisation model, empowering our staff and encouraging them to work in a highly collaborative way. We have a strong social conscience and actively support and finance humanitarian causes.

Job ID: 121637573