Search by job, company or skills

Senior Software Engineer - Site Reliability

Early Applicant
  • Posted 8 hours ago
  • Be among the first 10 applicants

Job Description

Company Description

Synthlane is a leading IT services and consulting firm that helps businesses, large enterprises, and government institutions solve complex technology challenges with strategic, scalable solutions. The company specializes in enterprise IT consulting, digital transformation, cybersecurity and risk management, cloud and infrastructure services, and custom software development. Synthlane focuses on aligning technology initiatives with business objectives, enabling clients to modernize operations, strengthen security, and improve productivity. Its experienced team of consultants, engineers, and strategists works closely with clients to design future-ready solutions that deliver measurable results. Organizations seeking a trusted partner for reliable, innovation-driven IT services rely on Synthlane to support long-term success.

Role Description

We are seeking an accomplished Senior Site Reliability Engineer (SRE) to lead the

design, implementation, and evolution of highly available, scalable, and resilient systems

across our multi-cloud infrastructure. In this senior role, you will drive architectural

decisions, establish reliability standards, and mentor teams while ensuring operational

excellence across complex distributed systems. You will partner with engineering

leadership, development teams, and product stakeholders to shape infrastructure

strategy, implement sophisticated automation, and champion a culture of reliability

engineering.

As a Senior SRE, you'll tackle sophisticated, large-scale challenges using cutting-edge

technologies across AWS and Azure platforms. You will lead critical initiatives that

impact system reliability at scale, architect solutions for complex infrastructure problems,

and guide teams in adopting industry-leading practices that drive meaningful

improvements across our entire technology ecosystem.

Requirements

● Architect and implement highly reliable, scalable, and cost-effective infrastructure

solutions for mission-critical applications across multi-cloud environments (AWS and

Azure).

● Lead the definition and refinement of service level objectives (SLOs), service level

indicators (SLIs), and error budgets, establishing reliability standards across the

organization.

● Design and implement sophisticated Infrastructure as Code (IaC) solutions using

Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.

● Drive automation strategies to eliminate toil, improve operational efficiency, and enable

self-service capabilities for development teams.

● Lead incident response efforts, conduct thorough post-incident reviews, and implement

systemic improvements to prevent recurrence.

● Champion cloud-native architectures and modern reliability practices, serving as a

technical advisor for infrastructure and platform decisions.

● Participate in and help optimize the on-call rotation, ensuring sustainable practices and

effective escalation procedures.

● Establish and maintain comprehensive documentation standards, runbooks, and

knowledge repositories that enable team autonomy and effective incident response.

● Design and implement advanced monitoring, logging, and alerting strategies using

observability platforms to enable proactive issue detection and resolution.

● Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement

sophisticated deployment strategies including blue-green, canary, and progressive

delivery patterns.

● Ensure security, compliance, and governance standards are embedded throughout the

infrastructure lifecycle, implementing security-as-code practices.

● Drive capacity planning, performance optimization, and cost management initiatives

across cloud platforms.

● Architect and implement highly reliable, scalable, and cost-effective infrastructure

solutions for mission-critical applications across multi-cloud environments (AWS and

Azure).

● Lead the definition and refinement of service level objectives (SLOs), service level

indicators (SLIs), and error budgets, establishing reliability standards across the

organization.

● Design and implement sophisticated Infrastructure as Code (IaC) solutions using

Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.

● Drive automation strategies to eliminate toil, improve operational efficiency, and enable

self-service capabilities for development teams.

● Lead incident response efforts, conduct thorough post-incident reviews, and implement

systemic improvements to prevent recurrence.

● Champion cloud-native architectures and modern reliability practices, serving as a

technical advisor for infrastructure and platform decisions.

● Participate in and help optimize the on-call rotation, ensuring sustainable practices and

effective escalation procedures.

● Establish and maintain comprehensive documentation standards, runbooks, and

knowledge repositories that enable team autonomy and effective incident response.

● Design and implement advanced monitoring, logging, and alerting strategies using

observability platforms to enable proactive issue detection and resolution.

● Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement

sophisticated deployment strategies including blue-green, canary, and progressive

delivery patterns.

● Ensure security, compliance, and governance standards are embedded throughout the

infrastructure lifecycle, implementing security-as-code practices.

● Drive capacity planning, performance optimization, and cost management initiatives

across cloud platforms.

● Collaborate with architecture and security teams to establish platform standards,

reference architectures, and best practices.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152473873

Beware of Scammers

We don’t charge money for job offers