Search by job, company or skills

Azure Chaos Studio_5 YEARS

  • Posted 6 hours ago
  • Be among the first 10 applicants

Job Description

We are looking for experienced SRE / Chaos Engineering / DevOps professionals with 5+ years of experience in cloud infrastructure, production reliability, observability, and resilience engineering. The ideal candidate will have hands-on experience with Azure, DevOps/CI-CD practices, monitoring, incident management, and failure/resilience testing.

Experience with Azure Chaos Studio is highly preferred. Candidates with strong SRE, Chaos Engineering, resilience engineering, or fault-injection experience and good Azure exposure will also be considered.

Key Responsibilities

  • Design and execute resilience, failure, and chaos engineering experiments to validate system reliability.
  • Work with Azure Chaos Studio to perform controlled fault-injection and resilience testing.
  • Identify system weaknesses related to availability, failover, recovery, and fault tolerance.
  • Develop and execute chaos experiments across cloud and containerized environments.
  • Collaborate with development, DevOps, infrastructure, and SRE teams to improve system reliability.
  • Monitor production environments and analyze system performance, availability, and reliability.
  • Support incident management, troubleshooting, root-cause analysis, and production recovery.
  • Implement and maintain CI/CD pipelines using GitHub Actions, Jenkins, or similar tools.
  • Automate infrastructure and deployment activities using Terraform and Ansible.
  • Work with Docker and Kubernetes/container orchestration platforms.
  • Use observability and monitoring tools such as Grafana, Prometheus, CloudWatch, Nagios, Azure Monitor, or similar platforms.
  • Participate in disaster recovery, failover, recovery, and business continuity testing.
  • Define and improve operational resilience, reliability, and recovery practices.
  • Document chaos experiments, findings, remediation actions, and reliability improvements.

Required Skills

  • 5+ years of experience in SRE, DevOps, Cloud Engineering, Reliability Engineering, or Production Engineering.
  • Strong experience with Azure cloud environments.
  • Knowledge or hands-on experience with Azure Chaos Studio is highly preferred.
  • Strong understanding of Chaos Engineering and resilience engineering concepts.
  • Experience with failover, recovery, fault tolerance, disaster recovery, and failure testing.
  • Hands-on experience with CI/CD tools such as GitHub Actions or Jenkins.
  • Experience with Terraform and/or Ansible.
  • Experience with Docker and Kubernetes.
  • Strong knowledge of monitoring and observability concepts.
  • Experience with tools such as Prometheus, Grafana, Azure Monitor, CloudWatch, Nagios, or equivalent.
  • Experience in production support and incident management.
  • Strong troubleshooting and problem-solving skills.

Good to Have

  • Azure Chaos Studio
  • Chaos Mesh / LitmusChaos / Gremlin / AWS Fault Injection Simulator
  • Azure Kubernetes Service (AKS)
  • Azure Monitor / Application Insights
  • SLO, SLA, SLI and error-budget concepts
  • Disaster Recovery and Business Continuity
  • GameDay / resilience testing
  • Python, PowerShell, or Bash scripting
  • Infrastructure as Code and automated reliability testing

Ideal Candidate Profile

Candidates From The Following Backgrounds Are Encouraged To Apply

  • Site Reliability Engineer (SRE)
  • Chaos Engineer
  • Cloud Reliability Engineer
  • Reliability Engineer
  • DevOps Engineer – SRE
  • Platform Engineer
  • Azure DevOps Engineer
  • Cloud Operations Engineer
  • Production Engineer

Skills: azure,sre,chaos

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 152619827

Beware of Scammers

We don’t charge money for job offers