Search Jobs

Search by job, company or skills

QA / Automation Engineer - Reliability & Validation Systems

QA / Automation Engineer - Reliability & Validation Systems

ciroos
Early Applicant
  • Posted 12 days ago
  • Be among the first 30 applicants

Job Description

We are looking for a QA/Automation Engineer who operates at the intersection of reliability engineering and system-level QA. You will define how reliability is validated, not just monitored. This includes building automation systems that continuously test, break, and verify complex distributed systems in production-like environments.

Responsibilities

  • Define and own the strategy for system-level validation.
  • Design and build scalable automation frameworks for API, integration, and end-to-end testing; Kubernetes and system-level validation; and regression and reliability pipelines.
  • Build systems that proactively detect failures before they reach production.
  • Drive chaos engineering and failure injection practices.
  • Establish CI/CD reliability gates with strong validation coverage.
  • Partner with SRE, platform, and backend teams to ensure systems are both observable and testable.
  • Lead incident analysis with a focus on improving validation and preventing recurrence.
  • Mentor engineers and raise the bar for system reliability and quality.
  • Kubernetes-based distributed systems at scale.
  • Observability and alerting pipelines.
  • AI-assisted incident investigation systems.
  • Multi-cloud (AWS, Azure, GCP) environments, networking deployments.
  • Reliability validation, chaos testing, and failure injection systems.
  • Infrastructure and deployment automation pipelines.

Requirements

  • Deep Expertise In Kubernetes internals, debugging, and multi-cluster systems; distributed systems behavior and failure modes; observability stacks and alerting frameworks; production incident handling and root cause analysis.
  • Strong Hands-on Experience With EKS, GKE, or managed Kubernetes platforms; networking concepts: VPC, load balancers, service communication, IAM, Chaos testing and reliability engineering practices, designing large-scale automation and validation systems.

Programming

  • Strong coding skills in Python and Go (mandatory).
  • Experience building automation frameworks and system-level tooling.
  • Proficiency in Shell scripting and infrastructure automation.

What Makes This Role Different

  • You are responsible for ensuring systems are provably reliable, not just operational.
  • Deep QA and validation engineering.
  • Focus on testing distributed systems, not just application features.
  • Work on failure scenarios, not just happy paths.

This job was posted by Umesh Pratap Singh from Ciroos.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

GKE

Alerting frameworks

EKS

GCP)

Failure injection

Observability stacks

Incident handling and root cause analysis

Multi-cloud (AWS

CI/CD reliability gates

System-level tooling

AI-assisted incident investigation

Chaos engineering

About Company