Search by job, company or skills

Senior DevOps Engineer/Lead

7-9 Years
  • Posted 4 hours ago
  • Be among the first 10 applicants

Job Description

Remote || Shift timing - 7:30 PM - 4:30 AM

We are looking for a Senior DevOps Engineer to join our Infrastructure team, focused on Kubernetes operations and automation at scale.

You will drive the strategy and own the health, scalability, and reliability of our Kubernetes clusters,while architecting and leading the development of tooling that platform and service teams rely on to manage infrastructure safely and efficiently. You will work closely with service owners, platform engineers, and tooling teams —

and mentor junior engineers

— to keep clusters running smoothly and to automate away manual, error-prone operational work.

WHAT YOU'LL DO

Own and drive the roadmap

for Kubernetes cluster operations — lead cluster upgrades, manage node groups (scaling,draining, replacement), and maintain overall cluster health across environments at scale

Lead troubleshooting

of complex production and non-production issues — use kubectl, logs, metrics, and other diagnosticsto identify root cause and resolve workload, node, and networking failures, often across multiple interdependent systems

Architect and lead development of internal tooling

— design, implement, and maintain automation (scripts, CLIs,controllers, operators) that helps teams manage infrastructure more reliably and with less manual effort

Drive automation strategy

for operational workflows — replace manual runbooks with scripts and tools that handle upgrades, remediation, scaling, and routine maintenance

Diagnose and resolve complex infrastructure blockers

— debug deployment failures, node/pod scheduling issues, resource constraints, and misconfigurations across clusters

Define and improve observability strategy

— instrument logging, metrics, and alerting to increase visibility into cluster health and reduce time-to-detect/resolve

Set standards for issue tracking and triage

— author detailed bug reports capturing root cause, repro steps, and impact; drive issues to resolution and improve team-wide reporting practices

Partner cross-functionally with service owners, platform teams, and leadership

— align on operational requirements, capacity planning, and upgrade/maintenance schedules

Lead on-call rotations and incident response

, own post-incident reviews, and drive continuous improvement initiatives across the infrastructure org

Mentor and upskill junior and mid-level engineers

, providing technical guidance and code/design reviews

WHAT WE'RE LOOKING FOR

Required

7-9 years

of DevOps, SRE, or Infrastructure Engineering experience

Deep, hands-on Kubernetes operations expertise

— cluster upgrades, node group management, and complex troubleshooting across large-scale, multi-cluster environments

Strong programming/scripting skills

(Python, Go, Bash, or similar) — proven track record building, scaling, and maintaining automation and internal tooling used by multiple teams

Advanced skills in reading and interpreting logs, metrics, and system state

to diagnose complex, cross-system infrastructure issues

Extensive experience with cloud infrastructure

(AWS, GCP, or Azure), including architecture decisions and cost/performance tradeoffs

Expert-level debugging skills across distributed systems

— able to trace failures from symptom through to root cause in highly complex environments

Excellent written and verbal communication

— able to produce clear status updates, bug reports, technical documentation, and influence technical direction across teams

Demonstrated experience mentoring engineers

and leading technical initiatives

Preferred

Deep experience with infrastructure-as-code tools (Terraform, Helm, Ansible), including designing reusable modules/patterns for org-wide use

Strong familiarity with CI/CD systems (Jenkins, GitHub Actions, Spinnaker, or similar), including pipeline architecture

Hands-on experience with GitOps workflows (ArgoCD, Flux) at scale

Advanced experience with observability stacks (Prometheus, Grafana, Datadog), including designing alerting/SLO frameworks

Proven experience operating Kubernetes at scale across multiple clusters, regions, or environments

, including capacity planning and disaster recovery

Experience contributing to or leading architectural decisions for infrastructure platforms

TECHNOLOGIES YOU'LL WORK WITH

Kubernetes · kubectl · Python · Go · Bash · Terraform · Helm · Prometheus/Grafana · GitHub · CI/CD tooling · Cloud infrastructure (AWS/GCP/Azure)

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 153790219

Similar Jobs

Coimbatore, India

Skills:

GithubPowerShellPrometheusBashGrafanaGitTerraformMicrosoft AzurePythonKubernetesInfrastructure as CodeAzure Key VaultAzure DevOps PipelinesrbacAzure PolicyAzure Monitor

Navi Mumbai, Mumbai

Skills:

Elastic Compute (EC2)Storages (S3EfsECSDevops EngineerAWS Cloud formationEBS

Bengaluru, India

Skills:

GitLinux Shell ScriptingTerraformDockerDatadogKubernetesPythonAWSKustomizeCI CD

Bengaluru, India

Skills:

DatadogPrometheusSplunkJavaPythonElkGrafanaShellAnsibleAppdynamics

Beware of Scammers

We don’t charge money for job offers