Search by job, company or skills

Senior Devops Engineer/lead

7-9 Years
18.5 - 28.5 LPA
Early Applicant
Quick Apply
  • Posted 10 hours ago
  • Be among the first 50 applicants

Job Description

ABOUT THE ROLEWe are looking for a Senior DevOps Engineer to join our Infrastructure team, focused on Kubernetes operations andautomation at scale. You will drive the strategy and own the health, scalability, and reliability of our Kubernetes clusters,while architecting and leading the development of tooling that platform and service teams rely on to manage infrastructuresafely and efficiently. You will work closely with service owners, platform engineers, and tooling teams — and mentor juniorengineers — to keep clusters running smoothly and to automate away manual, error-prone operational work.

WHAT YOU'LL DOOwn and drive the roadmap for Kubernetes cluster operations — lead cluster upgrades, manage node groups (scaling,draining, replacement), and maintain overall cluster health across environments at scaleLead troubleshooting of complex production and non-production issues — use kubectl, logs, metrics, and other diagnosticsto identify root cause and resolve workload, node, and networking failures, often across multiple interdependent systemsArchitect and lead development of internal tooling — design, implement, and maintain automation (scripts, CLIs,controllers, operators) that helps teams manage infrastructure more reliably and with less manual effortDrive automation strategy for operational workflows — replace manual runbooks with scripts and tools that handleupgrades, remediation, scaling, and routine maintenanceDiagnose and resolve complex infrastructure blockers — debug deployment failures, node/pod scheduling issues,resource constraints, and misconfigurations across clustersDefine and improve observability strategy — instrument logging, metrics, and alerting to increase visibility into clusterhealth and reduce time-to-detect/resolveSet standards for issue tracking and triage — author detailed bug reports capturing root cause, repro steps, and impact;drive issues to resolution and improve team-wide reporting practicesPartner cross-functionally with service owners, platform teams, and leadership — align on operational requirements,capacity planning, and upgrade/maintenance schedulesLead on-call rotations and incident response, own post-incident reviews, and drive continuous improvement initiativesacross the infrastructure orgMentor and upskill junior and mid-level engineers, providing technical guidance and code/design reviews

WHAT WE'RE LOOKING FORRequired7-9 years of DevOps, SRE, or Infrastructure Engineering experienceDeep, hands-on Kubernetes operations expertise — cluster upgrades, node group management, and complextroubleshooting across large-scale, multi-cluster environmentsStrong programming/scripting skills (Python, Go, Bash, or similar) — proven track record building, scaling, and maintainingautomation and internal tooling used by multiple teamsAdvanced skills in reading and interpreting logs, metrics, and system state to diagnose complex, cross-systeminfrastructure issuesExtensive experience with cloud infrastructure (AWS, GCP, or Azure), including architecture decisions andcost/performance tradeoffsExpert-level debugging skills across distributed systems — able to trace failures from symptom through to root cause inhighly complex environmentsExcellent written and verbal communication — able to produce clear status updates, bug reports, technical documentation,and influence technical direction across teamsDemonstrated experience mentoring engineers and leading technical initiativesPreferredDeep experience with infrastructure-as-code tools (Terraform, Helm, Ansible), including designing reusable modules/patternsfor org-wide useStrong familiarity with CI/CD systems (Jenkins, GitHub Actions, Spinnaker, or similar), including pipeline architectureHands-on experience with GitOps workflows (ArgoCD, Flux) at scaleAdvanced experience with observability stacks (Prometheus, Grafana, Datadog), including designing alerting/SLO frameworksProven experience operating Kubernetes at scale across multiple clusters, regions, or environments, includingcapacity planning and disaster recoveryExperience contributing to or leading architectural decisions for infrastructure platforms

More Info

Function:
Employment Type:

About Company

Artech is the largest Women & Minority owned IT staffing firm in the US, with US$ 800 million annual revenue run rate in 2021 and a footprint across the globe. With nearly three decades of experience, Artech empowers businesses through applied human intelligence and offers a spectrum of services that include Workforce Solutions (Contingent Staffing, Bulk/ Project Staffing, Master Vendor, RPO, Direct Hire and Payroll Transition) and Project-Based Solutions (Digital Experience, Technical Operations, Technical Development, Business Operations & Digital Platforms). Artech works with over 90 Fortune 500 clients across USA, Canada, India, and China.
At Artech, we are empowering talent by connecting potential with opportunities through applied human intelligence. We empower our teams to maximize the impact of their intellect, through a performance oriented, diverse, flexible, and inclusive work environment supported by our continuous learning and development focus.

Job ID: 153744835

Beware of Scammers

We don’t charge money for job offers