Search by job, company or skills

Senior DevOps Engineer (9+ yrs)

Senior DevOps Engineer (9+ yrs)

mindbrain
9-11 Years
Not Disclosed
Early Applicant
  • Posted 13 hours ago
  • Be among the first 10 applicants

Job Description

Senior DevOps Engineer – Production Stability

Experience: 9+ Years

Employment Type: Contract-to-Hire (C2H)

Location: Remote

Job Overview

We are looking for a highly experienced Senior DevOps Engineer with a strong production-operations mindset to support and stabilize business-critical 24x7 systems. The role involves working across hybrid cloud and on-premises environments, legacy and modern technologies, CI/CD pipelines, Kubernetes, networking, API gateways, and observability platforms.

This is a highly hands-on role (90% individual contributor) requiring strong troubleshooting and incident-response capabilities. The selected candidate will work as part of a 3–4 member contractor engineering pod and participate in an on-call rotation, primarily on weekdays.

Key Responsibilities

Production Operations & Incident Management

  • Support and maintain 24x7 production systems and integrations.
  • Participate in the on-call rotation and respond to production incidents.
  • Troubleshoot issues across CI/CD, Kubernetes, applications, networking, and API gateways.
  • Perform incident triage, mitigation, recovery, and root-cause analysis.
  • Ensure reliable deployments, rollback procedures, and production stability.

CI/CD & Deployment

  • Manage and troubleshoot CI/CD pipelines using GitHub Actions, Azure DevOps, Octopus, or similar tools.
  • Diagnose and resolve pipeline and deployment failures.
  • Support deployments across DEV, QA, PROD, and special environments.
  • Improve deployment reliability, automation, and consistency.

Kubernetes & Infrastructure

  • Operate and troubleshoot Kubernetes clusters across cloud and on-premises environments.
  • Perform cluster troubleshooting, patching, and basic upgrades.
  • Monitor and optimize CPU, memory, and other infrastructure resources.
  • Troubleshoot supporting services such as Redis, RabbitMQ, databases, and API gateways.

Cloud, API Gateway & Legacy Systems

  • Support AWS API Gateway and related cloud infrastructure.
  • Work with existing Terraform modules and configurations.
  • Troubleshoot routing, domain, and traffic-related issues involving Cloudflare/APIM.
  • Support critical legacy systems, including Windows-based services.

Monitoring & Observability

  • Use Prometheus, Grafana, logs, and monitoring platforms for production troubleshooting.
  • Analyze system performance and identify potential reliability issues.
  • Improve alerts, dashboards, and monitoring where required.

Documentation & Knowledge Sharing

  • Create and maintain deployment, rollback, recovery, and troubleshooting documentation.
  • Develop and maintain operational runbooks for critical systems.
  • Share technical knowledge with the contractor lead and engineering team.

Mandatory Skills

  • Strong hands-on Linux experience – Mandatory
  • 9+ years of experience in DevOps / SRE / Production Operations
  • Strong production troubleshooting and incident-management experience
  • Hands-on Kubernetes operations
  • Experience with CI/CD tools such as GitHub Actions, Azure DevOps, or Octopus
  • Experience with AWS and/or Azure
  • Working knowledge of Terraform
  • Experience with Prometheus, Grafana, and centralized logging
  • Experience supporting hybrid cloud + on-premises environments
  • Strong scripting skills using Bash, Python, PowerShell, or similar

Preferred Experience

  • AWS API Gateway / Azure API Management
  • Cloudflare and networking
  • Redis and RabbitMQ
  • Database troubleshooting
  • Windows-based legacy systems
  • Production on-call and 24x7 support environments
  • Incident response, RCA, and operational runbooks

What We're Looking For

  • Strong operator/SRE mindset
  • Comfortable working in complex and partially documented environments
  • Ability to quickly understand existing systems and take ownership
  • Calm and effective under production pressure
  • Strong troubleshooting and problem-solving skills
  • Focused on stability, reliability, and business continuity
  • Comfortable with hands-on production ownership and on-call responsibilities

First 90-Day Success Criteria

  • Independently support critical production systems.
  • Stabilize and improve CI/CD pipeline reliability.
  • Handle production incidents efficiently and minimize downtime.
  • Document critical deployment, rollback, and recovery workflows.
  • Establish effective collaboration with the contractor lead and engineering team.

Not a Good Fit If You

  • Prefer only greenfield or highly standardized environments.
  • Focus primarily on architecture rather than hands-on operations.
  • Prefer to avoid production ownership or on-call responsibilities.
  • Require fully documented systems before taking ownership.
  • Are uncomfortable troubleshooting under time-sensitive production conditions.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

About Company