Search by job, company or skills

Senior DevOps Engineer

Early Applicant
  • Posted 7 days ago
  • Be among the first 10 applicants

Job Description

You'll own our cloud infrastructure end-to-end and define the best practices the rest of engineering builds on, whether that's on AWS or standing up our own infrastructure where it makes sense. You're the person who decides what we run and where based on data, efficiency, and cost, and who gives us deep visibility into how the platform behaves through first-class observability, APM, tracing, and logging. As a senior owner of production, you'll also be part of our on-call rotation and incident response.

Responsibilities

  • Own AWS infrastructure architecture, provisioning (IaC), networking, security, scaling, and day-to-day reliability of a HIPAA-compliant platform.
  • Define infrastructure best practices and architecture standards, on cloud or self-hosted, and make build-vs-buy and where-to-run calls grounded in data, efficiency, and cost.
  • Drive cost optimization, right-sizing, usage visibility, and continuously tuning spend against performance.
  • Build first-class observability: APM, distributed tracing, metrics, custom logging, dashboards, SLOs, and actionable alerting.
  • Own messaging and streaming infrastructure, Kafka, AWS SQS, and the coordination systems behind them (e. g., Zookeeper) for reliable, scalable async and event-driven workloads.
  • Own the CI/CD pipeline and deployment automation safe, repeatable, fast releases across environments.
  • Harden security and compliance posture (PHI/HIPAA) secrets, access, encryption, audit, and incident readiness.
  • Participate in a rotational on-call schedule and provide production support, respond to incidents and alerts within SLA, including outside business hours when on rotation.
  • Lead incident response and post-mortems; drive blameless RCAs, track action items to closure, and reduce toil and MTTR through automation.
  • Set the standard for infra quality and mentor engineers on operability, reliability, cost-awareness, and healthy on-call practices.

Requirements

  • 7+ years in DevOps / SRE / infrastructure, with deep, proven hands-on AWS expertise running production at scale.
  • Strong infrastructure-as-code and automation background (e. g., Terraform / CloudFormation, config management, scripting).
  • Proven track record in defining infra architecture and best practices, and making decisions based on data, efficiency, and cost, not just standing up what's asked.
  • Deep observability expertise in APM, distributed tracing, monitoring, and custom logging with hands-on experience across Datadog, Grafana, and Prometheus.
  • Strong containerization and orchestration experience with Docker and Kubernetes (plus ECS), including running stateful workloads.
  • Hands-on with messaging/streaming infrastructure Kafka and/or AWS SQS (and related coordination tech such as Zookeeper).
  • Proven production-support experience comfortable owning uptime, participating in a rotational on-call schedule, and leading incident response under pressure.
  • Strong grasp of cloud security and, ideally, compliance-driven environments (HIPAA/PHI or similar).
  • A genuine ownership mindset goes the extra mile on complex, ambiguous infra problems rather than working around them.
  • Healthcare / health-tech experience, especially anything touching PHI and HIPAA compliance.
  • Experience setting up and operating self-hosted / on-prem infrastructure alongside cloud.
  • FinOps / cost-optimization tooling and practices at scale.
  • Experience building a golden-path platform tooling that other engineers adopt.
  • Experience establishing on-call / incident-management practices and tooling like PagerDuty, Incident.io, Grafana OnCall, and OpsGenie.

This job was posted by Muskaan Maini from Sciometrix.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151954341

Similar Jobs

Gurugram, Gurugram, India

Skills:

GithubCactiPrometheusKafkaApache TomcatGroovySvnHttp ServerShellDockerTerraformOpenshiftGitlabPythonAWSJavaRed HatPowerShellGangliaUbuntuJenkinsBitbucketAnsibleCentosNagiosPuppetAzureKubernetesChefMLflowKubeflowSageMakerHypericPulsarmunin

Gurugram, Gurugram, India

Skills:

BashDnsLinuxTerraformDockerIamLoad BalancingAzurePythonAWSVPCsGitHub ActionsrbacABACOIDC

Delhi, India

Skills:

Route 53VpcIso 27001Api GatewayNlbRDSAWSAzureECSIamTerraformPrivateLinkGitHub ActionsTACOSSOC 2IAM Identity CenterTransit GatewayAWS OrganizationsSCPs

Noida, India

Skills:

KafkaApache SparkElkGrafanaPrometheusRedisMySQLCassandraPythonKubernetesBashDockerApache KafkaJenkinsPromtailLokiGoHDFSCI CDakamai

Gurugram, Gurugram, India

Skills:

JenkinsTerraformBashArtifactoryPythonAWSHCLGitHub Enterprise

Beware of Scammers

We don’t charge money for job offers