Search Jobs

Search by job, company or skills

Senior DevOps Engineer

Senior DevOps Engineer

recrew ai
  • Posted an hour ago
  • Be among the first 10 applicants

Job Description

Role: Senior DevOps / Site Reliability Engineer

Function: DevOps / Site Reliability Engineering / Infrastructure

Location: Kolkata or Gurugram, India

Type: Full-time

Industry: AI, Medical SaaS, Dental Technology

About Company

The company builds AI-powered medical SaaS for the U.S. dental market.

Its all-in-one platform unifies analytics, scheduling, communications, payments, and marketing automation. Thousands of dental practices across North America rely on it daily.

The company was founded in 2015 by a practicing dentist, a data strategist, and a PhD technologist. It is mission-driven and execution-heavy, with a team of 88 people.

Engineers get high ownership, cross-functional collaboration, and clear, measurable impact.

Position Overview

This is a builder's role with full production ownership across the company's AWS infrastructure — not a monitoring seat or ticket queue. You will own Terraform modules, build and maintain GitHub Actions pipelines, and drive the reliability of a platform used by thousands of dental practices. The role offers flexibility between a normal and night shift, with the night window providing uninterrupted time for high-impact work: database migrations, pipeline rewrites, and staged cutovers.

Role & Responsibilities

  • Own and extend Terraform modules across AWS — ECS/Fargate or EKS, RDS, VPC networking, IAM, ALB/NLB, S3, Route 53, CloudWatch — managing state safely and eliminating manually created resources via imports and refactors.
  • Build and maintain GitHub Actions pipelines for build, test, containerization, and deployment; migrate legacy pipelines (Jenkins, CircleCI, GitLab CI) onto a single maintainable platform with staged rollouts, health-gated releases, and fast rollback.
  • Own the production change window on your shift: deploys, schema migrations against live systems, cutovers, and maintenance with minimal disruption.
  • Run and improve observability — dashboards, SLOs, and alerting calibrated to real user impact — and debug containerized applications in production across logs, metrics, traces, and resource limits.
  • Operate PostgreSQL in production: read query plans, identify slow queries, manage connections, locks, and replication; operate supporting services including Redis/ElastiCache and message brokers.
  • Write Python and shell automation to remove repetitive operational work; build self-service tooling, runbooks, and guardrails that make the wider engineering team faster.
  • Write clear, complete shift handover notes and maintain infrastructure documentation; participate in on-call rotation and drive blameless post-incident reviews.

Must Have Criteria

  • 6+ years in DevOps, SRE, Platform, or Infrastructure engineering with genuine production ownership.
  • Hands-on AWS experience across ECS/Fargate or EKS, RDS, VPC networking, IAM, ALB/NLB, S3, and CloudWatch.
  • Hands-on Terraform experience: writing modules, managing state, and reviewing plans in production environments.
  • Practical CI/CD experience with GitHub Actions (or GitLab CI, CircleCI, or Jenkins with demonstrated ability to migrate).
  • Docker proficiency with demonstrated ability to debug containerized applications in production.
  • Working PostgreSQL knowledge: query plans, slow query analysis, connections, locks, and replication.
  • Solid Linux fundamentals, shell scripting, and Python for operational automation.

Nice to Have

  • Experience with Celery, RabbitMQ, NATS, or Kafka in production.
  • Secrets management using AWS Secrets Manager, SOPS, or Vault.
  • Exposure to sharded or multi-tenant database architectures.
  • Experience in compliance-sensitive environments (HIPAA or SOC 2).
  • Kubernetes production experience.
  • Prior experience as the sole engineer on shift or in a night/off-hours rotation.
  • Telephony or VoIP infrastructure exposure.

What We Offer

  • Full infrastructure ownership with a protected change window for high-impact work — no ticket shuffling.
  • Shift flexibility: normal or night shift based on candidate preference and business coverage needs.
  • Modern stack: AWS, Terraform, GitHub Actions, Docker, PostgreSQL, and Kafka-compatible messaging.
  • Direct, measurable impact on the reliability of a platform serving thousands of healthcare practices.
  • Autonomy, cross-functional collaboration, and support for continuous learning and growth.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company

Similar Jobs

4-6 yrs
Gurugram, Gurugram, India
Skills:
DockerTerraformCloudformationAnsiblePrometheusElk StackGrafanaKubernetesGoogle CloudAWS
6-10 yrs
Delhi, India
Skills:
ElkCloudformationPrometheusBashGrafanaDatadogNew RelicJenkinsGcpDockerTerraformLinuxAnsibleShell scriptingAzureKubernetesPythonAWSGitHub ActionsociArgoCDGitLab CI CD
10-12 yrs
Noida, India
Skills:
PrometheusGrafanaDatadogJenkinsGcpDockerTerraformAnsibleAzureKubernetesPythonAWSChefPackerGitLab CIGitHub Actions
4-6 yrs
Noida, India
Skills:
containerization BashLinux AdministrationIncident ResponseDockerTerraformKubernetesPythonGCP ArchitectureWorkflow AutomationGPU Compute ManagementMulti-Cloud Enterprise IntegrationsAI ML Product DeploymentEnterprise Security CompliancePipeline EngineeringObservability
6-8 yrs
Gurugram, India
Skills:
PowerShellPrometheusBashGrafanaGitTerraformDockerIamMicrosoft AzurePythonKubernetesAWSAzure DevOpsCI CD PipelinesInfrastructure as CodeAzure Resource ManagerKey VaultGitHub ActionsAzure Security CenterBicepAzure MonitorApplication Insights