Search by job, company or skills

Site Reliability Engineer

Early Applicant
  • Posted 10 hours ago
  • Be among the first 10 applicants

Job Description

If it's down, it's on you. If it stays up, that's on you too.

At Techdome, you own production for real Healthcare, FinTech, AI, and SaaS products — not a ticket queue. You'll build pipelines, own incidents, ship zero-downtime releases, and be the person the team trusts at 2am.

What You'll Do

  • Keep production available, reliable, and performant across every environment.
  • Manage and optimize cloud environments (AWS, Azure, or GCP).
  • Build zero-downtime CI/CD pipelines using Blue-Green, Rolling, and Canary strategies.
  • Codify infrastructure with Terraform and Ansible.
  • Drive observability with Prometheus, Grafana, ELK, Datadog, and OpenTelemetry.
  • Define and own SLIs, SLOs, and error budgets.
  • Lead incident response, RCA, and post-incident reviews.
  • Optimize cloud cost and plan capacity.
  • Automate operational workflows, including AI-powered alert triage and incident summarization.
  • Join the on-call rotation.

What You Bring

  • 2+ years as an SRE, DevOps Engineer, Platform Engineer, or Cloud Engineer.
  • Hands-on production experience with AWS, Azure, or GCP.
  • Docker and Kubernetes experience under real load.
  • Infrastructure-as-Code expertise (Terraform, Ansible, or equivalent).
  • CI/CD pipelines built from scratch (Jenkins, GitHub Actions, GitLab CI, or similar).
  • Strong Linux, networking, and distributed-systems fundamentals.
  • Scripting skills in Python, Go, or Bash.
  • Experience with Blue-Green, Canary, and Rolling deployments in production.
  • Background in FinTech, Payments, Healthcare, or another high-availability domain.
  • Regular use of AI tools (Copilot, Claude, Cursor, ChatGPT, or similar) to move faster.

Extra credit: You've built AI-powered ops workflows (monitoring, alert triage, incident summarization), and you speak fluent SRE — SLOs, error budgets, chaos engineering.

Why Techdome

Work across AI, Healthcare, Payments, and SaaS products. Own critical infrastructure from day one. Work directly with founders and senior engineering leadership. Real stakes, fast decisions, real growth.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152466227

Similar Jobs

Hyderabad, Chennai, Pune

Skills:

SaasKafkaPythonSite Reliability EngineerIaaCRancher

Hyderabad, India

Skills:

TcpCassandraPrometheusElk StackDnsGrafanaCloud InfrastructureHttpsTerraformLinux Shell ScriptingNetworking ProtocolsPythonJavaHttpJenkinsMongoDBCouchbaseSplunkHelmKubernetesAws S3Infrastructure as CodeAlertmanagerSpinnaker

Hyderabad, India

Skills:

ElkPrometheusBashGrafanaDatadogJenkinsGcpTerraformDockerLinuxAnsibleAzurePythonKubernetesAWSGoGitLab CIGitHub ActionsOpenTelemetry

Hyderabad, India

Skills:

ServicenowIcmPowerShellGrafanaPythonKustoAzure Power BIKQLGenevaPagerDuty

Hyderabad, India

Skills:

GithubPrometheusBashAmazon CloudWatchGrafanaLinux AdministrationTerraformIamKubernetesPythonAWSAmazon EKSSecrets managementOpenSearchArgo CDOpenTelemetry

Beware of Scammers

We don’t charge money for job offers