Search by job, company or skills

Site Reliability Engineer (SRE)

Early Applicant
  • Posted 6 days ago
  • Be among the first 40 applicants

Job Description

About The Role -

Techdome is looking for a Site Reliability Engineer (SRE) to build, operate, and continuously improve highly available, secure, and scalable cloud infrastructure across our Healthcare, FinTech, AI, and SaaS products.

This isn't a typical DevOps role. You'll own production environments end-to-end — improving system reliability, automating operations, building resilient deployment pipelines, managing incidents, and shipping zero-downtime releases using strategies like Blue-Green and Rolling Deployments.

If you're passionate about automation, cloud infrastructure, Kubernetes, observability, and AI-powered operations, we'd love to talk to you.

Required Skills -

  • 3+ years of experience as an SRE, DevOps Engineer, Platform Engineer, or Cloud Engineer.
  • Hands-on experience with AWS, Azure, or GCP.
  • Strong experience with Docker and Kubernetes.
  • Expertise in Terraform, Ansible, or other Infrastructure as Code tools.
  • Experience building CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar.
  • Strong Linux administration, networking, and distributed systems knowledge.
  • Programming or scripting experience in Python, Go, or Bash.
  • Experience managing large-scale production environments.
  • Understanding of deployment strategies including Blue-Green, Canary, and Rolling Deployments.
  • Experience in FinTech, Payments, Healthcare, or other high-availability environments.
  • Working experience with AI tools such as GitHub Copilot, Claude, Cursor, ChatGPT, or similar developer productivity tools.

What You'll Do -

  • Maintain the availability, reliability, scalability, and performance of production systems.
  • Manage and optimize production environments across cloud platforms.
  • Design and automate deployment pipelines using CI/CD best practices.
  • Implement Blue-Green, Rolling, and Zero-Downtime deployment strategies.
  • Build Infrastructure as Code using Terraform and Ansible.
  • Implement observability using Prometheus, Grafana, ELK, Datadog, OpenTelemetry, and centralized logging.
  • Define and maintain SLIs, SLOs, and Error Budgets.
  • Lead production incident management, Root Cause Analysis (RCA), and post-incident reviews.
  • Perform cloud cost optimization and capacity planning.
  • Automate operational workflows using scripting and AI-powered tooling.
  • Participate in on-call rotations and production support.

Good to Have -

  • Experience building AI-powered operational workflows for monitoring, alert triage, incident summarization, or automation.
  • Knowledge of SRE principles including SLOs, SLIs, Error Budgets, Chaos Engineering, and Reliability Engineering.

Why Join Techdome -

  • Work on real-world AI, Healthcare, Payments, and SaaS products — not just one narrow domain.
  • Own critical production infrastructure from Day 1.
  • Build systems that support thousands of users and business-critical workflows.
  • Collaborate directly with founders and senior engineering leadership.
  • Experience a genuine AI-first engineering culture with modern tooling and automation.
  • Enjoy fast decision-making, real ownership, and accelerated career growth.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151361469

Similar Jobs

Hyderabad, India

Skills:

AppdynamicsGitDockerLinuxAnsiblePrometheusSplunkGrafanaKubernetesAWS

India, Remote

Skills:

GcpDatadogPrometheusAzureTerraformGrafanaJenkinsAnsibleGitHub ActionsAI-OpsGCP Operations SuiteAzure Monitor

Hyderabad, Bengaluru, Pune

Skills:

AutomationKubernetesAWSCI/CD Pipelines

Hyderabad, Pune

Skills:

UnixC++PerlData StructuresRubyPythonPerformance Tuning