About The Role -
Techdome is looking for a Site Reliability Engineer (SRE) to build, operate, and continuously improve highly available, secure, and scalable cloud infrastructure across our Healthcare, FinTech, AI, and SaaS products.
This isn't a typical DevOps role. You'll own production environments end-to-end — improving system reliability, automating operations, building resilient deployment pipelines, managing incidents, and shipping zero-downtime releases using strategies like Blue-Green and Rolling Deployments.
If you're passionate about automation, cloud infrastructure, Kubernetes, observability, and AI-powered operations, we'd love to talk to you.
Required Skills -
- 3+ years of experience as an SRE, DevOps Engineer, Platform Engineer, or Cloud Engineer.
- Hands-on experience with AWS, Azure, or GCP.
- Strong experience with Docker and Kubernetes.
- Expertise in Terraform, Ansible, or other Infrastructure as Code tools.
- Experience building CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar.
- Strong Linux administration, networking, and distributed systems knowledge.
- Programming or scripting experience in Python, Go, or Bash.
- Experience managing large-scale production environments.
- Understanding of deployment strategies including Blue-Green, Canary, and Rolling Deployments.
- Experience in FinTech, Payments, Healthcare, or other high-availability environments.
- Working experience with AI tools such as GitHub Copilot, Claude, Cursor, ChatGPT, or similar developer productivity tools.
What You'll Do -
- Maintain the availability, reliability, scalability, and performance of production systems.
- Manage and optimize production environments across cloud platforms.
- Design and automate deployment pipelines using CI/CD best practices.
- Implement Blue-Green, Rolling, and Zero-Downtime deployment strategies.
- Build Infrastructure as Code using Terraform and Ansible.
- Implement observability using Prometheus, Grafana, ELK, Datadog, OpenTelemetry, and centralized logging.
- Define and maintain SLIs, SLOs, and Error Budgets.
- Lead production incident management, Root Cause Analysis (RCA), and post-incident reviews.
- Perform cloud cost optimization and capacity planning.
- Automate operational workflows using scripting and AI-powered tooling.
- Participate in on-call rotations and production support.
Good to Have -
- Experience building AI-powered operational workflows for monitoring, alert triage, incident summarization, or automation.
- Knowledge of SRE principles including SLOs, SLIs, Error Budgets, Chaos Engineering, and Reliability Engineering.
Why Join Techdome -
- Work on real-world AI, Healthcare, Payments, and SaaS products — not just one narrow domain.
- Own critical production infrastructure from Day 1.
- Build systems that support thousands of users and business-critical workflows.
- Collaborate directly with founders and senior engineering leadership.
- Experience a genuine AI-first engineering culture with modern tooling and automation.
- Enjoy fast decision-making, real ownership, and accelerated career growth.