If it's down, it's on you. If it stays up, that's on you too.
This isn't a ticket-queue DevOps gig where you spin up infra and wait for the next request. At Techdome, you own production — for real Healthcare, FinTech, AI, and SaaS products, with real users depending on uptime that isn't negotiable.
You'll build the pipelines, own the incidents, ship the zero-downtime releases, and be the person the team trusts when something breaks at 2am. If automation, Kubernetes, observability, and AI-powered ops get you genuinely excited — not just as buzzwords on a resume — keep reading.
What You'll Actually Do
- Keep production up, fast, and stable — availability, reliability, scalability, and performance across every environment you touch.
- Run the cloud like it's yours — manage and optimize environments across AWS, Azure, or GCP.
- Build deployment pipelines that don't break things — CI/CD engineered for zero-downtime, with Blue-Green, Rolling, and Canary strategies as your default toolkit, not a slide in a deck.
- Codify the infrastructure — Terraform and Ansible so nothing important lives only in someone's head.
- See everything before it becomes a problem — observability with Prometheus, Grafana, ELK, Datadog, OpenTelemetry, and centralized logging.
- Own the numbers that matter — define and maintain SLIs, SLOs, and Error Budgets, not just watch dashboards.
- Lead when things go wrong — incident management, RCA, and post-incident reviews that actually prevent the repeat.
- Watch the bill, not just the uptime — cloud cost optimization and capacity planning.
- Automate the boring and the risky — scripting and AI-powered tooling for alert triage, incident summarization, and operational workflows.
- Carry the pager — participate in on-call rotations and production support, because ownership doesn't stop at 6pm.
What You Bring
- 3+ years as an SRE, DevOps Engineer, Platform Engineer, or Cloud Engineer.
- Real hands-on cloud experience — AWS, Azure, or GCP.
- Docker and Kubernetes fluency — not just I've used it once.
- Infrastructure as Code expertise — Terraform, Ansible, or equivalent.
- CI/CD pipelines built from scratch — Jenkins, GitHub Actions, GitLab CI, or similar.
- Solid Linux, networking, and distributed systems fundamentals — the stuff that doesn't show up in a tutorial.
- Scripting chops — Python, Go, or Bash.
- Experience running large-scale production environments where it's fine isn't good enough.
- Working knowledge of Blue-Green, Canary, and Rolling deployment strategies — in practice, not just in theory.
- Time in FinTech, Payments, Healthcare, or other high-availability environments where downtime has real consequences.
- A working AI toolkit — GitHub Copilot, Claude, Cursor, ChatGPT, or similar, used to actually get things done faster.
Extra Credit
- You've built AI-powered operational workflows — monitoring, alert triage, incident summarization, automation — not just used AI to write scripts faster.
- You speak fluent SRE — SLOs, SLIs, Error Budgets, Chaos Engineering, and Reliability Engineering aren't just terms you've heard.
Why Techdome
- No narrow lane. Work across real-world AI, Healthcare, Payments, and SaaS products — not one domain for years on end.
- Ownership from day one. Critical production infrastructure is yours, not something you inherit after two years of proving yourself.
- Scale that matters. Systems supporting thousands of users and business-critical workflows, not internal tools nobody depends on.
- Direct access to the top. Work with founders and senior engineering leadership — no six layers of management between your idea and a decision.
- An AI-first culture that's actually real. Modern tooling, real automation, not a mandate to use AI more with no support behind it.
- Speed and real ownership. Fast decisions, real stakes, and career growth that isn't stuck waiting for a title change.
If you want infrastructure work with actual teeth — let's talk.