

Search by job, company or skills

About organisation:-
We are a US Based Venture backed Digital Health Company. We enable Health Care Providers (HCP) to capture true Virtual Care Opportunities beyond Telehealth. We enable HCP to provide Proactive and Continuous Care and add new Recurring monthly revenue streams without any upfront cost. With our unique distribution and business model, we are seeing fast acceptance and great adaptation with our target customers.
We have built unique and Industry's first Integrated Hardware, Cloud & AI Technologies based Virtual care Platforms for HCP Market. We are a US-focused post revenue company with customers in 9 US States and growing fast. We provide an excellent opportunity to Innovate and work on cutting-edge product technologies in a very fast-moving dynamic and empowered environment.
Job brief
We're looking for a Senior Site Reliability Engineer to own the reliability, performance, and operational maturity of the platform. You'll sit at the intersection of engineering and operations — writing code, hardening infrastructure, and building the systems that let a small team run a HIPAA-regulated healthcare platform with confidence. This is a hands-on senior role: you'll set direction on SLOs, observability, incident response, and infrastructure-as-code while still shipping the work yourself.
Responsibilities
What We're Looking For
Nice to Have
• Exposure to new-age, open-source AI-native observability and AI-SRE tooling — e.g., SigNoz, HolmesGPT, IncidentFox, K8sGPT, Keep, or Coroot — and interest in using AI agents to accelerate root-cause analysis and reduce MTTR. Familiarity with commercial equivalents (e.g., Datadog Bits AI) is a plus, but we favor open-source, self-hostable options.
• Experience running systems under HIPAA, SOC 2, HITRUST, or similar regulatory regimes.
• Familiarity with GraphQL infrastructure (e.g., Hasura) and headless CMS platforms (e.g., Strapi).
• Background in healthcare, medical device, or other high-stakes/regulated environments.
• Experience with identity and access tooling (SSO/SAML/SCIM, Entra ID) and data/BI tools such as Metabase.
• Experience defining an on-call culture from scratch or maturing an early-stage one.
Job ID: 153429961
Skills:
Hadoop, Prometheus, Grafana, Datadog, Apache Airflow, Cloudwatch, Linux, Terraform, Spark, Splunk, Python, Kubernetes, AWS, AWS EMR, FinOps, Amazon EKS, OpenSearch
Skills:
Docker, Terraform, Prometheus, Bash, Splunk, Grafana, Python, Kubernetes, Aws S3, Go
Skills:
Proxies, Routing And Switching, Sda, Load Balancers, Shell, Ansible, Firewalls, Vmware Nsx, Python, Cisco ACI Fabrics, AI-assisted reliability workflows, SD-WAN
Skills:
Terraform, Kubernetes, Python, AWS, GenAI, GitOps, Go, Ai, AIOps, Observability
Skills:
Java, Node, Prometheus, Js, Pulumi, Grafana, C Sharp, Terraform, Typescript, Azure, Kubernetes, Go, OpenTelemetry