

Search by job, company or skills

About Ema
Ema is building the world's leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs.
We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver and Bangalore, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale.
About the Role
As a Site Reliability Engineer at Ema, you will own the stability, availability, and operational health of our agentic AI platform across customer environments. You'll work closely with Engineering and DevOps to provision infrastructure, drive deployment excellence, and keep production running at the quality bar our enterprise customers expect — 99.9%+ uptime, proactive incident response, and continuous improvement.
What You'll Do
Infrastructure & Deployment
Production Stability & Observability
Documentation & Knowledge Management
What We're Looking For
Nice to Have
Compensation
Compensation offered will be determined by factors such as location, level, job-related knowledge, skills, and experience. Certain roles may be eligible for variable compensation, equity, and benefits.
Ema Unlimited is an equal opportunity employer committed to providing equal employment opportunities to all employees and applicants without regard to race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or genetics.
Job ID: 150453151
Skills:
Hadoop, Prometheus, Grafana, Datadog, Apache Airflow, Cloudwatch, Linux, Terraform, Spark, Splunk, Python, Kubernetes, AWS, AWS EMR, FinOps, Amazon EKS, OpenSearch
Skills:
Unix, Elk, Prometheus, Bash, Grafana, Datadog, Gcp, Linux, Docker, Terraform, Ansible, Helm, Azure, Kubernetes, Python, AWS, GitOps, GitLab CI, GitHub Actions, OpenTelemetry, oci, ArgoCD
Skills:
Docker, Terraform, Prometheus, Bash, Splunk, Grafana, Python, Kubernetes, Aws S3, Go
Skills:
Proxies, Routing And Switching, Sda, Load Balancers, Shell, Ansible, Firewalls, Vmware Nsx, Python, Cisco ACI Fabrics, AI-assisted reliability workflows, SD-WAN
Skills:
Terraform, Kubernetes, Python, AWS, GenAI, GitOps, Go, Ai, AIOps, Observability