

Search by job, company or skills

What You'll Do:
-- Be the Force Behind Observability & Stability
-- Own & Automate Operations
-- Lead Incident Response & Operational Excellence
-- Collaborate & Mentor
What Makes You a Great Fit-
- Deep expertise in Logs, Metrics, and Alerting, with a strong focus on Alerting automation.
- Experience in hybrid SaaS environments spanning cloud-native and global infrastructure.
- Strong background in Kubernetes, Infrastructure-as-Code (Terraform), Golang, AWS/GCP, and networking observability.
- Proven track record of eliminating toil and improving operational efficiency through automation.
- Passion for deep observability, networking-scale analytics, and automation at the edge.
If you love solving reliability challenges at global scale, automating everything, and working in a hybrid cloud + networking environment, we want to talk to you!
The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.
Must-Have:
Observability & Alerting Expertise Strong experience with Logs, Metrics, and Alerts, with a focus on high-fidelity alerting and automation. Automation & Infrastructure as Code Deep knowledge of Terraform, ArgoCD, Helm, Kubernetes, and Golang for automation. Cloud & Hybrid SaaS Experience Hands-on experience managing cloud-native (AWS/GCP) and edge infrastructure. Incident Response & Reliability Engineering Strong on-call experience, with a track record of reducing MTTR through automation Kubernetes Mastery Hands-on experience deploying, managing, and troubleshooting Kubernetes in production environments.
Nice-to-Have:
Networking & Edge Observability Familiarity with monitoring routers, switches, and firewalls in a global PoP environment. Data & Analytics in Observability Experience with time-series databases (Prometheus, Grafana, OpenTelemetry, etc.). Security & Compliance Awareness Understanding of secure-by-design principles for monitoring & alerting. Mentorship & Collaboration Ability to mentor junior engineers and work cross-functionally with SREs, application teams, and network engineers. High Availability / Disaster Recovery: Experience with HA/DR and Migration
Qualifications
Job ID: 107296661
Skills:
Grafana, Saas, POP ENVIRONMENT