Search by job, company or skills

Sr. Principal Site Reliability Engineer

18-20 Years
Quick Apply
  • Posted 6 days ago
  • Over 50 applicants have applied

Job Description

What Youll Do:

 

Be the Force Behind Observability & Stability

  • Drive end-to-end Observability (Logs, Metrics, and Alerts) across our hybrid SaaS stack , spanning cloud, edge, and physical network devices.
  • Take ownership of Alerting strategy , cutting through noise while ensuring actionable, high-fidelity alerts.
  • Implement intelligent automation to reduce operational toil and enhance real-time visibility.

Own & Automate Operations

  • Design, build, and manage automation for self-healing infrastructure across cloud + global PoPs.
  • Develop automation for Kubernetes, ArgoCD, Helm Charts, Golang-based services, AWS, GCP, Terraform .
  • Improve networking observability , ensuring our routers, switches, and firewalls are monitored at scale.
  • Continuously eliminate manual ops work through automation and platform improvements.

Lead Incident Response & Operational Excellence

  • Participate in on-call rotations , ensuring rapid incident response across our cloud + edge stack.
  • Drive incident response automation , reducing MTTR and increasing system resilience .
  • Ensure security, compliance, and best practices in observability & automation .

Collaborate & Mentor

  • Work closely with application teams, network engineers, and SREs to improve reliability and performance.
  • Mentor junior engineers, fostering a culture of automation-first thinking and deep observability .

What Makes You a Great Fit

  • Deep expertise in Logs, Metrics, and Alerting, with a strong focus on Alerting automation.
  • Experience in hybrid SaaS environments spanning cloud-native and global infrastructure.
  • Strong background in Kubernetes, Infrastructure-as-Code (Terraform), Golang, AWS/GCP, and networking observability.
  • Proven track record of eliminating toil and improving operational efficiency through automation.
  • Passion for deep observability, networking-scale analytics, and automation at the edge.

Must-Have:

  • Observability & Alerting Expertise
  • Strong experience with Logs, Metrics, and Alerts, with a focus on highfidelity alerting and automation. Automation & Infrastructure as Code
  • Deep knowledge of Terraform, ArgoCD, Helm, Kubernetes, and Golang for automation. Cloud & Hybrid SaaS Experience
  • Handson experience managing cloudnative (AWS/GCP) and edge infrastructure. Incident Response & Reliability Engineering
  • Strong oncall experience, with a track record of reducing MTTR through automation Kubernetes Mastery
  • Handson experience deploying, managing, and troubleshooting Kubernetes in production environments.

Nice-to-Have:

  • Networking & Edge Observability - Familiarity with monitoring routers, switches, and firewalls in a global PoP environment
  • Data & Analytics in Observability - Experience with time-series databases (Prometheus, Grafana, OpenTelemetry, etc)
  • Security & Compliance Awareness - Understanding of secure-by-design principles for monitoring & alerting
  • Mentorship & Collaboration - Ability to mentor junior engineers and work cross-functionally with SREs, application teams, and network engineers
  • High Availability / Disaster Recovery: Experience with HA/DR and Migration

Qualifications

  • Typically, it requires at least 18 years of related experience with a bachelor s degree, 15 years and a master s degree, or a PhD with 12 years experience; or equivalent experience.
  • Excellent organizational agility and communication skills throughout the organization.

Job ID: 107297197

Beware of Scammers

We don’t charge money for job offers