Search by job, company or skills

Head of Platform Reliability

10-15 Years
  • Posted 10 hours ago
  • Be among the first 10 applicants

Job Description

About Nava

Nava is building next-generation AI infrastructure and inference platforms at global scale. We're looking for a Head of Platform Reliability to own the deployment, availability, and operational excellence of our AI infrastructure.

This leader will be responsible for ensuring that our GPU clusters, platform services, and infrastructure operate reliably in production, while continuously improving automation, incident response, and platform resilience.

What You'll Do

Platform Reliability & Operations

  • Own the day-to-day reliability, availability, and operational health of Nava's AI infrastructure platform.
  • Lead production operations across GPU clusters, networking, storage, orchestration, and platform services.
  • Define and drive Service Level Objectives (SLOs), Service Level Agreements (SLAs), and uptime targets.
  • Ensure production environments meet the highest standards of reliability, scalability, and operational excellence.

Deployment & Release Management

  • Own the deployment strategy for software, firmware, infrastructure updates, and platform releases.
  • Design and implement safe rollout mechanisms including phased deployments, canary releases, blue-green deployments, and rollback strategies.
  • Ensure production changes are executed with minimal customer impact and operational risk.
  • Establish deployment readiness reviews and operational change management processes.

Incident Management

  • Lead the incident response function for production systems.
  • Establish incident management processes, escalation frameworks, and post-incident reviews.
  • Drive Root Cause Analysis (RCA) and ensure corrective and preventive actions are implemented.
  • Build a culture of operational learning and continuous improvement.

Reliability Engineering

  • Identify recurring operational issues and drive long-term engineering fixes rather than temporary workarounds.
  • Improve system resilience through automation, observability, monitoring, alerting, and self-healing capabilities.
  • Partner with engineering teams to eliminate reliability bottlenecks and technical debt.
  • Drive capacity planning, performance optimization, and operational readiness.

Cross-Functional Leadership

  • Work closely with Platform Engineering, GPU Cluster Engineering, Networking, SRE, Product, and Customer Success teams.
  • Ensure operational requirements are embedded into system design from the earliest stages.
  • Drive operational excellence across multiple engineering functions.

Team Leadership

  • Build and lead a high-performing Platform Reliability and Site Reliability Engineering (SRE) organization.
  • Mentor engineering managers and technical leaders.
  • Foster a culture of accountability, ownership, operational discipline, and customer-first thinking.

What We're Looking For

  • 10–15 years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Operations, or Cloud Operations.
  • Proven experience leading large-scale production infrastructure teams.
  • Deep expertise in operating highly available distributed systems and cloud platforms.
  • Strong understanding of:
    • Kubernetes and container orchestration
    • Linux systems administration
    • Infrastructure automation
    • Monitoring and observability platforms
    • Incident management and production operations
    • High-availability architecture and disaster recovery
  • Experience managing large-scale production deployments and release management.
  • Strong knowledge of operational metrics including SLAs, SLOs, SLIs, MTTR, and incident response best practices.
  • Exceptional leadership, stakeholder management, and communication skills.
Nice to Have

  • Experience operating AI infrastructure, GPU clusters, or large-scale inference platforms.
  • Familiarity with NVIDIA GPU infrastructure, InfiniBand/RDMA networking, and distributed AI workloads.
  • Experience with GitOps, Infrastructure as Code (Terraform, Ansible), CI/CD pipelines, and production automation.
  • Exposure to cloud-native platforms, HPC environments, or hyperscale infrastructure.

Why Join Nava

  • Shape the future of AI infrastructure - Lead reliability and operational excellence for one of the world's most advanced AI platforms.
  • Scale with purpose - Build and grow production systems that power next-generation AI workloads for global customers.
  • Collaborate with world-class talent - Work alongside top-tier engineers, researchers, and operators solving some of the hardest infrastructure challenges in AI.
  • Build and lead at the frontier - Establish the operational foundations of Nava's global AI platform while growing a high-impact reliability engineering organization.

Skills: operational excellence,operations,drive,reliability engineering,management,reliability,infrastructure,automation,availability,platforms

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152344529

Beware of Scammers

We don’t charge money for job offers