Search by job, company or skills

Head of GPU Cluster Engineering

  • Posted 8 hours ago
  • Be among the first 10 applicants

Job Description

About Nava

Nava is building next-generation AI infrastructure and inference platforms designed to power the future of AI. We are looking for a Head of GPU Cluster Engineering to lead the architecture, deployment, and operations of large-scale GPU clusters.

This is a highly strategic leadership role responsible for defining the engineering vision, driving execution across multiple infrastructure domains, and ensuring our GPU platforms deliver world-class performance, scalability, and reliability.

What You'll Do

Platform Leadership

  • Own end-to-end engineering outcomes, from reference architecture through production operations.
  • Define the technical vision, architecture, and operating model for large-scale GPU infrastructure.
  • Establish engineering standards, design principles, and operational excellence across the platform.
  • Drive platform scalability, resiliency, and performance to support rapidly growing AI workloads.

Architecture & Technical Strategy

  • Lead the design and evolution of GPU cluster architectures across compute, networking, storage, orchestration, and observability.
  • Evaluate and drive technology decisions around GPUs, interconnects, storage, Kubernetes, scheduling, and AI infrastructure software.
  • Review and approve architecture decisions while balancing performance, reliability, cost, and operational complexity.
  • Drive infrastructure standardization and automation across deployments.

Cross-Functional Leadership

  • Lead execution across multiple engineering pillars:
    • GPU Compute
    • High-Speed Networking
    • Storage
    • Platform Engineering
    • Site Reliability Engineering (SRE)
    • Infrastructure Automation
  • Partner closely with Product, Supply Chain, Data Centre Operations, and Customer Success teams to ensure successful platform delivery.
  • Act as the technical escalation point for critical engineering decisions.
Delivery & Operational Excellence

  • Own engineering readiness gates from design through production deployment.
  • Drive release planning, operational reviews, risk assessments, and post-incident analysis.
  • Establish SLAs, SLOs, and operational metrics for platform health and reliability.
  • Champion automation, observability, incident management, and continuous improvement across the engineering organization.

Team Leadership

  • Build, mentor, and scale a high-performing engineering organization.
  • Develop technical leaders across infrastructure disciplines.
  • Foster a culture of engineering excellence, ownership, collaboration, and innovation.
  • Support hiring and talent development for critical infrastructure roles.

What We're Looking For

  • 12+ years of experience in infrastructure engineering, distributed systems, cloud platforms, or AI infrastructure.
  • Proven experience leading large-scale infrastructure or platform engineering teams.
  • Deep understanding of GPU clusters, AI infrastructure, HPC, or large-scale distributed systems.
  • Strong expertise across:
    • GPU Compute (NVIDIA ecosystem preferred)
    • High-performance networking (InfiniBand, RoCE, RDMA)
    • Kubernetes and container orchestration
    • Distributed storage systems
    • Infrastructure automation and observability
    • Linux systems and platform engineering
  • Experience designing highly available, scalable production infrastructure.
  • Strong architectural thinking with the ability to balance technical excellence and business priorities.
  • Excellent stakeholder management and cross-functional leadership skills.
Nice to Have

  • Experience building AI factories, GPU cloud platforms, or large-scale inference infrastructure.
  • Familiarity with CUDA, NCCL, Slurm, Ray, Kubeflow, or similar AI infrastructure technologies.
  • Experience working with hyperscalers, cloud providers, or AI-first technology companies.
  • Exposure to multi-region or global infrastructure deployments.

Why Join Nava

  • Build one of the world's leading AI infrastructure platforms.
  • Lead the engineering strategy behind next-generation GPU clusters and AI factories.
  • Work alongside world-class engineers solving some of the most challenging infrastructure problems in AI.
  • Shape the future of AI infrastructure from architecture through production at global scale.

Skills: leadership,automation,platforms,gpu,reliability,storage,architecture,cloud,infrastructure,cluster,drive

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152474509

Beware of Scammers

We don’t charge money for job offers